Missing data filling system based on set division and self-supervised learning

Through the methods of set division and self-supervised learning, the neural network model is trained by obtaining missing value blocks, which solves the problem that the missing value filling results are very different from the overall distribution characteristics of the original data in the existing technology, and achieves a higher filling accuracy.

CN120336733APending Publication Date: 2025-07-18SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510562887.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-05
Filing Date
2025-04-30
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the missing value filling method has insufficient representativeness of the local distribution characteristics, resulting in large differences between the filling results and the overall distribution characteristics of the original data, and the filling accuracy is low.

Method used

Through the method of set division and self-supervised learning, the missing value-free data blocks of the original data are obtained, the neural network model is trained, and the missing value is filled with the distribution characteristics of the missing value-free data blocks, and the data distribution characteristics are continuously updated and accumulated during the gradual filling process to improve the filling accuracy.

Benefits of technology

The accuracy of missing value filling is improved, making the filling result closer to the overall distribution characteristics of the original data, and reducing the gap between the filling result and the original data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336733A_ABST
    Figure CN120336733A_ABST
Patent Text Reader

Abstract

The invention discloses a missing data filling system based on set division and self-supervised learning, and relates to the technical field of medical data processing, and the system comprises a first obtaining module which is used for obtaining original data, converting the original data into a matrix form, and obtaining first data; the second acquisition module is used for recombining the first data by taking a row as a unit, and obtaining a candidate data subset according to a row without a missing value in each combination result; or recombining the first data by taking a column as a unit, and obtaining a candidate data subset according to rows without missing values in each combination result; calculating the number of non-missing values in each candidate data subset to obtain an effective information amount; the third obtaining module is used for obtaining a first non-missing data block according to the candidate data subset corresponding to the maximum value of the effective information amount, and each numerical value in the first non-missing data block is a non-missing value; and the first filling module is used for filling the first data according to the first missing-free data block to obtain second data so as to improve the filling accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of medical data processing, specifically to a missing data filling system based on set partitioning and self-supervised learning. Background Art

[0002] In real-world research, it is inevitable that there are missing values in the original data for statistical analysis. For example, patients intentionally or unintentionally conceal data, the accuracy of data collection devices, historical limitations, etc. may lead to incomplete data collection, or data may be lost during data storage. When analyzing and researching the original data, the existence of missing values reduces the utilization rate of effective data and increases the design difficulty of data modeling.

[0003] Currently, there are methods for filling missing values in the prior art. A method and system for filling missing values based on deep learning with a patent application number of CN201710358297.2 divides the original data into a complete data subset and a missing data subset. After learning the data distribution characteristics of the complete data subset through a neural network model, the neural network model is used to fill the missing values in the missing data subset. However, when the number of true values in the complete data subset is too small, the data distribution characteristics of the complete data subset are difficult to represent the overall distribution characteristics of the original data, resulting in inaccurate filling results. Summary of the Invention

[0004] The present invention provides a missing data filling system based on set partitioning and self-supervised learning to solve the technical problem in the prior art that filling missing values according to the local distribution characteristics of the original data results in a large difference between the filling result and the overall distribution characteristics of the original data, thereby achieving an improvement in filling accuracy.

[0005] The present invention claims protection for a missing data filling system based on set partitioning and self-supervised learning, and the system includes: The first acquisition module: acquires the original data, where the original data includes missing values and non-missing values, and converts the original data into a matrix form to obtain the first data; The second acquisition module: recombines the first data by row, and obtains candidate data subsets according to the columns without missing values in each combination result; or recombines the first data by column, and obtains the candidate data subsets according to the rows without missing values in each combination result; calculates the number of non-missing values in each candidate data subset to obtain the effective information amount; The third acquisition module: obtains the corresponding first complete data block according to the candidate data subset corresponding to the maximum value of the effective information amount, and each value in the first complete data block is a non-missing value; The first filling module: Fill the first data according to the first complete data block to obtain second data.

[0006] In an embodiment of the present application, the first filling module further includes: Obtain a first incomplete data block according to the first complete data block. The first incomplete data block has missing values, and the first incomplete data block includes data blocks in the first data that are in the same row but different columns from the first complete data block, and / or data blocks in the first data that are in the same column but different rows from the first complete data block; Fill the first incomplete data block according to the first complete data block to obtain a second complete data block, and obtain third data according to the second complete data block; Determine whether the third data has missing values. If so, use the third data as the new first data, return to S2 and continue to execute until the third data has no missing values, and obtain the second data according to the filling values corresponding to all the missing values.

[0007] In an embodiment of the present application, the third obtaining module further includes: Calculate the similarity between each candidate data subset and the historical filling data, where the historical filling data includes all the first complete data blocks and all the second complete data blocks in the historical filling operations, and the calculation method of the similarity includes: ; where Similarity is the similarity, N is the effective information amount corresponding to the candidate data subset, and N1 is the number of values in the candidate data subset that belong to the historical filling data; If there is a similarity with a value less than 1, remove the candidate data subset with a similarity value of 1; Take the candidate data subset corresponding to the maximum value of the effective information amount among all the current candidate data subsets as the first complete data block.

[0008] In an embodiment of the present application, the third obtaining module further includes: If the number of candidate data subsets corresponding to the maximum value of the effective information amount among all the current candidate data subsets is greater than 1, and the value of each corresponding similarity is less than 1, then take the candidate data subset corresponding to the minimum value of the similarity among them as the first complete data block.

[0009] In an embodiment of the present application, the first defective data block further includes a first defective data sub-block and a second defective data sub-block. The first defective data sub-block includes data blocks in the first data that are in the same row but different columns from the first non-defective data block, and the second defective data sub-block includes data blocks in the first data that are in the same column but different rows from the first non-defective data block: In the process of filling the first defective data block according to the first non-defective data block, it further includes: Filling the first defective data sub-block according to the first non-defective data block to obtain a first non-defective data sub-block, and obtaining the corresponding third data according to the first non-defective data sub-block; Filling the second defective data sub-block according to the first non-defective data block to obtain a second non-defective data sub-block, and obtaining the corresponding third data according to the second non-defective data sub-block; Respectively update the historical filling records corresponding to each third data. The historical filling records include the filling rounds and the filling positions corresponding to each filling round; Respectively determine whether there are missing values in each third data. If there are missing values in the third data, then use the corresponding third data as the new first data, return to S2 and continue to execute until there are no missing values in all the third data. Obtain the second data according to the filling values corresponding to the maximum values of the filling rounds in different historical filling records at each filling position.

[0010] In an embodiment of the present application, the historical filling record further includes the ratio of the number of true values in the first non-defective data block corresponding to the filling value obtained in each filling round to the number of true values in the original data, to obtain a first numerical value; When the filling rounds of the same filling position in different historical filling records are the same, select the filling value with the largest first numerical value as the filling value for the corresponding filling position.

[0011] In an embodiment of the present application, the second acquisition module further includes: S21: Respectively take each row or column of the first data as a first combined sub-result, and calculate the effective information amount of each first combined sub-result; take the first combined sub-result corresponding to the maximum value of the effective information amount as the standard combined result; S22: Respectively merge the remaining each row or column with the standard combined result to obtain a second combined sub-result; S23: Calculate the effective information amount of each second combined sub-result; S24: Compare the effective information amount of the standard combination result with the effective information amount of each second combination sub-result. If the effective information amount of the standard combination result is greater than the effective information amount of each second combination sub-result, obtain the candidate data block according to the standard combination result; if there exists a second combination sub-result whose effective information amount is not less than the effective information amount of the standard combination result, take the data block corresponding to the maximum value of the effective information amounts among the standard combination result and the second combination sub-result as the new standard combination result, return to S22 for execution, until the number of rows or columns in the second combination sub-result is equal to the number of rows or columns of the first data, and obtain the corresponding candidate data block according to the data block corresponding to the maximum value of the effective information amounts among the standard combination result and the second combination sub-result.

[0012] In an embodiment of the present application, the first filling module further includes: Set the values at random positions in the first non-defective data block as missing values multiple times to obtain corresponding multiple second defective data blocks; Obtain a training data set according to the second defective data blocks, and input each second defective data block in the training data set into a data filling model respectively to obtain corresponding predicted filling values; Calculate the loss values between all the predicted filling values and the corresponding true values in the first non-defective data block, and train the data filling model according to the loss values, and stop training when the loss value takes the minimum value; Input the first defective data block into the trained data filling model to obtain the second non-defective data block.

[0013] In an embodiment of the present application, the first filling module further includes: Perform random transformations on each second defective data block in the training data set, and the random transformation includes swapping random rows or columns in the same second defective data block.

[0014] In an embodiment of the present application, the missing rates of all the second defective data blocks are equal to the missing rate of the corresponding first defective data block, and the data filling model is constructed based on the Transformer model.

[0015] The present application has the following beneficial effects: 1. The overall distribution characteristics of the original data are affected by each true value. If the data used to train the neural network model contains more true values, the data distribution characteristics learned by the neural network model are closer to the overall distribution characteristics of the original data. The existence of missing values leads to uneven distribution of true values in the first data. Using the data block with missing values to train the neural network model will increase the design difficulty of the model. Therefore, the first complete data block corresponding to the maximum value of the effective data volume is selected to train the neural network model, so as to gather the scattered true values, so that the neural network model learns the data distribution characteristics closer to the overall distribution characteristics of the original data, thereby improving the filling accuracy.

[0016] 2. Each time, the data with the largest amount of effective information is selected from the first data as the first complete data block, and the neural network model is trained according to the first complete data block. The trained neural network model fills the missing values according to the learned data distribution characteristics, so that the filled values conform to the distribution characteristics of the first complete data block, that is, the local distribution characteristics of the original data. As the number of execution times of the filling step increases, the number of true values in the first complete data block gradually increases, and the data distribution characteristics learned by the neural network model gradually accumulate, thereby reducing the gap between the filling result and the overall distribution characteristics of the original data and improving the filling accuracy.

[0017] 3. During the step-by-step filling process, since the neural network model fills the missing values in the first defective data block based on the data distribution characteristics of the first complete data block and using the non-missing values in the first defective data block as the filling basis, the filling result conforms to the data distribution characteristics of the first complete data block. In this embodiment, each time of filling, the true values not used in the historical filling operations are incorporated into the first complete data block, that is, the true values that do not belong to the historical filling data. The data distribution characteristics learned by the neural network model are added with new data distribution characteristics on the basis of the data distribution characteristics learned during the historical filling operations. Thus, the data distribution characteristics learned by the neural network model gradually accumulate with the increase of the filling times. For example, in the second filling operation, samples 1, 2, 4, and 5 are selected as the first complete data block for the candidate data subsets corresponding to feature items 1, 3, and 4 respectively. The values of samples 2, 4, and 5 corresponding to feature items 1, 3, and 4 respectively conform to the data distribution characteristics of the first complete data block during the first filling operation, while the values of sample 1 corresponding to feature items 1, 3, and 4 respectively do not belong to the first complete data block and are not filled according to the data distribution characteristics of the first complete data block. Therefore, the values of sample 1 corresponding to feature items 1, 3, and 4 respectively belong to new data distribution characteristics, and the data distribution characteristics learned by the current neural network model accumulate new data distribution characteristics on the basis of the first time. Until all true values are used for training the neural network model and / or for serving as the basis for filling missing values, that is, when the values of all the similarities are 1, directly take the candidate data subset with the largest effective information amount as the first complete data block for training the neural network model. This first complete data block contains the distribution characteristics of part or all of the first complete data blocks in the historical filling operations. The filling result obtained by filling according to this first complete data block is closer to the overall distribution characteristics of the original data, thereby achieving the improvement of filling accuracy.

[0018] 4. Select the candidate data subset with the similarity not equal to 1 and the smallest similarity as the first complete data block from all the candidate data subsets in the order of the effective information amount from large to small. When more true values not used in the historical filling operations are involved, the corresponding value of the similarity is smaller, and the difference between the data distribution characteristics learned by the current neural network model and the data distribution characteristics learned by the neural network model during the historical filling operations is greater. The neural network model can learn all the local distribution characteristics scattered in the original data faster.

[0019] 5. Since different first data blocks with missing values are selected for filling in each filling operation, as the number of filling times increases, new true values will be added each time of filling. As a result, when filling the same missing value through different filling paths, the data features learned by the neural network model are different. For example, according to the first data shown in Table 1, the first complete data block shown in Table 3 is obtained. The data shown in Table 4 is selected as the first data sub-block with missing values for the first filling. When filling for the second time, according to the candidate data subset cd 2,14 fill the missing values in feature item 2 and feature item 5 of sample 1, sample 2, sample 4 and sample 5 respectively. That is, the filled value of the missing value corresponding to feature item 2 of sample 1 conforms to the data distribution characteristics of the candidate data subset cd 2,14 after filling. The data distribution characteristics are based on the distribution characteristics of the candidate data subset cd 1,22 at the first filling and new true values are added. According to the first data shown in Table 1, the first complete data block shown in Table 3 is obtained. The values corresponding to feature item 1, feature item 2 and feature item 3 of sample 1 and sample 3 are selected as the second data block with missing values for filling. The filled value of the missing value corresponding to feature item 2 of sample 1 only conforms to the distribution characteristics of the candidate data subset cd 1,22 After filling, the final filling result is obtained according to the maximum filling round corresponding to each filled value in the different third data without missing values, so that the overall filling result is closer to the overall distribution characteristics of the original data.

[0020] 6. When selecting a filled value for each missing position from multiple third data, when the filling rounds of the filled values obtained at the same filling position are the same, since in the filling operation, the larger the ratio of the number of true values in the first complete data block for training the data filling model to the number of true values in the original data, the closer the data distribution characteristics learned by the data filling model are to the overall distribution characteristics of the original data, and the obtained filled value is also more in line with the overall data distribution of the original data.

[0021] 7. The data filling model constructed according to the Transformer model can perceive the sequential characteristics of natural language data, enabling the model to not only learn the association relationships between different structured feature items and fill the missing values of structured feature items, but also learn the association relationships between structured feature items and natural language feature items, so as to fill the missing natural language feature items according to the known structured feature items and fill the missing non-natural language feature items according to the known natural language feature items. Brief Description of the Drawings

[0022] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, used to explain the principles of the present disclosure.

[0023] Figure 1 It is a schematic diagram of data processing in each stage of the missing data filling system related to the embodiments of the present application; Figure 2 It is a schematic diagram of data conversion in each stage of the missing data filling system related to the embodiments of the present application; Figure 3 It is an implementation manner of the missing data conversion in each stage related to the present application; Figure 4 It is a schematic diagram of the training of the data filling model related to the embodiments of the present application; Figure 5 It is another implementation manner of the missing data conversion in each stage related to the present application; Figure 6 It is a schematic diagram of the structure of the data filling model related to the embodiments of the present application; Figure 7 It is an overall structure diagram of the missing data filling system related to the embodiments of the present application; Figure 8 It is a schematic diagram of the structure of the electronic device related to the embodiments of the present application; Reference numeral in the figure: N - the number of hidden layers. Detailed implementation manners

[0024] To make the above objects, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, with reference to terms such as "one embodiment", "some embodiments", "implementation manners", "embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc., the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents that the specific features, structures, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0025] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, relational terms such as "first", "second", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0026] The present invention claims a missing data filling system based on set partitioning and self-supervised learning. Referring to the attached Figure 1 as shown, the system includes a module for performing the following steps: S1: Obtain the original data. The original data includes missing values and non-missing values. Convert the original data into matrix form to obtain the first data, denoted as S 1 : ; i = 1, 2, 3, ……, n; j = 1, 2, 3, ……, m; wherein, is the value at the i-th row and j-th column of the first data.

[0027] It should be noted that the data types of the data in the original data include continuous numerical values, discrete numerical values, and strings, but are not limited thereto. The missing value is the value to be filled in the original data. The non-missing values include true values, and the true values are the known data in the original data.

[0028] It should be noted that in step S1, preprocessing of the original data is also included. The preprocessing includes converting each non-numerical data in the original data into a continuous numerical value or a discrete numerical value. For example, converting the data corresponding to the feature item "whether suffering from diabetes" into 1 (representing not suffering from diabetes) and 0 (representing suffering from diabetes), that is, obtaining the corresponding . The above is only for illustration to assist in understanding the technical solution of this embodiment, and does not further limit the protection scope of the present application.

[0029] It should be noted that the preprocessing also includes converting the non-missing values into a preset value, denoted as U. The value of the preset value is preset according to the true values in the original data, and the preset value is used to distinguish the missing values and the non-missing values.

[0030] Among them, the first data includes the missing values and the non-missing values. The non-missing values are the known numerical values in the first data, including the true values and the filled values obtained by filling the missing values according to the true values.

[0031] In this embodiment, the original data includes n patient samples and m feature items. Denote the data of the i-th patient sample under the j-th feature item. The first data obtained from the original data is shown in Table 1 below: Table 1: First data example table

[0032] S2: Recombine the first data in units of rows, and obtain a candidate data subset according to the corresponding columns and / or rows without missing values in each combination result, denoted as CD k , CD k = {cd k,1 , cd k,2 , …, cd k,h , …, cd k,M}, where cd k,h represents the h-th candidate data subset, and M is the number of candidate data subsets. Calculate the number of non-missing values in each candidate data subset to obtain the effective information volume.

[0033] It should be noted that by recombining the first data in units of rows, K kinds of combination results can be obtained. Each combination result is to select a rows from the first data. The calculation method of the value of K is as follows: ; Among them, is the combination number of taking a columns from n rows, and the value of a is less than n.

[0034] It should be noted that all the combination results can be traversed, and the columns containing missing values in each combination result can be deleted to obtain the candidate data subset CD. Alternatively, according to different decision methods, the combination results can be screened, and some combination results can be selected, and the columns containing missing values in each combination result can be deleted to obtain the candidate data subset CD. For example, from all the partitioning methods, the combination results corresponding to the rows retaining specific one or more patient samples can be screened out. The above is only for illustration and does not constitute a limitation on the protection scope of the present application.

[0035] In another feasible implementation manner, the first data can also be recombined in units of columns, and the candidate data subset is obtained according to the rows without missing values in each combination result. The specific implementation manner is similar to the above and will not be described in detail here.

[0036] In this implementation manner, all the combination results are traversed to obtain a combination result example as shown in Table 2: Table 2: Combination result example

[0037] It should be noted that taking the combination result with serial number 22 as an example, this combination result is the data blocks corresponding to Sample 2, Sample 4, and Sample 5 for each feature item. Since there are missing values in this combination result, that is, the values corresponding to feature item 4 for Sample 2 and Sample 5 are missing values, and the value corresponding to feature item 5 for Sample 4 is a missing value. Therefore, the columns corresponding to non-missing values in this combination result are feature items 1 - 3, that is, the number of non-missing values in the data blocks corresponding to Sample 2, Sample 4, and Sample 5 for feature items 1 - 3 is 9, that is, the candidate data subset cd k,22 The effective information amount of

[0038] S3: Obtain the corresponding first complete data block according to the candidate data subset corresponding to the maximum value of the effective information amount, and each value in the first complete data block is a non-missing value.

[0039] In this embodiment, according to Table 2, it can be known that the effective information amount corresponding to the combination result with serial number 22 is the maximum value. Then, select the data blocks corresponding to Sample 2, Sample 4, and Sample 5 for feature items 1 - 3 as the first complete data block, as shown in Table 3: Table 3: Example of the first complete data block

[0040] S4: Fill the first data according to the first complete data block to obtain the second data, and the second data is the data without missing values after the first data is filled.

[0041] In this embodiment, the data filling model can be trained according to the first complete data block, so that the data filling model learns the data distribution characteristics of the first complete data block. Input the first incomplete data block into the trained data filling model, and the data filling model fills the missing values in the first data according to the data distribution characteristics of the first complete data block to obtain the filling values corresponding to the missing values, and obtain the second data. Among them, the data filling model can be constructed according to a neural network model.

[0042] It should be noted that the overall distribution characteristics of the original data are affected by each true value. The more true values are included in the data used to train the neural network model, the closer the data distribution characteristics learned by the neural network model are to the overall distribution characteristics of the original data. The existence of missing values results in uneven distribution of true values in the first data. Using the data block with missing values to train the neural network model will increase the design difficulty of the model. Therefore, the first complete data block corresponding to the maximum value of the effective data volume is selected to train the neural network model, so as to gather the scattered true values, so that the neural network model can learn the data distribution characteristics closer to the overall distribution characteristics of the original data, thereby improving the filling accuracy.

[0043] In a feasible implementation, referring to the attached Figure 2 As shown, in step S4, it further includes: S41: Obtain a first data block with missing values according to the first complete data block. The first data block with missing values has missing values. The first data block with missing values includes the data blocks in the first data that are in the same row but different columns as the first complete data block. The data block is a matrix composed of the numerical values at multiple specified positions in the data.

[0044] In this implementation, the first data block with missing values includes the data blocks corresponding to sample 2, sample 4, and sample 5 in feature item 4 and feature item 5. Refer to Table 4: Table 4: Example of the first data block with missing values

[0045] In another feasible implementation, the first data block with missing values may include the data blocks in the first data that are in the same column but different rows as the first complete data block, that is, the data blocks corresponding to sample 1 and sample 3 in feature items 1-3. The specific implementation is similar to the above and will not be described in detail here.

[0046] S42: Fill the first data block with missing values according to the first complete data block to obtain a second complete data block, and obtain the third data according to the second complete data block.

[0047] In this implementation, the neural network model is trained according to the numerical values corresponding to the first complete data block. The first data block with missing values is input into the trained neural network model, and the second complete data block is output. The second complete data block is the data after the neural network model fills the missing values in the first data block with missing values. Fill the corresponding positions of the first data according to the filled values in the second complete data block to obtain the third data.

[0048] S43: Determine whether there are missing values in the third data. If so, use the third data as the new first data, return to S2 and continue to execute until there are no missing values in the third data, and obtain the second data according to the filling values corresponding to all the missing values.

[0049] In this embodiment, after the first filling, the data blocks corresponding to each feature item of sample 2, sample 4, and sample 5 are all non-missing values, and the third data is further obtained. The values corresponding to feature item 2 and feature item 5 of sample 1 and the values corresponding to feature item 1, feature item 3, and feature item 4 of sample 3 are all missing values. Therefore, use this third data as the new first data, re-divide according to the positions of the remaining missing values to obtain the new first non-missing data block and the new first missing data block, train a new neural network model again, and fill in the missing values. Until the value corresponding to each feature item of each sample is non-missing, the second data is obtained.

[0050] In this embodiment, CD k is the set of the candidate data subsets during the k-th filling operation, and cd k,h represents the h-th candidate data subset in the set of the candidate data subsets during the k-th filling operation.

[0051] It should be noted that each time, select the data with the largest effective information amount from the first data as the first non-missing data block, train the neural network model according to the first non-missing data block, and the trained neural network model fills in the missing values according to the learned data distribution characteristics, so that the filling values conform to the distribution characteristics of the first non-missing data block, that is, the local distribution characteristics of the original data. As the number of execution times of the filling step increases, the true values in the first non-missing data block gradually increase, and the data distribution characteristics learned by the neural network model gradually accumulate, thereby reducing the gap between the filling result and the overall distribution characteristics of the original data and improving the filling accuracy.

[0052] In a feasible embodiment, step S2 further includes: Calculate the similarity between each candidate data subset cd k,h and the historical filling data H, where the historical filling data H includes all the first non-missing data blocks and all the second non-missing data blocks in the first to k-1 filling operations, and the calculation method of the similarity includes: ; where Similarity is the similarity, N is the effective information amount corresponding to the candidate data subset, and N1 is the number of values in the candidate data subset that belong to the historical filling data H. If there is a similarity value less than 1 among the similarity values, remove the candidate data subset with a similarity value of 1. Take the candidate data subset corresponding to the maximum value of the effective information amount among all current candidate data subsets as the first complete data block.

[0053] It should be noted that the initial value of the historical filled data H is a null value. Starting from the completion of the first filling, each time a filling operation is completed, update the historical filled data, incorporate each value in the first complete data block corresponding to the current filling operation and each value in the second complete data block into the historical filled data H. Before each filling operation starts, calculate the similarity between each candidate data subset and the historical filled data. When there is a candidate data subset with a similarity value less than 1 among all the similarity values, after removing the candidate data subset with a similarity value of 1, take the candidate data subset corresponding to the maximum value of the effective information amount among the remaining candidate data subsets as the first complete data block. When the similarity values of all are 1, take the candidate data subset corresponding to the maximum value of the effective information amount among all candidate data subsets as the first complete data block.

[0054] In this embodiment, according to the first data shown in Table 1, the first complete data block shown in Table 3 is obtained. After selecting the first incomplete data block shown in Table 4 for filling, incorporate the values corresponding to Sample 2, Sample 4, and Sample 5 in each feature item into the historical filled data, that is: ; When filling for the second time, calculate the similarity between each candidate data subset and the historical filled data. For example, in the second filling operation, for all combination results obtained in the same way, for the candidate data subset cd corresponding to the values of Sample 2, Sample 4, and Sample 5 in each feature item 2,22 in terms of which all values are non-missing values, the corresponding effective information amount is 15. Since all values in this candidate data subset belong to the historical filled data, the corresponding similarity is 15 / 15 = 1. For the candidate data subset cd corresponding to Sample 1, Sample 2, Sample 4, and Sample 5 in Feature Item 1, Feature Item 3, and Feature Item 4 respectively 2,14In this case, the corresponding effective information amount is 12. Since the values of sample 1 in the candidate data subset corresponding to feature item 1, feature item 3, and feature item 4 do not belong to the historical filled data, and the rest belong to the historical filled data, the similarity between this candidate data subset and the historical filled data is 9 / 12 = 0.75. Moreover, it is the candidate data subset with the largest effective information amount among all candidate data subsets during the current filling operation. Therefore, during the second filling, this candidate data subset is used as the first non-missing data block.

[0055] It should be noted that during the progressive filling process, since the neural network model fills the missing values in the first missing data block based on the data distribution characteristics of the first non-missing data block and using the non-missing values in the first missing data block as the filling basis, the filling result conforms to the data distribution characteristics of the first non-missing data block. In this embodiment, each time during filling, the true values not used in the previous historical filling operation are incorporated into the first non-missing data block, that is, the true values that do not belong to the historical filled data. The data distribution characteristics learned by the neural network model are added with new data distribution characteristics on the basis of the data distribution characteristics learned during the historical filling operation, so that the data distribution characteristics learned by the neural network model gradually accumulate with the increase in the number of filling times. For example, during the second filling operation, the candidate data subsets corresponding to sample 1, sample 2, sample 4, and sample 5 in feature item 1, feature item 3, and feature item 4 are used as the first non-missing data block. The values of sample 2, sample 4, and sample 5 corresponding to feature item 1, feature item 3, and feature item 4 conform to the data distribution characteristics of the first non-missing data block during the first filling operation, while the values of sample 1 corresponding to feature item 1, feature item 3, and feature item 4 do not belong to the first non-missing data block and are not filled according to the data distribution characteristics of the first non-missing data block. Therefore, the values of sample 1 corresponding to feature item 1, feature item 3, and feature item 4 belong to new data distribution characteristics, and the data distribution characteristics learned by the current neural network model have accumulated new data distribution characteristics on the basis of the first time. Until all true values are used for training the neural network model and / or as the basis for filling missing values, that is, when the values of all similarities are 1, the candidate data subset with the largest effective information amount is directly used as the first non-missing data block for training the neural network model. This first non-missing data block contains the distribution characteristics of part or all of the first non-missing data blocks in the historical filling operation, and the filling result obtained according to this first non-missing data block is closer to the overall distribution characteristics of the original data, thereby achieving an improvement in filling accuracy.

[0056] In a feasible embodiment, the system further includes: If the number of the candidate data subsets corresponding to the maximum value of the effective information amount in all the current candidate data subsets is greater than 1, and the value of each corresponding similarity is less than 1, then the candidate data subset corresponding to the minimum value of the similarity is taken as the first complete data block.

[0057] It should be noted that the candidate data subset with the similarity not equal to 1 and the minimum similarity is selected as the first complete data block from all the candidate data subsets in the order of the effective information amount from large to small. When more true values unused in the historical filling operation are used, the value of the corresponding similarity is smaller, and the difference between the data distribution characteristics learned by the current neural network model and the data distribution characteristics learned by the neural network model in the historical filling operation is greater. The neural network model can learn all the local distribution characteristics scattered in the original data faster.

[0058] In a feasible implementation manner, referring to the attached Figure 3 As shown, the first incomplete data block further includes a first incomplete data sub-block and a second incomplete data sub-block. The first incomplete data sub-block includes the data blocks in the first data that are in the same row but different columns from the first complete data block, and the second incomplete data sub-block includes the data blocks in the first data that are in the same column but different rows from the first complete data block: In the process of filling the first incomplete data block according to the first complete data block, it further includes: Filling the first incomplete data sub-block according to the first complete data block to obtain a first complete data sub-block, and obtaining the corresponding third data according to the first complete data sub-block; Filling the second incomplete data sub-block according to the first complete data block to obtain a second complete data sub-block, and obtaining the corresponding third data according to the second complete data sub-block; Updating the historical filling records corresponding to each third data respectively. The historical filling records include the filling rounds and the filling positions corresponding to each filling round; Judging whether each third data has a missing value respectively. If the third data has a missing value, then the corresponding third data is used as the new first data, and return to S2 to continue to execute until all the third data have no missing values. The second data is obtained according to the filling value corresponding to the maximum value of the filling rounds in the different historical filling records at each filling position.

[0059] Wherein, the filling round is the serial number of the filling operation in the process of filling the first data.

[0060] In this embodiment, the first complete data block shown in Table 3 is obtained from the first data shown in Table 1. The data shown in Table 4 is selected as the first incomplete data sub-block for filling, and the corresponding third data is used as the new first data for continuous filling. When the filling is completed, the corresponding third data without missing values is obtained, denoted as ; The first complete data block shown in Table 3 is obtained from the first data shown in Table 1. The values corresponding to Feature Item 1, Feature Item 2, and Feature Item 3 of Sample 1 and Sample 3 are selected as the second incomplete data block for filling, and the corresponding third data is used as the new first data for continuous filling. When the filling is completed, the corresponding third data without missing values is obtained, denoted as , according to each filling value in and the maximum value of the corresponding filling round to obtain the final filling result. For example, in the third data the filling value corresponding to Feature Item 2 of Sample 1 is obtained during the first filling operation, that is, the corresponding filling round is 1. In the third data the filling value corresponding to Feature Item 2 of Sample 1 is obtained during the second filling operation, that is, the corresponding filling round is 2. Therefore, the filling value corresponding to Feature Item 2 of Sample 1 in the third data is used as the final filling result of this missing value.

[0061] It should be noted that since different first incomplete data blocks are selected for filling during each filling operation, as the number of filling times increases, new true values are added during each filling, resulting in different data features learned by the neural network model when filling the same missing value through different filling paths. For example, the first complete data block shown in Table 3 is obtained from the first data shown in Table 1. The data shown in Table 4 is selected as the first incomplete data sub-block for the first filling. During the second filling, according to the candidate data subset cd 2,14 the missing values of Sample 1, Sample 2, Sample 4, and Sample 5 in Feature Item 2 and Feature Item 5 are filled, that is, the filling value corresponding to the missing value of Feature Item 2 of Sample 1 conforms to the data distribution characteristics of the candidate data subset cd 2,14 on the basis of the distribution characteristics of the candidate data subset cd 1,22 during the first filling, new true values are added. And the first complete data block shown in Table 3 is obtained from the first data shown in Table 1. The values corresponding to Feature Item 1, Feature Item 2, and Feature Item 3 of Sample 1 and Sample 3 are selected as the second incomplete data block for filling. The filling value corresponding to the missing value of Feature Item 2 of Sample 1 only conforms to the candidate data subset cd 1,22According to the maximum value of the filling round corresponding to each filling value in different third data without missing values, the final filling result is obtained, making the overall filling result closer to the overall distribution characteristics of the original data.

[0062] In a feasible implementation manner, the historical filling record further includes the ratio of the number of true values in the first non-missing data block corresponding to the filling value obtained in each filling round to the number of true values in the original data, obtaining a first value; When the filling rounds of the same filling position in different historical filling records are the same, select the filling value with the largest first value as the filling value for the corresponding filling position.

[0063] In this implementation manner, taking the third data as an example, for the filling value of sample 1 in feature item 2, since this filling value is obtained by training the data filling model according to the candidate data subset cd 2,14 in the second filling round to obtain this filling value. In the candidate data subset cd 2,14 except that the values corresponding to sample 2 and sample 5 in feature item 1 are the filling values obtained in the first filling round, the rest are all true values, totaling 10, and the number of true values in the original data is 14. Then the value of the first value corresponding to this filling value is 10 / 14 = 0.714.

[0064] It should be noted that when selecting filling values for each missing position from multiple third data, when the filling rounds for obtaining the corresponding filling values at the same filling position are the same, since during the filling operation, the larger the ratio of the number of true values in the first non-missing data block for training the data filling model to the number of true values in the original data, the closer the data distribution characteristics learned by the data filling model are to the overall distribution characteristics of the original data, and the obtained filling value is also more in line with the overall data distribution of the original data.

[0065] In a feasible implementation manner, in S2, it further includes: S21: Respectively take each row or column of the first data as a first combined sub-result, and calculate the effective information amount of each first combined sub-result; take the first combined sub-result corresponding to the maximum value of the effective information amount as the standard combined result; S22: Respectively combine the remaining each row or column with the standard combined result to obtain a second combined sub-result; S23: Calculate the effective information amount of each second combined sub-result; S24: Compare the effective information amount of the standard combination result with the effective information amount of each second combined sub-result. If the effective information amount of the standard combination result is greater than the effective information amount of each second combined sub-result, obtain the candidate data block according to the standard combination result; if there exists a second combined sub-result whose effective information amount is not less than the effective information amount of the standard combination result, take the data block corresponding to the maximum value of the effective information amount among the standard combination result and the second combined sub-result as the new standard combination result, return to execute S22 until the number of rows or columns in the second combined sub-result is equal to the number of rows or columns of the first data, and obtain the corresponding candidate data block according to the data block corresponding to the maximum value of the effective information amount among the standard combination result and the second combined sub-result.

[0066] In this embodiment, according to the first data shown in Table 1, the effective information amount of the first row is 3, the effective information amount of the second row is 4, the effective information amount of the third row is 2, the effective information amount of the fourth row is 4, and the effective information amount of the fifth row is 4. Take the data block corresponding to the second row as the standard combination result; respectively combine the data blocks corresponding to the first row, the third row, the fourth row, and the fifth row with the standard combination result, and the effective information amounts corresponding to the obtained second combined sub-results are 4, 4, 6, and 8 respectively. Take the second combined sub-result corresponding to the fifth row and the second row as the new standard combination result; respectively combine the data blocks corresponding to the first row, the third row, and the fourth row with the standard combination result, and the effective information amounts corresponding to the obtained second combined sub-results are 6, 6, and 9 respectively. Take the second combined sub-result corresponding to the fourth row, the fifth row, and the second row as the new standard combination result; respectively combine the data blocks corresponding to the first row and the third row with the standard combination result, and the effective information amounts corresponding to the obtained second combined sub-results are 8 and 4 respectively. Since the effective information amount corresponding to the standard combination result at this time is 9, that is, greater than the effective information amount of each second combined sub-result, the corresponding candidate data subset is obtained according to the second combined sub-results of the second row, the fourth row, and the fifth row.

[0067] It should be noted that when the original data is large-sample data, traversing each combination result corresponding to the first data in a traversal manner will result in a large amount of calculation, which is not conducive to filling missing data based on big data. Each time, the first combined sub-result or the second sub-result with the largest effective information amount is selected as the standard combined result, and based on the standard combined result, the standard combined results with more rows or columns are obtained respectively with each other row, so as to achieve faster searching for the first complete data block with more effective information amount from all combination results and reduce the search amount.

[0068] In a feasible implementation manner, referring to the appendix Figure 4 as shown, in the process of filling the first data according to the first complete data block, it further includes: Setting the values at random positions in the first complete data block as missing values for multiple times to obtain corresponding multiple second incomplete data blocks; Obtaining a training data set according to the second incomplete data blocks, and respectively inputting each second incomplete data block in the training data set into a data filling model to obtain corresponding predicted filling values; Calculating the loss values between all the predicted filling values and the corresponding true values in the first complete data block, training the data filling model according to the loss values, and stopping training when the loss value takes the minimum value; Inputting the first incomplete data block into the trained data filling model to obtain the second complete data block.

[0069] In this implementation manner, the data filling model is constructed according to a neural network model, including an input layer, a hidden layer and an output layer. During the training process, the Relu function is used as the activation function, and the loss values between all the predicted filling values and the corresponding true values in the first complete data block are calculated according to the mean square error loss function. The specific calculation method of the mean square error loss function MSE is: ; wherein, is the number of the second incomplete data blocks in the training data set, is the number of missing values in the i-th second incomplete data block, y i,j is the true value corresponding to the j-th missing value in the i-th second incomplete data block, is the predicted filling value corresponding to the j-th missing value in the i-th second incomplete data block.

[0070] In a feasible implementation manner, before respectively inputting each second incomplete data block in the training data set into the data filling model, it further includes: Randomly transform each of the second missing data blocks in the training dataset. The random transformation includes randomly swapping rows or columns in the same second missing data block to increase the number of training samples in the training dataset and reduce the probability of overfitting in the data filling model.

[0071] In a feasible implementation, referring to the appendix Figure 5 As shown, the missing rate of all the second missing data blocks is equal to the missing rate of the corresponding first missing data block, and the data filling model is constructed based on the Transformer model.

[0072] In this implementation, among the operations of setting the values at random positions in the first complete data block to missing values multiple times, the system further includes calculating the missing rate of the first missing data block, and setting the data at random positions in the first complete data block to missing values multiple times according to the missing rate of the first missing data block.

[0073] It should be noted that the missing rate is the proportion of the number of missing values in the total amount of data in the corresponding data block. The total amount of data is obtained by multiplying the number of samples and the number of features in the corresponding data block. For example, the missing rate of the first missing data block is equal to the ratio of the number of missing values in the first missing data block to the total amount of data, and the missing rate of the second missing data block is equal to the ratio of the number of missing values in the second missing data block to the total amount of data.

[0074] It should be noted that when training the data filling model, randomly missing the first complete data block according to the missing rate of the first missing data block enables the data filling model to learn the actual missing situation in the first missing data block, thereby improving the generalization ability of the model.

[0075] In this implementation, referring to the appendix Figure 6 As shown, the data filling model includes an embedding layer, a number of cascaded hidden layers, and a multi-layer perceptron, and of course it can be more than this. Among them, the embedding layer is used to encode the data block with missing data and then input it into the hidden layer. The hidden layer is used to extract the feature vectors of the input data features. The hidden layer includes a number of Transformer layers. Each Transformer layer includes a multi-head self-attention layer and a position feed-forward layer. The multi-head self-attention layer is used to independently perform corresponding self-attention calculations on the received input data according to the number of self-attention mechanisms, and then aggregate and normalize the calculation results of all self-attention mechanisms and output them. The position feed-forward layer is used to extract the position features in the output results of the corresponding multi-head self-attention layer, and aggregate and normalize the corresponding position features and the output of the corresponding multi-head self-attention layer. The multi-layer perceptron is used to convert the output of the hidden layer into a prediction result.

[0076] It should be noted that the hidden layer of the data filling model encodes each data item in the data to be filled into a feature vector of a fixed length. Especially in the encoding process of natural language feature items, the positional encoding is added to the word vector as the feature vector of this data item, enabling the data filling model to perceive the sequential features of natural language data. As a result, the model can not only learn the correlation relationships between different structured feature items and fill in the missing values of structured feature items, but also learn the correlation relationships between structured feature items and natural language feature items, so as to fill in the missing natural language feature items based on known structured feature items and fill in the missing non-natural language feature items based on known natural language feature items.

[0077] In a feasible implementation, taking the experimental data used to early identify whether a patient has diabetic nephropathy as an example for detailed description.

[0078] It should be noted that diabetic nephropathy is one of the main microvascular complications of type 2 diabetes. It not only seriously affects the quality of life of patients, but also significantly increases the risk of cardiovascular diseases and mortality of patients. Early identification of diabetic nephropathy is of great significance for the formulation of intervention measures, delaying the progression of the disease and reducing the medical burden of patients. Clinically, the diagnosis of diabetic nephropathy mainly relies on indicators such as urinary albumin / creatinine ratio (UACR) and estimated glomerular filtration rate (eGFR), etc. However, these indicators have low sensitivity in the early stage of diabetic nephropathy and are easily affected by various other factors.

[0079] In this implementation, to establish an efficient and accurate diabetic nephropathy prediction model, the real data records of 36,384 patients in 3,726 feature items were collected from multiple medical institutions. These feature items include basic information, diagnosis information, doctor's order information, test and examination information, etc. Among them, the basic information includes weight, age, etc., the doctor's order information includes medication information, etc., and the test and examination information includes urine routine examination, glycated hemoglobin, urinary microprotein / creatinine ratio, and serum creatinine, etc.

[0080] It should be noted that the collected data is preprocessed. The preprocessing includes deleting the feature items with all missing values in the whole column in the dataset, deleting the feature items with a single category greater than 80%, and the feature items with a coefficient of variation less than 0.1. After preprocessing, the dataset to be filled contains 36,384 patient samples and 3,699 feature items. Due to differences in data storage methods and data collection standards among different medical institutions, patient release, incomplete doctor records, etc., there are 141 feature items with missing values in the dataset.

[0081] In this embodiment, during the process of filling in missing data, in order to compare the effects of different methods for filling in missing data, a first filling model is established according to the KNN algorithm, a second filling model is established according to the fixed-value filling method, and a third filling model is established according to the present invention to fill in the missing data. The corresponding filled data is used to train diabetes nephropathy prediction models respectively constructed according to the XGBoost model and the Logistic model in advance. The AUC is used to evaluate the performance of each diabetes nephropathy prediction model. The higher the index value, the better the performance of the prediction model, and the closer the corresponding data filling result is to the true data distribution.

[0082] It should be noted that the preprocessed data set is represented in matrix form. The rows of the matrix represent patient samples, that is, all the data in each row represent all the features of a patient, and the columns of the matrix represent feature items, that is, all the data in each column represent the values of all patients for the corresponding feature items.

[0083] It should be noted that to ensure the reliability of the experimental results, the missing data is filled 5 times using 3 filling methods respectively to train the corresponding diabetes nephropathy prediction models, and the average value of 5 experiments is taken when calculating the AUC. The experimental results are shown in Table 5, where the present system refers to the missing data filling system described in the present invention.

[0084] Table 5 Performance comparison of different diabetes nephropathy prediction models

[0085] It should be noted that in the diabetes nephropathy prediction models constructed based on two models, the missing data filling steps adopted by the model with the optimal AUC performance are all the missing data filling steps described in the present invention. That is, among all the diabetes nephropathy prediction models constructed based on the XGBoost model, the AUC value of the diabetes nephropathy prediction model trained according to the output of the third filling model is the largest. At the same time, among all the diabetes nephropathy prediction models constructed based on the Logistic model, the AUC value of the diabetes nephropathy prediction model trained according to the output of the third filling model is also the largest.

[0086] In this embodiment, compared with the existing missing data filling methods, the technical solution of the present invention separates the complete data block to be filled area from the original data, and at the same time randomly deletes the complete data block according to the data missing rate of the area to be filled, and then trains a data filling model constructed according to the Transformer model, so that the data distribution state of the data set used to train the data filling model is closer to the true missing data distribution state, thereby learning the association relationship between different feature items of patients more closely to the truth, and further improving the missing data filling effect.

[0087] In a feasible embodiment, refer to the appendixFigure 7 As shown in Figure 7 , the missing data filling system based on set partitioning and self-supervised learning protected by the present invention includes: The first acquisition module: acquires the original data, where the original data includes missing values and non-missing values, and converts the original data into matrix form to obtain the first data; The second acquisition module: recombinates the first data by row, and obtains the candidate data subsets according to the columns without missing values in each recombination result; or recombinates the first data by column, and obtains the candidate data subsets according to the rows without missing values in each recombination result; calculates the number of non-missing values in each candidate data subset to obtain the effective information amount; The third acquisition module: obtains the corresponding first complete data block according to the candidate data subset corresponding to the maximum value of the effective information amount, and each value in the first complete data block is a non-missing value; The first filling module: fills the first data according to the first complete data block to obtain the second data.

[0088] In a feasible implementation manner, the first filling module further includes: Obtains the first incomplete data block according to the first complete data block, where the first incomplete data block has missing values, and the first incomplete data block includes the data block in the first data that is in the same row but different columns from the first complete data block, and / or the data block in the first data that is in the same column but different rows from the first complete data block; Fills the first incomplete data block according to the first complete data block to obtain the second complete data block, and obtains the third data according to the second complete data block; Judges whether the third data has missing values. If so, takes the third data as the new first data, returns to S2 to continue execution until the third data has no missing values, and obtains the second data according to the filling values corresponding to all missing values.

[0089] In a feasible implementation manner, the third acquisition module further includes: Calculates the similarity between each candidate data subset and the historical filling data, where the historical filling data includes all the first complete data blocks and all the second complete data blocks in the historical filling operations, and the calculation method of the similarity includes: ; where Similarity is the similarity, N is the effective information amount corresponding to the candidate data subset, and N1 is the number of values in the candidate data subset that belong to the historical filling data; If there is a similarity value less than 1, remove the candidate data subset with a similarity value of 1. Take the candidate data subset corresponding to the maximum value of the effective information amount in all current candidate data subsets as the first complete data block.

[0090] In a feasible implementation, the third acquisition module further includes: If the number of candidate data subsets corresponding to the maximum value of the effective information amount in all current candidate data subsets is greater than 1, and the corresponding similarity value of each is less than 1, then take the candidate data subset corresponding to the minimum similarity value among them as the first complete data block.

[0091] In a feasible implementation, the first incomplete data block further includes a first incomplete data sub-block and a second incomplete data sub-block. The first incomplete data sub-block includes the data blocks in the first data that are in the same row but different columns from the first complete data block, and the second incomplete data sub-block includes the data blocks in the first data that are in the same column but different rows from the first complete data block: In the process of filling the first incomplete data block according to the first complete data block, it further includes: Fill the first incomplete data sub-block according to the first complete data block to obtain a first complete data sub-block, and obtain the corresponding third data according to the first complete data sub-block; Fill the second incomplete data sub-block according to the first complete data block to obtain a second complete data sub-block, and obtain the corresponding third data according to the second complete data sub-block; Update the historical filling records corresponding to each third data respectively. The historical filling record includes the filling round and the filling position corresponding to each filling round; Judge whether each third data has a missing value respectively. If the third data has a missing value, then take the corresponding third data as the new first data, return to S2 and continue to execute until all third data have no missing values, and obtain the second data according to the filling value corresponding to the maximum value of the filling round in different historical filling records at each filling position.

[0092] In a feasible implementation, the historical filling record further includes the ratio of the number of true values in the first complete data block corresponding to the filling value obtained in each filling round to the number of true values in the original data, to obtain a first numerical value; When the filling rounds of different historical filling records at the same filling position are the same, select the filling value with the largest first numerical value as the filling value for the corresponding filling position.

[0093] In a feasible implementation, the second acquisition module further includes: S21: Respectively take each row or column of the first data as a first combined sub-result, and calculate the effective information amount of each first combined sub-result; take the first combined sub-result corresponding to the maximum value of the effective information amount as the standard combined result; S22: Respectively combine the remaining each row or column with the standard combined result to obtain a second combined sub-result; S23: Calculate the effective information amount of each second combined sub-result; S24: Compare the effective information amount of the standard combined result with the effective information amount of each second combined sub-result. If the effective information amount of the standard combined result is greater than the effective information amount of each second combined sub-result, then obtain the candidate data block according to the standard combined result; if there is a second combined sub-result whose effective information amount is not less than the effective information amount of the standard combined result, then take the data block corresponding to the maximum value of the effective information amount among the standard combined result and the second combined sub-result as the new standard combined result, return to execute S22 until the number of rows or columns in the second combined sub-result is equal to the number of rows or columns of the first data, and obtain the corresponding candidate data block according to the data block corresponding to the maximum value of the effective information amount among the standard combined result and the second combined sub-result.

[0094] In a feasible implementation, the first filling module further includes: Set the values at random positions in the first complete data block as missing values multiple times to obtain corresponding multiple second incomplete data blocks; Obtain a training data set according to the second incomplete data blocks, and respectively input each second incomplete data block in the training data set into a data filling model to obtain corresponding predicted filling values; Calculate the loss values between all the predicted filling values and the corresponding true values in the first complete data block, and train the data filling model according to the loss values. Stop training when the loss value takes the minimum value; Input the first incomplete data block into the trained data filling model to obtain the second complete data block.

[0095] In a feasible implementation, the first filling module further includes: Randomly transform each second incomplete data block in the training data set, and the random transformation includes swapping random rows or columns in the same second incomplete data block.

[0096] In a feasible implementation, the missing rate of all the second defective data blocks is equal to the missing rate of the corresponding first defective data blocks, and the data filling model is constructed based on the Transformer model.

[0097] Referring to the attached Figure 8 As shown, an embodiment of the present application provides an electronic device, including: a processor and a memory. The processor and the memory are interconnected and communicate with each other through a communication bus and / or other forms of connection mechanisms (not shown). The memory stores a computer program executable by the processor. When the computing device runs, the processor executes the computer program to execute the system in any optional implementation manner of the above embodiments.

[0098] An embodiment of the present application provides a storage medium. When the computer program is executed by the processor, it executes the system in any optional implementation manner of the above embodiments. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Read-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0099] In the embodiments provided by the present application, it should be understood that the disclosed system can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the system or unit can be in an electrical, mechanical or other form.

[0100] In addition, the unit described as a separated component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0101] Furthermore, the functional modules in each embodiment of the present application may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0102] Flowcharts are used herein to illustrate the steps of the embodiments through the present disclosure. It should be understood that the steps before or after do not necessarily need to be carried out precisely in sequence. On the contrary, they can be carried out in reverse order or evaluated simultaneously. At the same time, other operations can also be added to these processes.

[0103] Unless otherwise defined, all terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this disclosure belongs. It should also be understood that terms such as those defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense unless explicitly defined as such herein.

[0104] The missing data filling system based on set partitioning and self-supervised learning provided above has been introduced in detail. Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only for the embodiments of the present application, and is only used to help understand the missing data filling system based on set partitioning and self-supervised learning of the present application, and does not limit the protection scope of the present application; at the same time, for those skilled in the art, various changes and modifications can be made to the present application. Any modification or equivalent replacement made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A missing data filling system based on set partitioning and self-supervised learning, characterized in that, The system includes: The first acquisition module: acquiring original data, where the original data includes missing values and non-missing values, converting the original data into matrix form to obtain the first data; The second acquisition module: recombining the first data by row, obtaining candidate data subsets according to the columns without missing values in each combination result; or recombining the first data by column, obtaining the candidate data subsets according to the rows without missing values in each combination result; calculating the number of non-missing values in each candidate data subset to obtain the effective information amount; The third acquisition module: obtaining the corresponding first non-missing data block according to the candidate data subset corresponding to the maximum value of the effective information amount, where each value in the first non-missing data block is a non-missing value; The first filling module: filling the first data according to the first non-missing data block to obtain the second data.

2. The system according to claim 1, characterized in that, The first filling module further includes: Obtaining a first data block with missing values according to the first non-missing data block, where the first data block with missing values has missing values, and the first data block with missing values includes the data block in the first data that is in the same row but different columns from the first non-missing data block, and / or the data block in the first data that is in the same column but different rows from the first non-missing data block; Filling the first data block with missing values according to the first non-missing data block to obtain a second non-missing data block, and obtaining the third data according to the second non-missing data block; Judging whether the third data has missing values. If so, taking the third data as the new first data, returning to S2 to continue execution until the third data has no missing values, and obtaining the second data according to the filling values corresponding to all missing values.

3. The system according to claim 2, wherein The third acquisition module further includes: Calculating the similarity between each candidate data subset and historical filling data, where the historical filling data includes all the first non-missing data blocks and all the second non-missing data blocks in historical filling operations, and the calculation method of the similarity includes: ; where Similarity is the similarity, N is the effective information amount corresponding to the candidate data subset, and N1 is the number of values in the candidate data subset that belong to the historical filling data; If there is a similarity with a value less than 1, removing the candidate data subset with a similarity value of 1; Taking the candidate data subset corresponding to the maximum value of the effective information amount among all current candidate data subsets as the first non-missing data block.

4. The system according to claim 3, wherein The third acquisition module further includes: If the number of candidate data subsets corresponding to the maximum value of the effective information amount among all current candidate data subsets is greater than 1 and the value of each corresponding similarity is less than 1, taking the candidate data subset corresponding to the minimum value of the similarity as the first non-missing data block.

5. The system according to claim 4, characterized in that, The first defective data block further includes a first defective data sub-block and a second defective data sub-block. The first defective data sub-block includes data blocks in the first data that are in the same row but different columns from the first non-defective data block. The second defective data sub-block includes data blocks in the first data that are in the same column but different rows from the first non-defective data block: In the process of filling the first defective data block according to the first non-defective data block, it further includes: Filling the first defective data sub-block according to the first non-defective data block to obtain a first non-defective data sub-block, and obtaining the corresponding third data according to the first non-defective data sub-block; Filling the second defective data sub-block according to the first non-defective data block to obtain a second non-defective data sub-block, and obtaining the corresponding third data according to the second non-defective data sub-block; Updating the historical filling records corresponding to each of the third data respectively. The historical filling records include the filling round and the filling position corresponding to each filling round; Judging whether there is a missing value in each of the third data respectively. If there is a missing value in the third data, then taking the corresponding third data as the new first data, returning to S2 to continue execution until there is no missing value in all the third data, and obtaining the second data according to the filling value corresponding to the maximum value of the filling round in different historical filling records at each filling position.

6. The system according to claim 4, characterized in that, The historical filling record further includes the ratio of the number of true values in the first non-defective data block corresponding to the filling value obtained in each filling round to the number of true values in the original data, to obtain a first numerical value; When the filling rounds in different historical filling records at the same filling position are the same, selecting the filling value with the largest first numerical value as the filling value for the corresponding filling position.

7. The system according to any one of claims 1-6, characterized in that, The second acquisition module further includes: S21: Respectively taking each row or column of the first data as a first combined sub-result, and calculating the effective information amount of each first combined sub-result; taking the first combined sub-result corresponding to the maximum value of the effective information amount as the standard combined result; S22: Respectively merging the remaining each row or column with the standard combined result to obtain a second combined sub-result; S23: Calculating the effective information amount of each second combined sub-result; S24: Compare the effective information amount of the standard combination result with that of each second combination sub-result. If the effective information amount of the standard combination result is greater than that of each second combination sub-result, obtain the candidate data block according to the standard combination result; if there exists a second combination sub-result whose effective information amount is not less than that of the standard combination result, take the data block corresponding to the maximum value of the effective information amounts among the standard combination result and the second combination sub-result as the new standard combination result, and return to S22 for execution until the number of rows or columns in the second combination sub-result is equal to the number of rows or columns of the first data, and obtain the corresponding candidate data block according to the data block corresponding to the maximum value of the effective information amounts among the standard combination result and the second combination sub-result.

8. The system according to claim 7, wherein The first filling module further includes: Set the values at random positions in the first complete data block as missing values multiple times to obtain corresponding multiple second incomplete data blocks; Obtain a training data set according to the second incomplete data blocks, and input each second incomplete data block in the training data set into a data filling model respectively to obtain corresponding predicted filling values; Calculate the loss values between all the predicted filling values and the corresponding true values in the first complete data block, and train the data filling model according to the loss values, and stop training when the loss value reaches the minimum; Input the first incomplete data block into the trained data filling model to obtain the second complete data block.

9. The system according to claim 8, characterized in that, The first filling module further includes: Perform random transformations on each second incomplete data block in the training data set, and the random transformations include swapping random rows or columns in the same second incomplete data block.

10. The system according to claim 8, wherein The missing rates of all the second incomplete data blocks are equal to the missing rate of the corresponding first incomplete data block, and the data filling model is constructed based on the Transformer model.

Citation Information

Patent Citations

  • A method and system for filling missing values ​​based on deep learning

    CN107273429B