Establishment method of skin melanoma standard data set
Through the improved Generative Adversarial Network (GAN) model, the problem of high-dimensional, non-random missing values in the skin melanoma dataset is solved, and the data is efficiently completed and quality improvement is achieved, supporting more accurate disease research and clinical decision-making.
Patent Information
- Application Number
- CN202510336945.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
The existing technology is difficult to effectively solve the problem of high-dimensional and non-random missing values in the skin melanoma data set, resulting in incomplete and inaccurate data, and lack of standard and mature data resource acquisition and integration technology.
An improved generative adversarial network (GAN) is introduced to build a data completion model, and dynamically generate fill values that conform to the real data distribution through adversarial learning to fill in missing parts in high-dimensional data.
It significantly improves the reliability of data quality and downstream knowledge maps, provides standardized, standardized, and continuously collected skin melanoma data resources, and supports research such as disease pattern recognition, risk assessment and treatment effect monitoring.
Smart Images

Figure CN120220940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of establishing a knowledge graph dataset. Background Art
[0002] Cutaneous melanoma is a malignant tumor caused by the malignant transformation of skin melanocytes, characterized by rapid growth rate, high malignancy, easy metastasis in the early stage, poor prognosis, etc. Its incidence shows a continuous growth trend in all regions of the world, posing a serious threat to the life and health of patients. Therefore, early diagnosis and early treatment are ideal strategies to improve the prognosis of cutaneous melanoma. Establishing a standardized cutaneous melanoma dataset is of great significance for research in aspects such as disease pattern recognition, risk assessment, and treatment effect monitoring. However, currently, medical data is incomplete, non-standard, and inaccurate, lacking standards and mature technical routes for data resource collection and integration.
[0003] In traditional cutaneous melanoma research and clinical practice, data acquisition and management face many difficulties, and there is no standardized cutaneous melanoma dataset. On the one hand, data sources are extremely scattered, and different medical institutions use different electronic medical record systems, with greatly different data formats, recording methods, and standards. On the other hand, although the data sample size of cutaneous melanoma accumulated in clinical practice is of a certain scale, the data distribution is severely unbalanced. To solve these problems, it is necessary to construct a standardized, standard, and continuously collected cutaneous melanoma dataset with the cutaneous melanoma knowledge graph as the core, providing comprehensive, accurate, efficient, and continuously updated data resources.
[0004] As a powerful deep learning model, the Generative Adversarial Network (GAN) shows great potential in the field of data generation. Through the adversarial training of the generator and discriminator, it can generate data samples similar to the real data distribution. Traditional missing value processing methods (such as mean filling, interpolation) are applicable to low-dimensional and randomly missing data, but have limited effects on high-dimensional and non-randomly missing complex data (such as gene mutation data, pathological examination data), and are prone to introducing biases. Therefore, this application constructs a data completion model by introducing an improved generative adversarial network. This model uses adversarial learning to dynamically generate filling values that conform to the real data distribution, significantly improving data quality and the reliability of the downstream knowledge graph. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method for establishing a standardized cutaneous melanoma dataset, and the steps of this method are as follows:
[0006] Step 1: Data collection: Obtain cutaneous melanoma sample data, including cutaneous melanoma clinical guidelines, relevant literature, expert consensus, and clinical case system data;
[0007] Step 2: Data preprocessing: Perform preliminary cleaning and formatting on the original data;
[0008] Step 3: Data completion: Through multiple rounds of adversarial training with a generative adversarial network model, generate filling values that highly match the characteristics of the original data and have a reasonable distribution, effectively filling the missing parts in high-dimensional data to make the data more complete and continuous;
[0009] Among them, the specific indicators of the clinical case system data in the data collection step include:
[0010] (1) Basic patient information: Name, ID number, age, gender, ethnicity, place of origin, current residence address, contact phone number, marital status, admission date, discharge date, length of hospital stay;
[0011] (2) Medical history: Chief complaint (main symptoms, duration, accompanying symptoms), current medical history (onset time, incentive, rash location, rash nature, rash size, whether the boundary is clear, presence of ulceration, presence of itching and pain, treatment process), past medical history (past disease history, medication history, surgical history, surgical time, surgical site, surgical name, radiotherapy and chemotherapy history, blood transfusion history, trauma history), personal history (smoking history, daily smoking amount, whether smoking cessation, smoking cessation time limit, drinking history, drinking years, average daily drinking amount (g / day), whether alcohol abstinence, alcohol abstinence time limit), family history (skin tumor history, genetic disease history), reproductive and menstrual history;
[0012] (3) Physical examination: Body temperature (°C), respiration (times / min), pulse (times / min), heart rate (times / min), systolic blood pressure (mmHg), diastolic blood pressure (mmHg), weight (kg), height (cm), skin lesion morphology, color, long and short axes, boundary, ulcer, infection, lymph node examination;
[0013] (4) Dermoscopy: Blue-white structure, morphological asymmetry, irregular boundary, uneven color, atypical pigment network, irregular vascular structure, erosion and ulcer;
[0014] (5) Pathological examination: Pathological type, diameter, presence of ulcer, Breslow thickness, Clark grade, infiltration of subcutaneous fat / muscle / bone, resection margin, inflammatory infiltration, immunohistochemical results, BRAFV600E mutation;
[0015] (6) Laboratory examination: Blood type, blood and urine routine, coagulation examination, biochemical examination (liver function, kidney function, blood lipid, ions, blood glucose), immune routine, tumor markers;
[0016] (7) Imaging examination: Lymph node ultrasound examination, abdominal ultrasound examination, head CT, chest CT, abdominal CT, PET-CT;
[0017] (8) Surgical treatment: surgical procedure, surgical grading, anesthesia method, surgical resection scope, defect size, resection depth;
[0018] (9) Drug treatment: targeted / chemotherapy drug names, usage of targeted / chemotherapy drugs / interferon, treatment time, treatment cycle;
[0019] (10) Radiation therapy: treatment dose, treatment frequency, treatment time;
[0020] (11) Prognosis: surgical wound healing, complications, functional impairment, adverse reactions of drugs and radiation therapy;
[0021] (12) Follow-up: disease progression, recurrence, metastasis, death.
[0022] In the above data preprocessing step, first, the original data is marked with missing value masks, and a mask matrix is used to identify the missing positions. In the mask matrix, 1 represents non-missing, and 0 represents missing. Then, for numerical features, they are uniformly normalized to the interval [0, 1]. For categorical and text features, they are uniformly encoded using the word vector embedding method to form preprocessed data.
[0023] In the above data completion step, the generative adversarial network model includes:
[0024] 1. Generator: used to take a random noise vector as input, and through a pre-constructed complex neural network structure, perform layer-by-layer transformation and combination on the input information, simulate the distribution characteristics of real skin melanoma data, and generate synthetic data that is similar to real skin melanoma data in appearance, features, and data structure.
[0025] 2. Discriminator: used to extract and analyze features of the synthetic data output from the generator and real skin melanoma data through an internally carefully designed multi-layer neural network structure.
[0026] 3. Data fusion module: used to fuse the filled values generated by the generator with the real part in the original data. First, according to the missing value mask matrix, determine the positions that need to be filled. Then, fill the generated filled values into the corresponding missing positions in the original data according to the instructions of the mask matrix to obtain the completed data.
[0027] 4. Generator loss function: used to generate as realistic filled values as possible, making it difficult for the discriminator to distinguish between generated data and real data. Discriminator loss function: used to accurately distinguish between real data and generated data.
[0028] Advantages of the present invention:
[0029] To solve the problem of high-dimensional and non-random missing values in complex datasets of cutaneous melanoma, the present invention introduces an improved generative adversarial network (GAN) to construct a data completion model, and specifically optimizes the architectures and training mechanisms of its generator and discriminator. During the continuous iterative training process, the model can accurately identify the feature patterns in complex datasets, and then achieve precise completion of high-dimensional and non-random missing values, providing a reliable data basis for subsequent data analysis, model training and other tasks.
[0030] By introducing an improved generative adversarial network to construct a data completion model, the present invention can complete the missing values in multi-source and multi-dimensional data, thereby establishing a standardized, normalized, and continuously collected data resource for cutaneous melanoma, providing a reliable data resource for the research of this disease. Brief Description of the Drawings
[0031] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. They are used together with the embodiments of the present application to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps. In the drawings:
[0032] Figure 1 It is a block diagram of a method for establishing a standardized dataset of cutaneous melanoma according to an embodiment of the present application.
[0033] Figure 2 It is a flowchart of the generator in the data completion module in a method for establishing a standardized dataset of cutaneous melanoma according to an embodiment of the present application.
[0034] Figure 3 It is a flowchart of the data fusion module and the discriminator in the data completion module in a method for establishing a standardized dataset of cutaneous melanoma according to an embodiment of the present application. Detailed Description of the Embodiments
[0035] The embodiments of the present application will be described in more detail below with reference to the drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0036] It should be understood that the various steps described in the method embodiments of the present application can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.
[0037] As used herein, the term "comprising" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0038] It should be noted that the concepts such as "first", "second", etc. mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0039] It should be noted that the modifications of "one" and "plural" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0040] Cutaneous melanoma is a malignant tumor caused by the malignant transformation of cutaneous melanocytes, which is characterized by a fast growth rate, high malignancy, easy metastasis in the early stage, and poor prognosis. Although cutaneous melanoma accounts for only 10% of all skin cancers, it accounts for 80% of the mortality of skin cancers. It can occur in any part of the body skin, and the clinical manifestations are diverse. The typical manifestations often include uneven rash color, asymmetric shape, irregular boundary, etc. As the disease progresses, the tumor gradually enlarges, forming nodules or ulcers, and some patients are accompanied by symptoms such as pain, itching and bleeding. Due to the extremely aggressive nature and high fatality rate of cutaneous melanoma, early diagnosis and treatment are particularly important, and delaying the disease will increase the risk of death.
[0041] At present, the diagnosis and treatment models of cutaneous melanoma in various regions and hospitals at all levels in China are not yet unified, and there is a large heterogeneity. Although most hospitals have adopted the medical record information management system, it cannot meet the deeper scientific research requirements. Nowadays, with the help of medical big data, diseases can be analyzed in more detail and depth, enabling medical research to focus more precisely on major diseases with high incidence, high mortality and serious impact on human health. However, due to incomplete, non-standard, inaccurate medical data, lack of association, lack of standards for data resource collection and integration, and lack of mature technical routes, it is difficult to form high-quality big data for clinical reference and disease research.
[0042] Therefore, based on the above problems, the technical concept of this application is in the technical field of establishing a skin melanoma dataset. Specifically, it involves taking the skin melanoma knowledge graph as the core, introducing an improved generative adversarial network model, and establishing a standardized, normalized, and continuously collected skin melanoma data resource.
[0043] Figure 1 It is a block diagram of a method for establishing a skin melanoma standard dataset according to an embodiment of the present application. As Figure 1 shown, a method 100 for establishing a skin melanoma standard dataset according to an embodiment of the present application includes: a data acquisition module 110, configured to acquire skin melanoma data to obtain skin melanoma sample data; a data preprocessing module 120, configured to perform preliminary cleaning and formatting on the original data, providing a high-quality, standardized data foundation for subsequent data mining, model training, and knowledge graph construction, significantly improving the efficiency and reliability of the entire data processing process; a data completion module 130, configured to fill high-dimensional and non-random missing data. Through multiple rounds of adversarial training with a generative adversarial network, it generates filling values that highly match the characteristics of the original data and have a reasonable distribution, effectively filling the missing parts in high-dimensional data and making the data more complete and continuous; a knowledge graph construction module 140: configured to establish a knowledge system of skin melanoma, perform reasoning and mining on the existing data, and discover potential knowledge and rules, where the concept layer of the melanoma knowledge graph guides the relationships between datasets; the entity attributes of the knowledge graph guide the establishment of the attribute dimensions of the data layer; a model interpretability module 150, configured to improve the transparency and credibility of the entire data processing process, display the dynamic optimization of the model during the data completion process of different batches, help researchers evaluate the quality of data completion, enhance trust in the data processing results, and provide a reliable basis for subsequent data-based research, diagnosis, and treatment plan formulation.
[0044] In an embodiment of the present application, the data acquisition module 110 is configured to acquire skin melanoma data to obtain skin melanoma sample data. It should be understood that it shows different sources of skin melanoma data: skin melanoma treatment guidelines, expert consensus, relevant medical literature, clinical case management systems. At the same time, it also shows the rich dimensions of the collected data, including the patient's basic information, medical history, physical examination, dermoscopy, pathological examination, laboratory examination, imaging examination, surgical treatment, drug treatment, radiotherapy, prognosis, and follow-up records, etc. Through a rigorous and standardized data acquisition process, on the premise of obtaining relevant authorizations and permissions, it ensures the legality and reliability of the data source, and finally obtains sample data with high representativeness and research value.
[0045] In an embodiment of the present application, the data preprocessing module 120 is configured to perform preliminary cleaning and formatting on the original data.
[0046] (1) Mask and mark missing values based on the original data, and use a mask matrix to identify the missing positions. In the mask matrix, 1 indicates not missing, and 0 indicates missing. For example, for a two-dimensional data matrix X, create a mask matrix M of the same size. If X ij exists, then M ij = 1; if X ij is missing, then M ij = 0.
[0047] (2) For numerical data, uniformly normalize it to the interval [0, 1]. In this embodiment, the maximum-minimum normalization method is adopted to map the data to a specific interval and eliminate the influence of dimensions. The formula is:
[0048]
[0049] where x is the original data, X min is the minimum value, and X max is the maximum value; after processing, the data is in the interval [0, 1].
[0050] (3) For text data, use the word vector embedding method to uniformly encode it and convert the text into a numerical form; for missing values in text data, special markers can be filled (such as " <unk>” indicates unknown) or is filled in by speculation based on the context information of the text. In this embodiment, the filling method by speculation can be as follows:
[0051] 1. For time series data, such as January, February,?, April, May. It is speculated that the missing one in the middle is March;
[0052] 2. For example, if the surgical method column is missing, the surgical method can be speculated based on the preoperative diagnosis and pathological examination in the context;
[0053] 3. For example, if the treatment method is missing, based on the disease type, such as diabetes, referring to other diabetes cases, the treatment method may be insulin injection.
[0054] Figure 2 It is a flowchart of the data completion module in a method for establishing a standardized dataset of cutaneous melanoma according to an embodiment of the present application. It can be understood that first, the original data needs to go through the key link of data preprocessing. In this link, feature extraction is the first step. Through professional data mining algorithms and technologies, the most representative and essential feature information is accurately extracted from a large amount of original data, laying a solid foundation for subsequent data processing. The normalization operation follows immediately. It unifies data with different ranges and magnitudes to a standard scale, eliminates the influence of data dimension differences, and makes the data more fair and accurate in subsequent analysis and processing. At the same time, the missing values in the data are clearly marked so that they can be clearly identified and properly processed in the subsequent processing process.
[0055] The preprocessed data then enters the dataset division stage and is reasonably divided into a training set and a test set. Among them, the training set is used to train two core components, the generator and the discriminator. The generator takes a randomly generated noise vector as the initial input, combines with some real data selected from the training set, and through a carefully designed neural network structure, generates predicted values for filling in the missing values. The task of the discriminator is to distinguish the filling values generated by the generator from the real data in the training set. It also uses a complex neural network structure. Through feature extraction and pattern recognition of the input data, it outputs a probability value indicating the possibility that the input data is real data. The goal of the discriminator is to maximize the ability to distinguish real data from generated data, thereby prompting the generator to continuously improve the generation quality. The generator and the discriminator interact and co-evolve through an adversarial learning mechanism.
[0056] After the generator generates the filling values, these filling values will enter the data fusion module. In this module, through a reasonable data fusion algorithm design, the generated filling values are organically combined with the non-missing part of the original data. After data fusion, complete filled data is obtained. The filled data is evaluated through a series of scientific and rigorous evaluation metrics. These evaluation metrics cover multiple dimensions to comprehensively measure the quality of data filling.
[0057] If the evaluation result shows that the expected satisfactory standard is not reached, that is, the data filling effect is not ideal, then the model optimization process will be immediately started. The model will, based on the evaluation result, deeply analyze the problems and deficiencies existing in the generator and discriminator during the training process, adjust the model parameters and optimize the model structure accordingly, and then retrain the generator and discriminator. This process is repeated until the filled data reaches a satisfactory effect and can meet the requirements of various application scenarios such as subsequent data analysis and model training.
[0058] Specifically, in the embodiment of the present application, in the data filling module 130, an improved generative adversarial network model is constructed based on the preprocessed data. The model consists of a generator, a discriminator, and a data fusion module, and is used to fill high-dimensional and non-randomly missing data. Through multiple rounds of adversarial training of the generative adversarial network, filling values that highly match the characteristics of the original data and have a reasonable distribution are generated, effectively filling the missing part in the high-dimensional data and making the data more complete and continuous. It includes:
[0059] 1. Generator: It is used to take a random noise vector as input and, through a pre-constructed complex neural network structure, perform layer-by-layer transformation and combination on the input information. Simulate the distribution characteristics of real skin melanoma data and generate synthetic data that is similar to real skin melanoma data in appearance, features, and data structure.
[0060] Specifically, the generator adopts a fully connected neural network structure and combines a self-attention mechanism to enhance the global perception ability of the generator. It includes:
[0061] 1.1 Input layer
[0062] Concatenate the vector x containing missing values missing and the noise vector z as inputs, both with a dimension of n x . Among them, introducing the noise vector z can break through the limitations of the data. By injecting additional random noise, the potential distribution of the data can be fully explored. The calculation formula is:
[0063] X input =[z; x missing
[0064] 1.2 The first fully connected layer
[0065] Set the number of fully connected neurons in the first layer to 1024, and map the input vector to a hidden layer vector with a dimension of 1024. The calculation formula is:
[0066] h1 = LeakyReLU(W1 × X input + b1)
[0067] where W1 is a weight matrix with a dimension of 1024×2n x and b1 is a bias vector with a dimension of 1024. LeakyReLU is the activation function, and its formula is:
[0068]
[0069] 1.3 Second fully connected layer
[0070] Set the number of fully connected neurons in the second layer to 2048, and map the input vector to a hidden layer vector with a dimension of 2048. The calculation formula is:
[0071] h2 = LeakyReLU(W2 × h1 + b2)
[0072] where W2 is a learnable weight matrix of 2048×1024, and b2 is a bias vector with a dimension of 2048.
[0073] 1.4 Third fully connected layer
[0074] (1) Incorporate the attention mechanism to capture the importance between different features, generate an attention weight matrix, and use the weight matrix to perform weighted summation on the vector, thereby enhancing important features and suppressing unimportant features.
[0075] First, for the input feature vector X[h2], it is mapped into Query (Q), Key (K), and Value (V) through three linear transformations or fully connected layers respectively:
[0076] Q = XW Q
[0077] K = XW K
[0078] V = XW V
[0079] where W Q 、W K and W V are learnable weight matrices, W Q 、W K have a dimension of 128×2048, and W V has a dimension of 2048×2048; Q, K, and V represent the query matrix, key matrix, and value matrix respectively;
[0080] Next, calculate the attention score matrix. Its calculation is based on the dot product of the query matrix Q and the transpose of the key matrix K:
[0081]
[0082] where K T is the transpose of the key matrix K; QK T is the dot product of the query matrix Q and the transpose of the key matrix; to prevent the dot product value from being too large and causing unstable gradients, is used as a scaling factor to control the range of the dot product, d k is the dimension of h2;
[0083] Then, use the Softmax function to normalize the scaled attention scores to obtain the attention weight matrix (Attention Weights):
[0084] Attention Weights = Softmax(A)
[0085] The softmax function is calculated row by row to ensure that the sum of the weights in each row is 1, thus converting the attention score matrix into a probability distribution, enabling the model to more clearly represent the relative importance of each eigenvalue to other eigenvalues:
[0086]
[0087] where A ij represents the attention score of the i-th eigenvalue to the j-th eigenvalue; exp(A ij ) is the exponential function value of A ij , that is, e A ij ; ∑ n k=1 exp(A ik ) is the sum of the exponential values of all elements in the i-th row.
[0088] Finally, use the attention weights to perform a weighted sum on the value matrix V to obtain the final output feature Output:
[0089] Output = AttentionWeights × V
[0090] (2) Fully connected calculation
[0091] Set the number of neurons in the third fully connected layer to 4096, and map the input vector to a hidden layer vector with a dimension of 4096. The calculation formula is:
[0092] h3 = LeakyReLU(W3 × Output + b3)
[0093] Among them, W3 is a learnable weight matrix of 4096×2048, and b3 is a bias vector with a dimension of 4096.
[0094] 1.5 Output layer
[0095] The number of neurons in the output layer is the same as the dimension of the input layer, denoted as n. x . The calculation formula is as follows:
[0096] X generated = Sigmoid(W4×h3 + b4)
[0097] Among them, W4 is a learnable weight matrix of n x ×4096, and b3 is a bias vector with a dimension of n x . The Sigmoid activation function is used to map the output to between [0, 1], matching the range of the normalized data.
[0098] 2. Discriminator: Used to extract and analyze features of the synthetic data from the output of the generator and the real skin melanoma data through a carefully designed multi-layer neural network structure inside.
[0099] Specifically, the discriminator also adopts a fully connected neural network structure. It includes:
[0100] 2.1 Input layer
[0101] Its input is the original data (the part without missing values) X true and the filled value X generated generated by the generator, and the concatenated vector [X true ; X generated . The dimension of the concatenated vector is 2n x .
[0102] 2.2 First fully connected layer
[0103] The first fully connected layer maps the input vector to a hidden layer with a dimension of 1024. The calculation formula is:
[0104] l1 = LeakyReLU(W5×[X true ; X generated + b5)
[0105] Among them, W5 is a learnable weight matrix with a dimension of 1024×2n x , and b5 is a bias vector with a dimension of 1024, and LeakyReLU is the activation function.
[0106] 2.3 Second fully connected layer
[0107] The second fully connected layer maps l1 to a hidden layer vector with a dimension of 2048, and the calculation formula is:
[0108] l2 = LeakyReLU(W6 × l1 + b6)
[0109] Where W6 is a learnable weight matrix of 2048×1024, and b6 is a bias vector with a dimension of 2048.
[0110] 2.4 The third fully connected layer
[0111] The third fully connected layer maps l2 to a hidden layer vector with a dimension of 4096, and the calculation formula is:
[0112] l3 = LeakyReLU(W7 × l2 + b7)
[0113] Where W7 is a learnable weight matrix of 4096×2048, and b7 is a bias vector with a dimension of 4096.
[0114] 2.5 Output layer
[0115] The number of neurons in the output layer is 1, which is used to output a scalar value D([X true ; X generated ) to determine whether the input data is real data or generated data. Sigmoid is used as the activation function, and the calculation formula is:
[0116] D([X true ; X generated ) = Sigmoid(W8 × l3 + b8)
[0117] Where W8 is a learnable weight matrix of 1×4096, and b8 is a bias vector with a dimension of 1. The output value is between [0, 1]. The closer it is to 1, the more likely it is to be real data, and the closer it is to 0, the more likely it is to be generated data.
[0118] 3. Data fusion module:
[0119] The data fusion module is used to fuse the filled values generated by the generator with the real part in the original data. First, according to the missing value mask matrix, determine the positions that need to be filled. Then, fill the generated filled values into the corresponding missing positions of the original data according to the indication of the mask matrix to obtain the completed data.
[0120] 4. Loss function:
[0121] The loss function L G of the generator is defined as:
[0122] L G = -E Z [log(D(G(z; x_missing)))]
[0123] where E Z denotes expectation, z is the noise vector, x_missing is the vector with missing values, G(z, x_missing) is the filled value generated by the generator, and D(G(z, x_missing)) is the judgment result of the discriminator on the generated data.
[0124] The loss function L of the discriminator D is defined as:
[0125] L D = -E xture [log(D(x ture ))] - E Z [log(1 - D(G(z; x_missing)))]
[0126] where x true is the real data sampled from the real data distribution.
[0127] 5. Optimizer
[0128] The Adam optimizer is used to update the parameters of the generator and the discriminator. The initial learning rate is set to 1e-4, the L2-norm regularization weight decay value is 1e-5, the exponential decay rate of the first moment estimate is 0.9, and the exponential decay rate of the second moment estimate is 0.999.
[0129] Model training: The training set data is input into the model for training according to a certain batch size.
[0130] This is used to enable the model to learn the features and patterns of skin melanoma data in stages and efficiently, gradually converge in multiple iterations, thereby improving the overall performance and accurately simulating the real skin melanoma data distribution. In each round of training, first fix the parameters of the discriminator and update the parameters of the generator. By calculating the loss function of the generator and backpropagating, update the weights of the generator. Then fix the parameters of the generator and update the parameters of the discriminator. By calculating the loss function of the discriminator and backpropagating, update the weights of the discriminator. Alternate in this way until the generator and the discriminator reach a better balance state, that is, the filled value generated by the generator can better fit the real data distribution, and the discriminator is difficult to accurately distinguish between real data and generated data. Record the changes in the loss values of the generator and the discriminator during the training process to observe the training situation and convergence trend of the model.
[0131] Model evaluation metrics: Multiple evaluation metrics are used to evaluate the completed data, including important metrics such as accuracy, recall, F1-score, mean squared error (MSE), mean absolute error (MAE), etc. Among them, the mean squared error (MSE): It is used to measure the mean of the squared errors between the predicted values and the true values. The calculation formula is:
[0132]
[0133] where n is the number of samples, x i is the true value, is the predicted value (completed value). The smaller the MSE value, the smaller the error between the completed data and the true data, and the better the completion effect.
[0134] Mean absolute error (MAE): It is used to measure the mean of the absolute errors between the predicted values and the true values. The calculation formula is:
[0135]
[0136] The smaller the MAE value, the smaller the average error between the completed data and the true data, and the higher the prediction accuracy of the model.
[0137] Based on the above implementation process, the unbiasedness of the data is ensured.
[0138] Specifically, in the embodiment of the present application, in the model interpretability module 150, it is used to improve the transparency and credibility of the entire data processing process, display the dynamic optimization of the model during the data completion process of different batches, help researchers evaluate the data completion quality, enhance the trust in the data processing results, and provide a reliable basis for subsequent data-based research, diagnosis, and treatment plan formulation. It includes:
[0139] 1. Calculate the importance of the intermediate layer features of the generator: By analyzing the feature outputs of the intermediate layer of the generator, calculate the average value of the marginal contributions of the feature under all possible feature combinations to determine the contribution degree of each feature to the generation result. Using the feature attribution method, assign a Shapley value to each feature, and this value reflects the importance of the feature in the generation process. If the Shapley value of a certain feature is large, it indicates that it plays a key role in the process of generating the filled value and has a significant impact on the generation result; conversely, the feature with a smaller Shapley value has a relatively weak impact on the generation result. By analyzing the Shapley values of all features, we can clarify which features the generator mainly relies on when generating the filled value and which features play a relatively minor role, so as to deeply understand the internal working mechanism of the generator. The calculation formula is:
[0140]
[0141] Where N is the set of all features, φ(i) represents the Shapley value of feature i in set N, S is a subset of N that does not contain feature i, |S| represents the number of elements in subset S, and f is the prediction function of the original model.
[0142] 2. Visualize the discriminator decision boundary:
[0143] Using the visualization technique t-SNE (t-Distributed Stochastic Neighbor Embedding), map the input data (real data and generated data) of the discriminator to a low-dimensional space to visualize the decision boundary of the discriminator. By observing the shape and position of the decision boundary, we can understand how the discriminator distinguishes between real data and generated data, and further infer the differences and similarities between the filled values generated by the generator and the real data.
[0144] First, define a grid in the low-dimensional space. Each point y in the grid can be regarded as a test sample. For each point y in the grid, input it into the discriminator D to obtain the output D(y) of the discriminator, which represents the probability that the point is judged as real data.
[0145] The discriminator calculates the output range through the sigmoid function as [0,1]. Set a threshold t (usually take t = 0.5). When D(y) ≥ t, it is considered that the point belongs to the real data area; when D(y) < t, it is considered that the point belongs to the generated data area.
[0146] By discriminating all points in the grid, we can obtain a two-dimensional probability distribution Z, where Z m,n = D(y m,n ), y m,n is the point in the m-th row and n-th column of the grid. Then, use a contour plot to visualize this probability distribution, where the contour line with a probability value equal to the threshold t is the decision boundary of the discriminator.
[0147] The above description is only the preferred embodiments of the present application and the explanation of the applied technical principles. The description of the above embodiments is only used to help understand the method and concept of the present invention; at the same time, without departing from the scope and spirit of the described implementations, many modifications and changes are obvious to those of ordinary skill in the art. In summary, the content of this specification should not be construed as a limitation to the present invention.< / unk>
Claims
1. A method for establishing a skin melanoma normative dataset, characterized in that: The steps of this method are as follows: Step 1: Data collection: Obtain skin melanoma sample data; Step 2: Data preprocessing: Perform preliminary cleaning and formatting of the original data; first, perform mask marking of missing values based on the original data, and use a mask matrix to identify the missing positions; then normalize the numerical data, and use the word vector embedding method to uniformly encode the text data; for the missing values in the text data, fill in special marks or infer filling based on the context information of the text; Step 3: Data completion: Generate a generative adversarial network model through multiple rounds of adversarial training to generate filling values that are highly matched with the original data features and reasonably distributed, effectively filling in the missing parts of the high-dimensional data, making the data more complete and continuous; The generative adversarial network model consists of a generator, a discriminator and a data fusion module, wherein the generator consists of an output layer, three fully connected layers and an output layer, and an attention mechanism is introduced in the third fully connected layer to capture the importance of different features, enhance important features, and suppress unimportant features. A random noise vector is used as input, and the input information is transformed and combined layer by layer to simulate the distribution characteristics of real skin melanoma data, and generate synthetic data similar to real skin melanoma data in appearance, characteristics and data structure; the discriminator consists of an input layer, three fully connected layers and an output layer, and the synthetic data output from the generator and the real skin melanoma data are extracted by the discriminator to determine whether the data is real; the data fusion module determines the position to be filled according to the missing value mask matrix, and fills the filling value generated by the generator into the corresponding missing position of the original data to obtain the completed data.
2. The method for establishing a skin melanoma normative data set according to claim 1, characterized in that: The specific indicators of clinical case system data in the data collection step include: (1) Basic information of the patient: name, ID number, age, gender, ethnicity, place of origin, current address, contact number, marital status, date of admission, date of discharge, and number of days of hospitalization; (2) Medical history: chief complaint, current medical history, past medical history, personal history, family history, marital history, and menstrual history; (3) Physical examination: body temperature (°C), respiration (bpm), pulse (bpm), heart rate (bpm), systolic blood pressure (mmHg), diastolic blood pressure (mmHg), weight (kg), height (cm), morphology, color, length and short axis of skin lesions, borders, ulcers, infection, and lymph node examination; (4) Dermatoscopy: blue-white structure, asymmetric shape, irregular borders, uneven color, atypical pigment network, irregular vascular structure, and erosion ulcers; (5) Pathological examination: pathological classification, diameter, presence of ulcer, Breslow thickness, Clark grade, infiltration of subcutaneous fat / muscle / bone, resection margin, inflammatory infiltration, immunohistochemistry results, and BRAFV600E mutation; (6) Laboratory tests: blood type, routine blood and urine tests, coagulation tests, biochemical tests, routine immune tests, and tumor markers; (7) Imaging examinations: lymph node ultrasound, abdominal ultrasound, head CT, chest CT, abdominal CT, and PET-CT; (8) Surgical treatment: surgical procedure, surgical classification, anesthesia method, surgical incision range, defect size and resection depth; (9) Drug treatment: name of targeted / chemotherapeutic drug, usage of targeted / chemotherapeutic drug / interferon, treatment time and treatment cycle; (10) Radiotherapy: treatment dose, number of treatments and treatment duration; (11) Prognosis: surgical wound healing, complications, functional impairment, adverse reactions to drugs and radiotherapy; (12) Follow-up: disease progression, recurrence, metastasis and death.
3. The method for establishing a skin melanoma normative data set according to claim 1, characterized in that: The generator model contains: 1.1 Input Layer Concatenate vector x containing missing values missing and noise vector z as input, both with dimension n x ; The introduction of the noise vector z can break through the limitations of the data and fully explore the potential distribution of the data by injecting additional random noise; the calculation formula is: X input =[z;x missing ] 1.2 The first fully connected layer The number of fully connected neurons in the first layer is set to 1024, and the input vector is mapped to a hidden layer vector with a dimension of 1024; the calculation formula is: h1=LeakyReLU(W1×X input +b1) Where W1 is of dimension 1024×2n x The weight matrix, b1 is the bias vector with a dimension of 1024, and LeakyReLU is the activation function, whose formula is: 1.3 Second fully connected layer The number of fully connected neurons in the second layer is set to 2048, and the input vector is mapped to a hidden layer vector with a dimension of 2048; the calculation formula is: h2=LeakyReLU(W2×h1+b2) Where W2 is a 2048×1024 learnable weight matrix, and b2 is a bias vector with a dimension of 2048; 1.4 The third fully connected layer (1) Combined with the attention mechanism, it is used to capture the importance of different features, generate an attention weight matrix, and use the weight matrix to perform weighted summation on the vectors, thereby enhancing important features and suppressing unimportant features; First, for the input feature vector X[h2], it is mapped to Query (Q), Key (K), and Value (V) respectively through three linear transformations (or fully connected layers): Q=XW Q K=XW K V=XW V Among them, W Q , W K and W V is the learnable weight matrix, W Q , W K The dimension is 128×2048, W V The dimension is 2048×2048; Q, K, V represent the query matrix, key matrix, and value matrix respectively; Next, the attention score matrix is calculated; it is calculated based on the dot product of the query matrix Q and the transposed key matrix K: Among them, K T is the transpose of the key matrix K; QK T is the dot product of the query matrix Q and the transpose of the key matrix; in order to prevent the dot product value from being too large and causing gradient instability, we use As a scaling factor, used to control the range of the dot product, d k is the dimension of h2; Then, the scaled attention scores are normalized using the Softmax function to obtain the attention weights matrix: Attention Weights=Softmax(A) The softmax function is calculated row by row to ensure that the sum of the weights of each row is 1, thereby converting the attention score matrix into a probability distribution, allowing the model to more clearly represent the relative importance of each eigenvalue to other eigenvalues: Among them, A ij represents the attention score of the i-th eigenvalue to the j-th eigenvalue; exp(A ij ) is A ij The exponential function value, that is, e A ij ; ∑ n k=1 exp(A ik ) is the sum of the exponential values of all elements in the i-th row; Finally, the attention weights are used to perform a weighted sum on the value matrix V to obtain the final output feature Output: Output = AttentionWeights × V (2) Fully connected computing The number of fully connected neurons in the third layer is set to 4096, and the input vector is mapped to a hidden layer vector with a dimension of 4096; the calculation formula is: h3=LeakyReLU(W3×Output+b3) Where W3 is a 4096×2048 learnable weight matrix, and b3 is a bias vector with a dimension of 4096; 1.5 Output Layer The number of neurons in the output layer is the same as the dimension of the input layer, denoted as n x ; The calculation formula is as follows: X generated =Sigmoid(W4×h3+b4) Where W4 is n x ×4096 learnable weight matrix, b3 is a matrix with dimension n x The bias vector is used to map the output to [0,1] using the Sigmoid activation function, matching the normalized data range.
4. The method for establishing a skin melanoma normative dataset according to claim 1, characterized in that: The discriminator model structure is as follows: 2.1 Input Layer Its input is the original data (without missing values) X true and the padding value X generated by the generator generated The concatenated vector [X true ;X generated ], the vector dimension after concatenation is 2n x ; 2.2 First fully connected layer The first fully connected layer maps the input vector to a hidden layer with a dimension of 1024; The calculation formula is: l1=LeakyReLU(W5×[X true ;X generated ]+b5) Where W5 is of dimension 1024×2n x The learnable weight matrix, b5 is a bias vector with a dimension of 1024, and LeakyReLU is the activation function; 2.3 Second fully connected layer The second fully connected layer maps l1 to a hidden layer vector of dimension 2048, calculated as: l2=LeakyReLU(W6×l1+b6) Where W6 is a learnable weight matrix of 2048×1024, and b6 is a bias vector of dimension 2048; 2.4 The third fully connected layer The third fully connected layer maps l2 to a hidden layer vector of dimension 4096, and the calculation formula is: l3=LeakyReLU(W7×l2+b7) Where W7 is a 4096×2048 learnable weight matrix, and b7 is a bias vector with a dimension of 4096; 2.5 Output Layer The number of neurons in the output layer is 1, which is used to output a scalar value D([X true ;X generated ]) to determine whether the input data is real data or generated data; Sigmoid is used as the activation function, and the calculation formula is: D([X true ;X generated ])=Sigmoid(W8×l3+b8) Where W8 is a 1×4096 learnable weight matrix, b8 is a bias vector with a dimension of 1; the output value is between [0,1], the closer to 1, the more likely it is real data, and the closer to 0, the more likely it is generated data.
5. The method for establishing a skin melanoma normative data set according to claim 1, characterized in that: The data fusion module first determines the positions that need to be filled according to the missing value mask matrix; then, the generated filling values are filled into the corresponding missing positions of the original data according to the instructions of the mask matrix to obtain the completed data.
6. The method for establishing a skin melanoma normative dataset according to claim 1, characterized in that: The loss function L of the generator G Defined as: L G =-E Z [log(D(G(z;x_missing)))] Where E Z represents expectation, z is the noise vector, x_missing is the vector with missing values, G(z,x_missing) is the filling value generated by the generator, and D(G(z,x_missing)) is the judgment result of the discriminator on the generated data.
7. The method for establishing a skin melanoma normative data set according to claim 1, characterized in that: The loss function L of the discriminator D Defined as: L D =-E xture [log(D(x ture ))]- E Z [log(1-D(G(z;x_missing)))] where x true is real data sampled from the real data distribution.
8. The method for establishing a skin melanoma normative data set according to claim 1, characterized in that: The Adam optimizer is used to update the parameters of the generator and discriminator; the initial learning rate is set to 1e-4, the L2 norm regularization weight decay value is 1e-5, the exponential decay rate of the first-order moment estimate is 0.9, and the exponential decay rate of the second-order moment estimate is 0.999.