A method, system, electronic device, and medium for symbol data quantization processing
The proposed method addresses the challenge of uncovering latent symbol data associations through data standardization, clustering, and a symbol data mapping model, enhancing quantization quality and supporting tasks like drug development and clinical diagnostics.
Patent Information
- Application Number
- CN202211365507.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-02
AI Technical Summary
The existing quantitative processing methods of symbol data cannot effectively mine potential correlations between symbol data features in the field of medical data processing, and there are problems of poor quantification quality and dimensional disasters.
After standardizing the scalar data, the symbolic data is classified and clustered, the Euclidean distance and Euclidean distance are calculated, and the symbolic data mapping model is used for quantization, including the combination processing of the generator and the calculation module, and the adversarial generation model is built for training to obtain the symbolic quantization results.
It has achieved efficient mining of potential connections between symbol data features, improved quantitative quality, and facilitated the application of downstream tasks such as drug research and development and clinical diagnosis.
Smart Images

Figure CN115858611B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to a symbolic data quantization processing method, system, electronic equipment and medium. Background Art
[0002] In big data mining, in addition to scalar data, the mined raw data often also contains some non-scalar data - symbolic data, such as the two opposing attributes "negative" and "positive" in medical standard table data, or symbolic identifiers such as "too low", "standard" and "too high" with continuous attributes. In addition, there is a special kind of symbolic data - discrete symbolic data, such as drug names, physiological mechanism entity nouns, etc. This kind of discrete symbolic data seems to have nothing to do with each other, but in fact has a certain potential correlation in the real world. In order to achieve the measurement of symbolic data in the raw data, and then obtain the potential correlation between the features of the raw data, it is usually necessary to quantify the symbolic data. In the prior art, the commonly used symbolic data quantization processing methods in the industry can be roughly divided into the following two types:
[0003] a. The assignment method, also known as the mapping method, is to sort the symbol data by size or time first, and then assign values to them with continuous numbers, such as OneHot Encoding. This method performs better on symbol data with continuous attributes.
[0004] b. Embedding method, which is suitable for quantifying some symbols with active, passive, opposing and other relationships. It is currently commonly used for text quantization, such as Word2Vec (text-to-text quantization), BERT and other algorithms.
[0005] However, in the process of using the above prior art, the inventors found that the prior art has at least the following problems:
[0006] In the field of medical data processing, the current existing symbolic data quantization processing methods cannot mine the potential correlation between symbolic data features. Specifically, in the process of symbolic quantization using the assignment method, it is impossible to reasonably quantize data with implicit meaning or potential correlation; and when quantizing symbolic data with uncertain magnitude, the quality of quantization cannot be guaranteed. In the process of symbolic quantization using the embedding method, when quantizing symbolic data, large dimensional information is often generated, which can easily cause the dimensional disaster of data. In addition, the interpretability of quantization using this method is extremely poor, and it is also impossible to mine the potential correlation between symbolic data features. Summary of the invention
[0007] The present invention aims to solve the above technical problems at least to a certain extent. The present invention provides a method, system, electronic device and medium for symbol data quantization processing.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] In a first aspect, a method for quantifying symbol data is provided, including:
[0010] Extracting original data from a database, where the original data includes scalar data and symbol data corresponding to the scalar data;
[0011] Performing standardization processing on the scalar data to obtain standardized data;
[0012] Classifying the symbol data to obtain multiple symbol data classes;
[0013] Performing clustering processing on the standardized data according to the symbol data classes to obtain multiple clustering results with the same number as the symbol data classes and multiple cluster centers corresponding to the multiple clustering results; wherein, any one of the clustering results includes multiple standardized data;
[0014] Calculating the Euclidean distance α between each of the standardized data and its corresponding cluster center;
[0015] Calculating the mean value of all the standardized data in the clustering results corresponding to each symbol data class, and then calculating the Euclidean distance β between each standardized data and the mean value corresponding to the symbol data class to which it belongs;
[0016] Inputting the standardized data, the Euclidean distance α, and the Euclidean distance β into a symbol data mapping model to obtain a symbol quantization result.
[0017] The present invention can facilitate the efficient discovery of potential connections between symbol data features, and is conducive to better serving downstream tasks such as drug research and development and clinical diagnosis. Specifically, in the implementation process of the present invention, after extracting original data from a database, standardization processing is performed on the scalar data in the original data to obtain standardized data; then, according to the symbol data classes, clustering processing is performed on the standardized data to obtain multiple clustering results with the same number as the symbol data classes and multiple cluster centers corresponding to the multiple clustering results; then, the Euclidean distance α between each of the standardized data and its corresponding cluster center is calculated, and then the symbol data is classified to obtain multiple symbol data classes, and the Euclidean distance β between each standardized data and the mean value corresponding to the symbol data class to which it belongs is calculated; finally, the standardized data, the Euclidean distance α, and the Euclidean distance β are input into a symbol data mapping model to obtain a symbol quantization result.
[0018] In a possible design, performing standardization processing on the scalar data includes:
[0019] Calculate the mean and standard deviation of the scalar data, and then obtain the standardized data corresponding to the scalar data according to the mean and standard deviation of the scalar data; wherein, the standardized data is:
[0020]
[0021] In the formula, x is the scalar data, μ is the mean of the scalar data, and σ is the standard deviation of the scalar data.
[0022] In a possible design, perform clustering processing on the standardized data according to the symbol data class to obtain a plurality of clustering results with the same number as the symbol data class and a plurality of cluster centers corresponding to the plurality of clustering results, including:
[0023] Randomly select a plurality of initial cluster centers with the same number as the symbol data class;
[0024] Successively calculate the distances of the standardized data to the plurality of initial cluster centers, and classify the current standardized data into the cluster of the initial cluster center with the smallest distance until all the standardized data are classified into the corresponding clusters of the initial cluster centers with the smallest distances;
[0025] Calculate the mean of all the standardized data in the cluster of each initial cluster center, and use this mean to replace the initial cluster center of the corresponding cluster to obtain a plurality of updated cluster centers;
[0026] Reset the updated initial cluster centers as the initial cluster centers, and successively calculate the distances of the standardized data to the plurality of initial cluster centers again until the change amount of the mean of all the standardized data in the cluster of each initial cluster center is less than the threshold. At this time, the cluster corresponding to the mean is the clustering result, and the mean at this time is the cluster center corresponding to the clustering result.
[0027] In a possible design, the symbol data mapping model includes a generator and a calculation module connected in sequence;
[0028] Input the standardized data, the Euclidean distance α, and the Euclidean distance β into the symbol data mapping model to obtain a symbol quantization result, including:
[0029] Input the standardized data into the generator, and input the Euclidean distance α and the Euclidean distance β into the calculation module;
[0030] The generator processes the standardized data to obtain an initial symbol quantization value, and then inputs the initial symbol quantization value into the calculation module;
[0031] The calculation module multiplies the initial symbol quantization value by the Euclidean distance α to obtain a multiplication result, and then adds the multiplication result to the Euclidean distance β to obtain the symbol quantization result of the symbol data.
[0032] In a possible design, the generator includes a first input layer, a first hidden layer, and a first output layer connected in sequence, and a first normalization layer and an activation layer are sequentially connected between the first input layer and the first hidden layer and between the first hidden layer and the first output layer; the generator further includes a random module and a heterogeneous random field module. The heterogeneous random field module has two input ends. The first hidden layer is connected to the first input end of the heterogeneous random field module, and the output end of the activation layer at the next level of the first hidden layer is connected to the second input end of the heterogeneous random field module through the random module; the output end of the first output layer and the output end of the heterogeneous random field module are connected to the calculation module.
[0033] In a possible design, the generator processes the normalized data to obtain the initial symbol quantization value, including:
[0034] The normalized data is input into the generator through the first input layer for processing, and a first generation result is obtained through the first output layer;
[0035] The result K output by the first hidden layer is input into the heterogeneous random field module from the first input end of the heterogeneous random field module. The result output by the activation layer at the next level of the first hidden layer is randomly shuffled through the random module to obtain a feature vector Q, which is then input into the heterogeneous random field module from the second input end of the heterogeneous random field module;
[0036] According to a preset Gaussian random matrix, a result V can be obtained, and the result V is input into the heterogeneous random field module from the second input end of the heterogeneous random field module;
[0037] The heterogeneous random field module combines and processes the result K, the feature vector Q, and the result V to obtain a second generation result;
[0038] The first generation result and the second generation result are added to obtain the initial symbol quantization value.
[0039] In a possible design, the steps for obtaining the symbol data mapping model are as follows:
[0040] Construct an initial adversarial generation model;
[0041] Extract a training data set from the database;
[0042] Use the training data set to train the initial adversarial generation model, and obtain the symbol data mapping model after training.
[0043] In a second aspect, a symbol data quantization processing system is provided for implementing the symbol data quantization processing method described in any one of the above; the symbol data quantization processing system includes:
[0044] A data extraction module for extracting raw data from a database, the raw data including scalar data and symbol data corresponding to the scalar data;
[0045] A data processing module communicatively connected to the data extraction module for performing normalization processing on the scalar data to obtain normalized data; and also for classifying the symbol data to obtain a plurality of symbol data classes;
[0046] A clustering processing module communicatively connected to the data processing module for performing clustering processing on the normalized data according to the symbol data classes to obtain a plurality of clustering results having the same number as the number of symbol data classes and a plurality of cluster center points corresponding to the plurality of clustering results; wherein, any one of the clustering results includes a plurality of normalized data;
[0047] An Euclidean distance calculation module communicatively connected to the clustering processing module for calculating the Euclidean distance α between each of the normalized data and its corresponding cluster center point; and also for calculating the mean value of all the normalized data in the clustering results corresponding to each symbol data class, and then calculating the Euclidean distance β between each of the normalized data and the mean value corresponding to the symbol data class to which it belongs;
[0048] A quantization processing module communicatively connected to the Euclidean distance calculation module for inputting the normalized data, the Euclidean distance α and the Euclidean distance β into a symbol data mapping model to obtain a symbol quantization result.
[0049] In a third aspect, an electronic device is provided, including:
[0050] A memory for storing computer program instructions; and,
[0051] A processor for executing the computer program instructions to complete the operations of the symbol data quantization processing method described in any one of the above.
[0052] In a fourth aspect, a computer-readable storage medium is provided for storing computer-readable computer program instructions, the computer program instructions being configured to perform the operations of the symbol data quantization processing method described in any one of the above when running. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flowchart of the symbol data quantization processing method in Embodiment 1;
[0054] Figure 2 is a flowchart of the symbol data quantization processing method in Embodiment 1;
[0055] Figure 3 is a flowchart of obtaining the symbol quantization result in Embodiment 1;
[0056] Figure 4 is a schematic structural diagram of the symbol data quantization processing module in Embodiment 1;
[0057] Figure 5 is a schematic structural diagram of the heterogeneous random field module in Embodiment 1;
[0058] Figure 6 is a schematic structural diagram of the discriminator in Embodiment 1;
[0059] Figure 7 is a flowchart of training the adversarial generation model in Embodiment 1;
[0060] Figure 8 is a block diagram of the symbol data quantization processing module in Embodiment 2. Detailed implementation manners
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the present invention will be briefly introduced below in combination with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structure of the drawings is only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. It should be noted here that the description of these embodiment modes is used to help understand the present invention, but does not constitute a limitation to the present invention.
[0062] Embodiment 1:
[0063] This embodiment discloses a symbol data quantization processing method, which can be, but is not limited to, executed by a computer device or a virtual machine with certain computing resources, such as an electronic device such as a personal computer, a smart phone, a personal digital assistant or a wearable device, or executed by a virtual machine.
[0064] As Figure 1 and 2 shown, a symbol data quantization processing method can, but is not limited to, include the following steps:
[0065] S1. Extract the original data from the database. The original data includes scalar data and symbolic data corresponding to the scalar data. In this embodiment, the database can be, but is not limited to, a clinical database. When extracting the original data from the database, the original data is extracted in the order of the data in the database, and each piece of data is extracted only once. Specifically, in this embodiment, it is assumed that the total amount of scalar data and symbolic data in the original data is N, where the number of scalar data is Ω, and the number of symbolic data classes corresponding to the symbolic data is r. It should be understood that one or more scalar data are matched under each symbolic data class, that is, the scalar data corresponds to the symbolic data class.
[0066] It should be noted that, taking the original data in the Figure 2 original data table as an example, 4.21 and 2.86 used to represent "(blood glucose content)", and 19 and 18 used to represent "(aspartate aminotransferase content)" are both scalar data, and "negative" and "positive" used to represent "HBeAg (hepatitis B e antigen) qualitative", and "entecavir" and "polyethylene glycol" used to represent "prescribed drugs" are both symbolic data.
[0067] S2. Perform standardization processing on the scalar data to obtain the standardized data. It should be noted that data standardization is a method of feature scaling and is a key step in preprocessing scalar data. Specifically, different evaluation indicators often have different dimensions and dimension units, and such a situation will affect the results of data analysis. In order to eliminate the influence of dimensions between indicators, data standardization processing is required to solve the comparability between data indicators. After the original data is processed by data standardization, each indicator is at the same order of magnitude, which is suitable for comprehensive comparison and evaluation, can facilitate the acceleration of the solution speed of gradient descent, and then improve the convergence speed of the subsequent model. It can eliminate the influence of magnitude and dimension, and then improve the accuracy of the model. It can also simplify the calculation amount.
[0068] In this embodiment, performing standardization processing on the scalar data includes:
[0069] Calculate the mean and standard deviation of the scalar data, and then obtain the standardized data corresponding to the scalar data according to the mean and standard deviation of the scalar data. Among them, the standardized data is:
[0070]
[0071] In the formula, x is the scalar data, μ is the mean of the scalar data, and σ is the standard deviation of the scalar data. It should be noted that the standardized data after standardization conforms to the standard normal distribution, that is, the mean of multiple standardized data is 0 and the standard deviation is 1; in this embodiment, the standardized data obtained after standardization can retain the useful information in the outliers of the scalar data in the original data, and can avoid the problem of the algorithm being sensitive to outliers in the subsequent data processing process.
[0072] Specifically, Figure 2 For example, the standardized data corresponding to 4.21 and 2.86 for "blood sugar (content)" are 0.823 and 0.417 respectively, and the standardized data corresponding to 19 and 18 for "aspartate aminotransferase (content)" are 0.882 and 0.845 respectively.
[0073] S3. Classify the symbol data to obtain multiple symbol data categories; it should be noted that the symbol data is classified, that is, the symbol data using the same symbol is divided into one category, such as Figure 2 The symbolic data "negative" and "positive" in the data are divided into the symbolic data category of "HBeAg qualitative". Figure 2 The symbol data "entecavir" and "polyethylene glycol" are classified into the symbol data category of "prescribed drugs".
[0074] S4. According to the symbol data class, the standardized data is clustered to obtain multiple clustering results the same as the number of the symbol data classes and multiple cluster center points corresponding to the multiple clustering results; wherein, any clustering result includes multiple standardized data; it should be understood that each standardized data is divided into the clustering result where the cluster center point closest to it is located.
[0075] In this embodiment, clustering is performed on the standardized data according to the symbol data class to obtain a plurality of clustering results having the same number as the symbol data class and a plurality of cluster center points corresponding to the plurality of clustering results, including:
[0076] S401. Randomly select a number of initial cluster center points that is the same as the number of the symbol data classes.
[0077] S402. Calculate the distances of the standardized data to multiple initial cluster center points in turn, and classify the current standardized data into the cluster with the smallest distance to the initial cluster center point, until all standardized data are classified into the corresponding cluster with the smallest distance to the initial cluster center point.
[0078] S403. Calculate the mean of all the standardized data in the clustering cluster of each initial cluster center point, and use this mean to replace the initial cluster center point of the corresponding clustering cluster, obtaining multiple updated cluster center points. It should be noted that in this embodiment, by calculating the mean of all the standardized data in the clustering cluster and using this mean as the new cluster center point, the update of the cluster center point is thus realized.
[0079] S404. Reset the updated initial cluster center points as the initial cluster center points, and re-calculate the distances from the standardized data to the multiple initial cluster center points in sequence until the change amount of the mean of all the standardized data in the clustering of each initial cluster center point is less than the threshold. At this time, the clustering cluster corresponding to the mean is the clustering result, and the mean at this time is the cluster center point corresponding to the clustering result. It should be noted that the change amount of the mean is the difference between the current mean and the mean in the previous step, and is used to measure the change amplitude between the original cluster center point and the updated cluster center point.
[0080] It should be noted that in this embodiment, when performing clustering processing on the standardized data, the K-Means (K-means clustering) algorithm is used for execution. Its clustering process is relatively easy, the convergence speed is fast, the clustering effect is excellent, and at the same time, the interpretability is relatively strong.
[0081] S5. Calculate the Euclidean distance α between each standardized data and its corresponding cluster center point. It should be understood that any standardized data is located in a certain clustering result, and this standardized data is near the cluster center point in the corresponding clustering result. It also needs to be noted that the cluster center point corresponding to any standardized data is the cluster center point of the clustering result where this standardized data is located.
[0082] S6. Calculate the mean of all the standardized data in the clustering results corresponding to each symbol data class, and then calculate the Euclidean distance β between each standardized data and the mean corresponding to the symbol data class to which it belongs.
[0083] In this embodiment, both the Euclidean distance α and the Euclidean distance β can be used as global variables for the subsequent symbol quantization process.
[0084] S7. Input the standardized data, the Euclidean distance α, and the Euclidean distance β into the symbol data mapping model to obtain the symbol quantization result. It should be noted that the symbol data mapping model is pre-set and trained for performing symbol quantization.
[0085] In this embodiment, the symbol data mapping model includes a generator and a calculation module connected in sequence.
[0086] Specifically, input the standardized data, the Euclidean distance α, and the Euclidean distance β into the symbol data mapping model to obtain the symbol quantization result, as Figure 3 shown, including:
[0087] S701. Input the standardized data into the generator, and input the Euclidean distance α and the Euclidean distance β into the calculation module;
[0088] S702. The generator processes the standardized data to obtain the initial symbol quantization value, and then inputs the initial symbol quantization value into the calculation module;
[0089] S703. The calculation module multiplies the initial symbol quantization value by the Euclidean distance α to obtain a multiplication result, and then adds the multiplication result to the Euclidean distance β to obtain the symbol quantization result of the symbol data.
[0090] In this embodiment, as Figure 4 and 5 shown, the generator includes a first input layer, a first hidden layer, and a first output layer connected in sequence, and a first normalization layer and an activation layer are sequentially connected between the first input layer and the first hidden layer and between the first hidden layer and the first output layer; the generator further includes a random module and a heterogeneous random field module, and there are two input ends of the heterogeneous random field module. The first hidden layer is connected to the first input end of the heterogeneous random field module, and the output end of the activation layer at the next level of the first hidden layer is connected to the second input end of the heterogeneous random field module through the random module; the output end of the first output layer and the output end of the heterogeneous random field module are connected to the calculation module. Specifically, in this embodiment, the number of discriminators is greater than 3 and less than the number of scalar data in the original data.
[0091] In this embodiment, the output value of the heterogeneous random field module is added to the output result of the first output layer to obtain the output value of the generator, that is, the initial symbol quantization value.
[0092] In this embodiment, in step S702, when the generator processes the standardized data to obtain the initial symbol quantization value, it includes:
[0093] A1. The standardized data is input into the generator through the first input layer for processing, and a first generation result is obtained through the first output layer;
[0094] A2. The result K output by the first hidden layer is input into the heterogeneous random field module from the first input end of the heterogeneous random field module. The result output by the activation layer at the next level of the first hidden layer is randomly shuffled through a random module to obtain a feature vector Q, which is then input into the heterogeneous random field module from the second input end of the heterogeneous random field module;
[0095] A3. According to a preset Gaussian random matrix, a result V can be obtained, and the result V is input into the heterogeneous random field module from the second input end of the heterogeneous random field module;
[0096] A4. The heterogeneous random field module combines and processes the result K, the feature vector Q, and the result V to obtain a second generation result;
[0097] A5. Add the first generation result and the second generation result to obtain an initial value of symbol quantization.
[0098] Specifically, when inputting the normalized data, the Euclidean distance α, and the Euclidean distance β into the symbol data mapping model, the normalized data is input into the generator from the first input layer of the generator for processing to obtain an initial value of symbol quantization. Then, through a calculation module, the initial value of symbol quantization, the Euclidean distance α, and the Euclidean distance β are combined to obtain the symbol quantization result corresponding to the symbol data. Among them, the normalized data is processed in each module according to the connection relationship of each module in the generator, and the first generation result obtained can be output through the first output layer of the generator; during this process, the result K output by the first hidden layer is input into the heterogeneous random field module from the first input end of the heterogeneous random field module. The result output by the activation layer at the next level of the first hidden layer can be randomly shuffled through a random module to obtain a feature vector Q, which is then input into the heterogeneous random field module from the second input end of the heterogeneous random field module. In addition, according to a preset Gaussian random matrix, a result V can be obtained, and this result V is also input into the heterogeneous random field module from the second input end of the heterogeneous random field module. The heterogeneous random field module can combine and process K, Q, and V to obtain a second generation result; after adding the first generation result and the second generation result, the initial value of symbol quantization is obtained.
[0099] It should be noted that the generator in this embodiment is designed by the inventor for the quantization of symbolic data in the standard data in the medical field. The heterogeneous random field module it contains has an efficient non-linear expression logic and plays a role in emphasizing global features during the model inference process. In this embodiment, the heterogeneous random field module can first perform signal processing on Q, specifically, perform Fourier series expansion processing on it, and then perform differential amplification processing and peak suppression processing on the data after Fourier series expansion processing respectively. The result of differential amplification processing is log. Differential amplification processing means using the square method to amplify the feature vector Q. When performing peak suppression processing, it is achieved by using the Peak to Average Power Ratio (PAPR); then, the heterogeneous random field module obtains the hidden Markov chain result of K; after that, the heterogeneous random field module obtains the recursive factorial of V; finally, multiply the result obtained by peak suppression processing by the hidden Markov chain result, then add it to the recursive factorial, and divide the obtained result by log to obtain the second generation result.
[0100] During this process, the peak suppression processing of Q can avoid the exposure of features when multiplying by the hidden Markov chain result of K later, that is, overemphasizing a certain feature; the recursive factorial of V can be used as an intermediate result for correcting features in subsequent addition operations to play a role in activating features; the differential amplification processing of Q can play a switching function to control which activated features to retain and which to discard.
[0101] In addition, in this embodiment, the steps for obtaining the symbolic data mapping model are as follows:
[0102] B1. Construct an initial adversarial generation model;
[0103] B2. Extract a training data set from the database;
[0104] B3. Use the training data set to train the initial adversarial generation model, and obtain the symbolic data mapping model after training.
[0105] Specifically, as Figure 7 shown, the adversarial generation model includes a generator and a discriminator connected in sequence. The generator in the adversarial generation model has the same structure as the generator in the symbolic data mapping model, that is, the generator includes a first input layer, a first hidden layer, and a first output layer connected in sequence, and a first normalization layer and an activation layer are connected in sequence between the first input layer and the first hidden layer and between the first hidden layer and the first output layer; as Figure 6 shown, the discriminator includes a second input layer, a second hidden layer, and a second output layer connected in sequence, and a second normalization layer is connected between the second input layer and the second hidden layer and between the second hidden layer and the second output layer.
[0106] Among them, the generator is used to combine the Euclidean distance α and the Euclidean distance β to generate the symbol quantization result corresponding to the symbol data;
[0107] The discriminator is used to verify the quality of the symbol quantization result generated by the generator and output the verification result for the user to further confirm the quantization quality of the generator; specifically, the verification result is any value between 0 and 1.
[0108] Specifically, in this embodiment, an adversarial connection relationship is formed between the generator and the discriminator. The combination of the generator and the discriminator can meet the conditions of adversarial learning and achieve the effect of competing with and improving each other.
[0109] In this embodiment, after the first hidden layer of the generator in the adversarial generation model and the symbol data mapping model receives the data input by the first input layer, and the second hidden layer of the discriminator receives the data input by the second input layer, a preset hidden layer formula is called to perform abstraction processing on the received data. Specifically, in this embodiment, the hidden layer formula is as follows:
[0110] FullyLayer i (x) = x·W i +b i ;
[0111] In the formula, x is the data received by the first hidden layer or the second hidden layer; W i is the weight of the data x; b i is the bias term of the data x; FullyLayer i (x) is the data output by the first hidden layer or the second hidden layer; i is 1 or 2, indicating that the current formula is the hidden layer formula used by the first hidden layer or the hidden layer formula used by the second hidden layer.
[0112] In this embodiment, the first hidden layer and the second hidden layer can perform abstraction processing on the features of the input data to better linearly divide different types of data.
[0113] In addition, in this embodiment, after the first normalization layer of the generator in the adversarial generation model and the symbol data mapping model and the second normalization layer of the discriminator receive the data, a preset layer normalization function is called to perform normalization processing on the received data. Specifically, in this embodiment, the layer normalization function is as follows:
[0114]
[0115] In the formula, y is the data received by the first normalization layer or the second normalization layer; E(y) is the average value of the data y; γ iis the random weight of data y; b is the bias term of data y, and var(y) is the sample variance of data y; LayerNorm i (y) is the data output by the first normalization layer or the second normalization layer; i is 1 or 2, indicating that the current formula is the layer normalization function used by the first normalization layer or the layer normalization function used by the second normalization layer.
[0116] In this embodiment, after the activation layers of the adversarial generation model and the symbolic data mapping model receive the data, a preset activation function is called to process the data. Specifically, in this embodiment, the activation function formula is as follows:
[0117]
[0118] In the formula, z is the data received by the activation layer; a is a preset limit parameter; LeakyReLU(z,a) is the data output by the activation layer.
[0119] In the generator of this embodiment, the number of weights of the first input layer is Ω, which is the same as the number of scalar data; in this embodiment, three first hidden layers are provided, and each first hidden layer is also connected by a first normalization layer and an activation layer, and the number of weights of each first hidden layer is Ω*4; the number of weights of the first output layer is 1. In this embodiment, the output of the first output layer is set as δ. In this embodiment, the activation layer uses Leaky ReLU as the activation function.
[0120] In the discriminator, the number of weights of the second input layer is Ω; in this embodiment, three second hidden layers are also provided, and each second hidden layer is connected by a second normalization layer, and the number of weights of each second hidden layer is Ω*4; the number of weights of the second output layer is 1.
[0121] In this embodiment, since the larger the number of discriminators, the greater the computational overhead, in order to balance both the computational overhead and the model accuracy, the number of discriminators is greater than 3 and less than the number Ω of scalar data in the original data.
[0122] In this embodiment, each discriminator uses its own loss value as the optimization target, and the optimization function is Adam (first-order optimization function); the loss value of the generator is the average of the discriminator loss values, and Adam is used as the optimization function.
[0123] Specifically, in this embodiment, the training data set includes a plurality of scalar data and symbol data corresponding to the scalar data. It should be noted that the plurality of scalar data can be randomly extracted from a database to be used as training data for the subsequent model, and the model is trained to obtain a trained symbol data mapping model for directly quantifying the symbol data. In this embodiment, the scalar data in the training data set is also extracted from the database and is the same as the scalar data in the original data extracted from the database mentioned above, except that one is used for symbol data quantization and the other is used for model training.
[0124] Correspondingly, the initial adversarial generation model is trained using the training data set, and after training, a symbol data mapping model is obtained, including:
[0125] B301. Standardize the scalar data to obtain standardized data;
[0126] B302. According to the symbol data classes, perform clustering processing on the standardized data to obtain a plurality of clustering results equal in number to the number of symbol data classes and a plurality of cluster center points corresponding to the plurality of clustering results;
[0127] B303. Calculate the Euclidean distance α between each piece of the standardized data and its corresponding cluster center point;
[0128] B304. Classify the symbol data to obtain a plurality of symbol data classes;
[0129] B305. Calculate the mean value of all the standardized data in the clustering results corresponding to each symbol data class, and then calculate the Euclidean distance β between each piece of the standardized data and the mean value corresponding to the symbol data class to which it belongs;
[0130] B306. Input the standardized data, the Euclidean distance α, and the Euclidean distance β into the generator of the initial adversarial generation model to obtain a plurality of symbol quantization results equal in number to the number of the standardized data;
[0131] B307. The discriminator of the initial adversarial generation model verifies the quality of the symbol quantization results generated by the generator and outputs a verification result for the user to further confirm the quantization quality of the generator, and based on the verification result, obtains a plurality of loss function values corresponding to the plurality of symbol quantization results;
[0132] B308. To ensure the reliability of the trained model, it is also possible to determine whether the model converges according to the sum of the plurality of loss function values. If not, reclassify the symbol data and obtain the Euclidean distance β until the loss function value drops to an appropriate value, such as 1. At this time, save the generator, that is, obtain the symbol data mapping model.
[0133] Specifically, in this embodiment, the initial adversarial generation model includes a generator and a discriminator connected in sequence, that is, the same as the structure of the symbol data mapping model. For example, let the scalar data be [x1, x2, x3, x4, x5], and let the number of discriminators be 5, that is, the discriminators are divided into discriminator 1 to discriminator 5. Then, the standardized data, the Euclidean distance α, and the Euclidean distance β are input into the initial adversarial generation model to obtain a plurality of symbol quantization results with the same quantity as the standardized data, including:
[0134] The standardized data is input into the generator to obtain a symbol quantization initial value, and the symbol quantization initial value is multiplied by the Euclidean distance α and then added to the Euclidean distance β to obtain an initial quantization result;
[0135] [x1, x2, x3, x4, initial quantization result] is input into discriminator 1 to obtain the symbol quantization result of scalar data x5;
[0136] [x1, x2, x3, x5, initial quantization result] is input into discriminator 2 to obtain the symbol quantization result of scalar data x4;
[0137] [x1, x2, x4, x5, initial quantization result] is input into discriminator 3 to obtain the symbol quantization result of scalar data x3;
[0138] [x1, x3, x4, x5, initial quantization result] is input into discriminator 4 to obtain the symbol quantization result of scalar data x2;
[0139] [x2, x3, x4, x5, initial quantization result] is input into discriminator 5 to obtain the symbol quantization result of scalar data x1;
[0140] In this embodiment, the loss function value of any discriminator is:
[0141]
[0142] where y i is the scalar data, such as x1, x2, x3, x4 or x5, is the symbol quantization result of the scalar data y i ;
[0143] The sum of multiple loss function values is:
[0144]
[0145] where n is the total amount of scalar data.
[0146] This embodiment facilitates the efficient discovery of potential connections between symbolic data features, which is conducive to better serving downstream tasks such as drug research and development and clinical diagnosis. Specifically, during the implementation of this embodiment, after extracting the original data from the database, the scalar data in the original data is standardized to obtain the standardized data; then, according to the symbolic data classes, the standardized data is clustered to obtain multiple clustering results with the same number as the symbolic data classes and multiple cluster centers corresponding to the multiple clustering results; then, the Euclidean distance α between each standardized data and its corresponding cluster center is calculated, and then the symbolic data is classified to obtain multiple symbolic data classes, and the Euclidean distance β between each standardized data and the mean value corresponding to the symbolic data class it belongs to is calculated; finally, the standardized data, the Euclidean distance α, and the Euclidean distance β are input into the symbolic data mapping model to obtain the symbolic quantization result.
[0147] Embodiment 2:
[0148] This embodiment discloses a symbolic data quantization processing system for implementing the symbolic data quantization processing method in Embodiment 1; as Figure 8 shown, the symbolic data quantization processing system includes:
[0149] A data extraction module, configured to extract original data from a database, where the original data includes scalar data and symbolic data corresponding to the scalar data;
[0150] A data processing module, communicatively connected to the data extraction module, configured to standardize the scalar data to obtain standardized data; and is also configured to classify the symbolic data to obtain multiple symbolic data classes;
[0151] A clustering processing module, communicatively connected to the data processing module, configured to cluster the standardized data according to the symbolic data classes to obtain multiple clustering results with the same number as the symbolic data classes and multiple cluster centers corresponding to the multiple clustering results; wherein, any one of the clustering results includes multiple standardized data;
[0152] A Euclidean distance calculation module, communicatively connected to the clustering processing module, configured to calculate the Euclidean distance α between each standardized data and its corresponding cluster center; and is also configured to calculate the mean value of all standardized data in the clustering results corresponding to each symbolic data class, and then calculate the Euclidean distance β between each standardized data and the mean value corresponding to the symbolic data class it belongs to;
[0153] A quantization processing module, communicatively connected to the Euclidean distance calculation module, is configured to input the standardized data, the Euclidean distance α, and the Euclidean distance β into a symbol data mapping model to obtain a symbol quantization result.
[0154] Embodiment 3:
[0155] Based on Embodiment 1 or 2, this embodiment discloses an electronic device, which may be a smart phone, a tablet computer, a notebook computer, a desktop computer, or the like. The electronic device may be referred to as a terminal device, a portable terminal device, a desktop terminal device, etc. The electronic device includes:
[0156] A memory, configured to store computer program instructions; and,
[0157] A processor, configured to execute the computer program instructions to complete the operations of the symbol data quantization processing method described in any one of Embodiment 1.
[0158] Embodiment 4:
[0159] Based on any one of Embodiments 1 to 3, this embodiment discloses a computer-readable storage medium, configured to store computer-readable computer program instructions, where the computer program instructions are configured to perform the operations of the symbol data quantization processing method described in Embodiment 1 when running.
[0160] It should be noted that if the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0161] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
[0163] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for quantifying symbolic data, characterized in that: including: extracting raw data from a database, where the raw data includes scalar data and symbolic data corresponding to the scalar data; performing standardization processing on the scalar data to obtain standardized data; classifying the symbolic data to obtain multiple symbolic data classes; performing clustering processing on the standardized data according to the symbolic data classes to obtain multiple clustering results equal in number to the number of the symbolic data classes and multiple cluster centers corresponding to the multiple clustering results; wherein, any one of the clustering results includes multiple standardized data; calculating the Euclidean distance α between each of the standardized data and its corresponding cluster center; calculating the mean value of all the standardized data in the clustering results corresponding to each symbolic data class, and then calculating the Euclidean distance β between each standardized data and the mean value corresponding to the symbolic data class to which it belongs; inputting the standardized data, the Euclidean distance α, and the Euclidean distance β into a symbolic data mapping model to obtain a symbolic quantization result.
2. The symbol data quantization processing method according to claim 1, wherein: Performing standardization processing on the scalar data includes: calculating the mean value and standard deviation of the scalar data, and then obtaining the standardized data corresponding to the scalar data according to the mean value and standard deviation of the scalar data; wherein, the standardized data is: where x is the scalar data, μ is the mean value of the scalar data, and σ is the standard deviation of the scalar data.
3. A method for quantifying and processing symbolic data according to claim 1, characterized in that: Performing clustering processing on the standardized data according to the symbolic data classes to obtain multiple clustering results equal in number to the number of the symbolic data classes and multiple cluster centers corresponding to the multiple clustering results includes: randomly selecting multiple initial cluster centers equal in number to the number of the symbolic data classes; successively calculating the distances of the standardized data to the multiple initial cluster centers, and classifying the current standardized data into the clustering cluster of the initial cluster center with the minimum distance until all the standardized data are classified into the corresponding clustering clusters of the initial cluster centers with the minimum distances; calculating the mean value of all the standardized data in the clustering cluster of each initial cluster center, and replacing the initial cluster center of the corresponding clustering cluster with the mean value to obtain multiple updated cluster centers; resetting the updated initial cluster centers as the initial cluster centers, and successively calculating the distances of the standardized data to the multiple initial cluster centers again until the change amount of the mean value of all the standardized data in the clustering of each initial cluster center is less than a threshold value. At this time, the clustering cluster corresponding to the mean value is the clustering result, and the mean value at this time is the cluster center corresponding to the clustering result.
4. A method for quantifying symbol data according to claim 1, characterized in that: The symbolic data mapping model includes a generator and a calculation module connected in sequence; Inputting the standardized data, the Euclidean distance α, and the Euclidean distance β into the symbolic data mapping model to obtain a symbolic quantization result includes: inputting the standardized data into the generator, and inputting the Euclidean distance α and the Euclidean distance β into the calculation module; the generator processes the standardized data to obtain an initial value of symbolic quantization, and then inputs the initial value of symbolic quantization into the calculation module; The computing module multiplies the initial symbol quantization value by the Euclidean distance α to obtain a multiplication result, and then adds the multiplication result to the Euclidean distance β to obtain the symbol quantization result of the symbol data.
5. A method for quantifying and processing symbolic data according to claim 4, characterized in that: The generator includes a first input layer, a first hidden layer, and a first output layer connected in sequence. A first normalization layer and an activation layer are sequentially connected between the first input layer and the first hidden layer and between the first hidden layer and the first output layer; the generator further includes a random module and a heterogeneous random field module. The heterogeneous random field module has two input ends. The first hidden layer is connected to the first input end of the heterogeneous random field module, and the output end of the activation layer at the next level of the first hidden layer is connected to the second input end of the heterogeneous random field module through the random module; the output end of the first output layer and the output end of the heterogeneous random field module are connected to the computing module.
6. A method for quantifying symbol data according to claim 5, characterized in that: The generator processes the normalized data to obtain the initial symbol quantization value, including: The normalized data is input into the generator through the first input layer for processing, and a first generation result is obtained through the first output layer; The result K output by the first hidden layer is input into the heterogeneous random field module from the first input end of the heterogeneous random field module. The result output by the activation layer at the next level of the first hidden layer is randomly shuffled through the random module to obtain a feature vector Q, and then input into the heterogeneous random field module from the second input end of the heterogeneous random field module; According to a preset Gaussian random matrix, a result V can be obtained, and the result V is input into the heterogeneous random field module from the second input end of the heterogeneous random field module; The heterogeneous random field module combines and processes the result K, the feature vector Q, and the result V to obtain a second generation result; The first generation result and the second generation result are added to obtain the initial symbol quantization value.
7. A method for quantifying symbol data according to claim 1, characterized in that: The obtaining steps of the symbol data mapping model are as follows: Construct an initial adversarial generation model; Extract a training data set from the database; Use the training data set to train the initial adversarial generation model, and obtain the symbol data mapping model after training.
8. A symbol data quantization processing system, characterized in that: For implementing the symbol data quantization processing method according to any one of claims 1 to 7; the symbol data quantization processing system includes: A data extraction module for extracting raw data from the database, where the raw data includes scalar data and symbol data corresponding to the scalar data; A data processing module communicatively connected to the data extraction module for performing normalization processing on the scalar data to obtain normalized data; and also for classifying the symbol data to obtain multiple symbol data classes; A clustering processing module communicatively connected to the data processing module for performing clustering processing on the normalized data according to the symbol data classes to obtain multiple clustering results having the same number as the symbol data classes and multiple cluster centers corresponding to the multiple clustering results; wherein, any one of the clustering results includes multiple normalized data; The Euclidean distance calculation module, which is communicatively connected to the clustering processing module, is used to calculate the Euclidean distance α between each of the standardized data and its corresponding cluster center point; it is also used to calculate the mean value of all the standardized data in the clustering result corresponding to each symbol data class, and then calculate the Euclidean distance β between each standardized data and the mean value corresponding to the symbol data class to which it belongs. The quantization processing module, which is communicatively connected to the Euclidean distance calculation module, is used to input the standardized data, the Euclidean distance α, and the Euclidean distance β into the symbol data mapping model to obtain a symbol quantization result.
9. An electronic device, characterized in that: Comprising: A memory for storing computer program instructions; And, A processor for executing the computer program instructions to complete the operations of the symbol data quantization processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium for storing computer-readable computer program instructions, characterized in that: The computer program instructions are configured to perform the operations of the symbol data quantization processing method as described in any one of claims 1 to 7 when running.
Citation Information
Patent Citations
Data object clustering, data processing and data identification method
CN110363206A
Human Emotion Assessment Based on Physiological Data Using Semiotic Analysis
US20170042463A1