Method, System, Device and Medium for Data Clustering and Generation of Simulation Service Platform
By processing data in parallel with the neural network, the clustering and discrimination process is optimized and the learning ability of the neural network is used for data generation, the problems of inefficient traditional methods and insufficient generalization ability are solved, and efficient and accurate data clustering and generation are achieved.
Patent Information
- Application Number
- CN202411370818.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-09-29
AI Technical Summary
Traditional data clustering and generation methods rely on a single algorithm model, which is inefficient or lacks generalization capabilities, and lacks effective combination with the data generation process, limiting its potential in application scenarios such as data enhancement and privacy protection.
The nested lattice and neural network are used to process data in parallel, and the clustering and discrimination process of neural networks is optimized through nested lattice coding, and the learning ability of neural networks is used to generate data.
The efficiency and accuracy of simulation data processing are improved, combined with random smoothing technology, the resistance to noise and outliers is enhanced, the clustering and generation process is achieved, and the pertinence and effectiveness of data generation is improved.
Smart Images

Figure CN119226830B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of simulation data processing, and particularly to a method, system, electronic device and storage medium for data clustering and generation in a simulation service platform. Background Art
[0002] With the rapid development of information technology, data processing and simulation technologies have become an indispensable part of the industrial design and manufacturing fields. Especially driven by artificial intelligence (AI) technology, the efficiency and accuracy of data processing and simulation have been continuously improved, providing strong support for industrial innovation and design optimization.
[0003] Data clustering is an important process of data processing, which can divide the samples in a data set into several groups or "clusters" that are similar to each other. In the fields of industrial design and simulation services, traditional data clustering and generation methods usually rely on a single algorithm model, such as a single neural network model, and these models may face problems of low efficiency or insufficient generalization ability when dealing with complex data structures. At the same time, traditional data clustering methods often lack an effective combination with the data generation process, restricting their potential in application scenarios such as data augmentation and privacy protection. Summary of the Invention
[0004] To solve or partially solve the problems existing in the related technologies, the present application provides a method, system, electronic device and storage medium for data clustering and generation in a simulation service platform, which can process data based on a nested lattice and a neural network in parallel, optimize the clustering discrimination process of the neural network through nested lattice encoding, and simultaneously use the learning ability of the neural network for data generation, improving the efficiency and accuracy of simulation data processing.
[0005] The first aspect of the present application provides a method for data clustering and generation in a simulation service platform, including:
[0006] Obtain input data, and add the input data to a preset smoothing vector to generate smoothed data;
[0007] Compress the smoothed data based on a preset nested lattice system to obtain encoded information;
[0008] Input the encoded information into a preset discrimination network and a preset generation network respectively;
[0009] The preset discrimination network performs clustering discrimination based on the encoded information to generate a clustering result;
[0010] The preset generation network generates new data samples based on the clustering result and the encoded information.
[0011] Preferably, the preset nested lattice system includes a fine lattice and a coarse lattice. Compressing the smooth data based on the preset nested lattice system to obtain encoded information includes:
[0012] Adjust the smooth data based on a preset estimation coefficient to generate adjusted data;
[0013] Quantize the adjusted data into the fine lattice to generate quantization information;
[0014] Map the quantization information to the coarse lattice to obtain the encoded information.
[0015] Preferably, the steps after compressing the smooth data based on the preset nested lattice system to obtain encoded information include:
[0016] Decompress the encoded information according to the preset estimation coefficient and the preset smooth vector to generate reconstructed data;
[0017] Calculate a distortion metric based on the preset target data and the reconstructed data; the distortion metric is used to evaluate the reconstruction quality of the reconstructed data.
[0018] Preferably, inputting the encoded information into a preset discriminant network and a preset generation network respectively includes:
[0019] Determine the compression distance from the input data to the preset target data according to the encoded information;
[0020] Input the compression distance into the preset discriminant network and the preset generation network respectively.
[0021] Preferably, the preset discriminant network performs clustering discrimination based on the encoded information to generate a clustering result, including:
[0022] The preset discriminant network performs clustering discrimination based on the compression distance to generate the clustering result;
[0023] The preset generation network generates new data samples based on the clustering result and the encoded information, including:
[0024] The preset generation network generates the new data samples based on the clustering result and the compression distance.
[0025] Preferably, after adding the input data and the preset smooth vector to generate smooth data, it further includes:
[0026] Add or subtract a noise vector to the smooth data.
[0027] Preferably, the method further includes:
[0028] Adjust the preset generation network according to the clustering result; or,
[0029] Adjust the preset discriminant network according to the new data sample; or,
[0030] Adjust the preset smoothing vector or the preset estimation coefficient according to the new data sample.
[0031] A second aspect of the present application provides a data clustering and generation system for a simulation service platform, including:
[0032] A preprocessing module, configured to obtain input data, add the input data to a preset smoothing vector, and generate smoothed data;
[0033] A compression module, configured to compress the smoothed data based on a preset nested lattice system to obtain encoded information;
[0034] An input module, configured to input the encoded information into a preset discriminant network and a preset generation network respectively;
[0035] A clustering module, configured to perform clustering discrimination on the encoded information by the preset discriminant network to generate a clustering result;
[0036] A generation module, configured to generate a new data sample by the preset generation network based on the clustering result and the encoded information.
[0037] A third aspect of the present application provides an electronic device, including:
[0038] A processor; and
[0039] A memory, storing executable code thereon, which when executed by the processor, causes the processor to execute the method as described above.
[0040] A fourth aspect of the present application provides a computer-readable storage medium, storing executable code thereon, which when executed by a processor of an electronic device, causes the processor to execute the method as described above.
[0041] The technical solution provided by this application may include the following beneficial effects: The embodiments of this application disclose a method for data clustering and generation of a simulation service platform, including obtaining input data, adding the input data to a preset smoothing vector to generate smoothed data, compressing the smoothed data based on a preset nested lattice system to obtain encoded information, inputting the encoded information into a preset discriminant network and a preset generation network respectively. The preset discriminant network performs clustering discrimination on the encoded information to generate a clustering result, and the preset generation network generates new data samples based on the clustering result and the encoded information. It can process data in parallel based on the nested lattice and neural network, optimize the clustering discrimination process of the neural network through nested lattice encoding, and at the same time utilize the learning ability of the neural network for data generation, improving the efficiency and accuracy of simulation data processing. Combining with the random smoothing technology, it improves the resistance to noise and outliers.
[0042] The technical solution of this application can also: use the nested lattice encoding technology to calculate the compression distance between data and use it as the input feature for neural network discrimination and generation, realizing the close combination of the clustering and generation processes, and improving the pertinence and effectiveness of data generation.
[0043] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] By describing the exemplary embodiments of this application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of this application will become more obvious. Among them, in the exemplary embodiments of this application, the same reference numerals generally represent the same components.
[0045] Figure 1 is a schematic flowchart of the method for data clustering and generation of the simulation service platform shown in the embodiments of this application;
[0046] Figure 2 is another schematic flowchart of the method for data clustering and generation of the simulation service platform shown in the embodiments of this application;
[0047] Figure 3 is a flowchart of the method for data clustering and generation of the simulation service platform shown in the embodiments of this application;
[0048] Figure 4 is another flowchart of the method for data clustering and generation of the simulation service platform shown in the embodiments of this application;
[0049] Figure 5 is a schematic structural diagram of the data clustering and generation system of the simulation service platform shown in the embodiments of this application;
[0050] Figure 6It is a schematic structural diagram of an electronic device shown in an embodiment of the present application. Detailed implementation manners
[0051] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application will be more thorough and complete, and can fully convey the scope of the present application to those skilled in the art.
[0052] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0053] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0054] Data clustering is an important process in data processing, which can divide the samples in a dataset into several groups or "clusters" that are similar to each other. In the field of industrial design and simulation services, traditional data clustering and generation methods usually rely on a single algorithm model, such as a single neural network model, and these models may face problems of low efficiency or insufficient generalization ability when dealing with complex data structures. At the same time, traditional data clustering methods often lack an effective combination with the data generation process, limiting their potential in application scenarios such as data augmentation and privacy protection.
[0055] In view of the above problems, the embodiments of the present application provide a method for data clustering and generation of a simulation service platform, which can process data in parallel based on a nested lattice and a neural network, optimize the clustering discrimination process of the neural network through nested lattice encoding, and at the same time use the learning ability of the neural network for data generation, improving the efficiency and accuracy of simulation data processing.
[0056] The technical solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0057] Figure 1 It is a schematic flow chart of the method for data clustering and generation of the simulation service platform shown in the embodiments of the present application.
[0058] See Figure 1 , the method includes:
[0059] Step 101, obtain input data, add the input data to a preset smoothing vector to generate smoothed data;
[0060] The input data of the present application can be training samples for model training, can also be raw data to be encrypted, and can also be data to be detected for anomalies. The input data can be obtained through sensor measurement, database query, user input, etc.
[0061] For the obtained input data, the input data can be preprocessed first, and the input data is added to a preset smoothing vector to generate smoothed data, aiming to reduce errors caused by noise or outliers and improve the adaptability to the input data. In one example, the input data can be raw data X, or can also be observed data Y. For example, noise interference can be added to the raw data X, and the raw data X is added with a noise vector N to obtain the observed data Y.
[0062] Step 102, compress the smoothed data based on a preset nested lattice system to obtain encoded information;
[0063] In the embodiments of the present application, a preset nested lattice system is introduced to compress the smoothed data, which improves the efficiency of high-dimensional data compression and also improves the accuracy of high-dimensional data compression. By compressing the smoothed data, encoded information can be obtained.
[0064] Step 103, input the encoded information into a preset discriminant network and a preset generation network respectively;
[0065] The discriminant network is a machine learning model mainly used to judge which category or the probability of a category the input data belongs to. The generation network is a neural network used to generate new data samples.
[0066] The encoded information is input into a preset discriminant network and a preset generation network respectively to perform subsequent processing on the data.
[0067] Step 104, the preset discriminant network performs clustering discrimination based on the encoded information to generate a clustering result;
[0068] The preset discriminant network performs clustering discrimination based on the encoded information to generate a clustering result. In one example, when the input data is X, the clustering result can be the category of X; when the input data is Y, the clustering result can be the category of Y; when the input data is Z and W, the clustering result can be whether Z and W belong to the same category.
[0069] Step 105, the preset generation network generates new data samples based on the clustering result and the encoding information.
[0070] According to the clustering result and the encoding information, the preset generation network can learn and generate new data samples that conform to specific clustering characteristics. The new data samples are related to the input data. If the input data is a picture, then the new data samples are also pictures; if the input data is a formula, then the new data samples are also formulas; if the input data is a question and answer, then the new data samples are also questions and answers.
[0071] In the preset generation network, a reward mechanism and a punishment mechanism are also set. If the new data sample and the input data do not belong to the same category, negative feedback can be generated and fed back to the preset generation network and the preset discrimination network; if the new data sample and the input data belong to the same category, positive feedback can be generated and fed back to the preset generation network and the preset discrimination network. By setting the reward mechanism and the punishment mechanism, the neural networks can perform cyclic iteration, continuously optimizing and improving the clustering effect and the quality of the generated data.
[0072] The embodiments of the present application can be applied in the machine learning training process, where new samples can be generated to enhance the training data, thereby improving the generalization ability of the model; it can also be applied in the data encryption process to compress and encrypt the data while retaining the key information of the data, protecting the privacy of the original data; it can also perform anomaly detection on the data to improve the accuracy and security of the data.
[0073] The embodiments of the present application disclose a method for data clustering and generation of a simulation service platform, including obtaining input data, adding the input data to a preset smoothing vector to generate smoothed data, compressing the smoothed data based on a preset nested lattice system to obtain encoding information, inputting the encoding information into a preset discrimination network and a preset generation network respectively, the preset discrimination network performing clustering discrimination on the encoding information to generate a clustering result, and the preset generation network generating new data samples based on the clustering result and the encoding information. It can process data in parallel based on the nested lattice and the neural network, optimize the clustering discrimination process of the neural network through nested lattice encoding, and at the same time utilize the learning ability of the neural network for data generation, improving the efficiency and accuracy of simulation data processing. Combining with the random smoothing technology, it improves the resistance to noise and outliers.
[0074] Figure 2 It is another flow schematic diagram of the method for data clustering and generation of the simulation service platform shown in the embodiments of the present application.
[0075] See Figure 2 This method includes:
[0076] Step 201, obtain input data, add the input data to a preset smoothing vector to generate smoothed data;
[0077] The input data of this application can be training samples for model training, can also be raw data to be encrypted, or can also be data to be detected for anomalies. The input data can be obtained through sensor measurement, database query, user input, etc.
[0078] For the obtained input data, the input data can be preprocessed first, and the input data is added to a preset smoothing vector to generate smoothed data, aiming to reduce errors caused by noise or outliers and improve the adaptability to the input data. In an example, the input data can be raw data X or observed data Y. For example, noise interference can be added to the raw data X, and the raw data X is added with a noise vector N to obtain the observed data Y.
[0079] In an optional embodiment of this application, after step 201, it further includes:
[0080] Step S11, add or subtract a noise vector based on the smoothed data.
[0081] If the input data is raw data X, if there is noise interference in the raw data, after generating the smoothed data, a noise vector N can also be added based on the smoothed data.
[0082] If the input data is observed data Y, if there is noise interference in the observed data, after generating the smoothed data, the noise vector N can also be subtracted based on the smoothed data.
[0083] Step 202, compress the smoothed data based on a preset nested lattice system to obtain encoded information;
[0084] In the embodiment of this application, a preset nested lattice system is introduced to compress data, which improves the efficiency of high-dimensional data compression and also improves the accuracy of high-dimensional data compression. By compressing the smoothed data, encoded information can be obtained.
[0085] In an optional embodiment of this application, the preset nested lattice system includes a fine lattice and a coarse lattice, and step 202 includes:
[0086] Sub-step 2021, adjust the smoothed data based on a preset estimation coefficient to generate adjusted data;
[0087] The preset nested lattice system includes a fine lattice Lambda_F and a coarse lattice Lambda_C, and the coarse lattice Lambda_C is a subset of the fine lattice Lambda_F.
[0088] Multiply the preset estimation coefficient α by the smoothed data to perform a scaling adjustment on the smoothed data, thereby generating adjusted data.
[0089] Sub-step 2022: Quantize the adjusted data to a fine lattice to generate quantization information;
[0090] A quantization function can be used to perform a quantization operation on the adjusted data to obtain quantization information. When the input data is the original data X, the quantization information is the quantization value quantized_value; when the input data is the original data Y, the quantization information is the quantization vector quantized_vector.
[0091] Sub-step 2023: Map the quantization information to a coarse lattice to obtain encoded information.
[0092] Perform a modulo operation of the quantization information with respect to the coarse lattice Lambda_C to map the quantization information to the coarse lattice Lambda_C, thereby obtaining encoded information. The quantization value is further compressed into the coarse lattice through the modulo operation for easy storage and transmission.
[0093] In an alternative embodiment of the present application, after step 202, it further includes:
[0094] Step S21: Decompress the encoded information according to the preset estimation coefficient and the preset smoothing vector to generate reconstructed data;
[0095] According to the preset estimation coefficient α and the preset smoothing vector U, the encoded information can be decompressed to generate reconstructed data.
[0096] Step S22: Calculate a distortion metric based on the preset target data and the reconstructed data; the distortion metric is used to evaluate the reconstruction quality of the reconstructed data and is also used to adjust the preset estimation coefficient.
[0097] When the input data is X, the target data can be Y; when the input data is Y, the target data can be X. The preset target data is known or obtained in some way, such as real data, training data, or other reference data, etc. Calculate the Euclidean distance between the preset target data and the reconstructed data through the Euclidean distance function, and use the Euclidean distance between the preset target data and the reconstructed data as the distortion metric of the preset target data and the reconstructed data. The distortion metric can evaluate the reconstruction quality of the reconstructed data. If the distortion metric is within a preset range, it can be considered that the reconstruction quality of the reconstructed data is high. If the distortion metric is not within the preset range, the quality of the reconstructed data is not high, and it can be considered that the decompression performance does not meet the requirements, and the decompression performance is adjusted by adjusting the preset estimation coefficient.
[0098] Step 203: Determine the compression distance from the input data to the target data according to the encoded information;
[0099] If the distortion metric is within a preset range, the length of the encoded information or a certain complexity metric is used as the compression distance from the input data to the target data.
[0100] Step 204: Input the compression distance into a preset discrimination network and a preset generation network respectively.
[0101] Input the compression distance into the preset discrimination network and the preset generation network simultaneously. The preset discrimination network and the preset generation network can process the compression distance in parallel to improve the efficiency of data processing.
[0102] Step 205: The preset discrimination network performs clustering discrimination based on the compression distance and generates a clustering result.
[0103] The preset discrimination network uses the distance information as input features for clustering discrimination, generates a clustering result, and outputs data such as clustering labels, clustering centers, or clustering quality metrics.
[0104] Step 206: The preset generation network generates new data samples based on the clustering result and the compression distance.
[0105] The preset generation network then generates data based on the clustering result and the compression distance, generating new data samples similar to the clustering center. The compression distance can be used as an indicator to evaluate the clustering effect and the data compression efficiency. A lower compression distance usually means a better clustering effect and a higher data compression efficiency.
[0106] By processing data in parallel with nested lattice encoding and neural networks, the compression distance metric based on nested lattice encoding is closely combined with the clustering discrimination and generation processes based on neural networks, improving the efficiency and accuracy of data processing.
[0107] In an optional embodiment of the present application, the method further includes:
[0108] Step S31: Adjust the preset generation network according to the clustering result; or,
[0109] Adjust the preset discrimination network according to the new data sample; or,
[0110] Adjust the preset smoothing vector or preset estimation coefficient according to the new data sample.
[0111] The clustering results generated according to the preset discrimination network can adjust the parameters of the preset generation network. The clustering algorithm of the preset discrimination network can be adjusted according to the generation quality of new data samples. Additionally, the preset smoothing vector or preset estimation coefficient can be adjusted according to new data samples. By continuously optimizing and adjusting the neural network, the clustering effect and data generation quality can be improved. By continuously adjusting the preset smoothing vector or preset estimation coefficient, the best data processing effect can be achieved. Moreover, the step sizes of the fine lattice and the coarse lattice can be adjusted to meet the data processing requirements in different fields and scenarios, featuring high flexibility and scalability.
[0112] To better understand the technical solutions of the embodiments of the present application, the following separately takes the input data as the original data X and the observed data Y as examples to elaborate on the solutions of the embodiments of the present application in detail. Refer to Figure 3 , which is the flowchart of the data clustering and generation method of the simulation service platform shown in the embodiments of the present application. When the input data is the original data X, the method includes:
[0113] Step 301: Obtain the input data X, add the input data X to the preset smoothing vector U to generate the smoothed data X_smooth;
[0114] Obtain the original data X, which can be a vector or matrix of any dimension. Use a random number generator to generate a preset smoothing vector U with the same dimension as X. The element distribution of the preset smoothing vector U can be a Gaussian distribution or other suitable distributions. Add the preset smoothing vector U to the original data X to obtain the smoothed data X_smooth = X + U.
[0115] In an optional embodiment of the present application, after step 301, it further includes:
[0116] Step S31: Add the noise vector N to the smoothed data X_smooth.
[0117] If there is noise interference in the original data, after generating the smoothed data, the noise vector N can also be added to the smoothed data to obtain the smoothed data X 1 = X + U + N.
[0118] Step 302: Adjust the smoothed data X_smooth based on the preset estimation coefficient α to generate the adjusted data X t .
[0119] If there is no noise interference in the original data, multiply the preset estimation coefficient α by the smoothed data to perform scaling adjustment on the smoothed data X_smooth, thereby generating the adjusted data X t= α × (X + U). If there is noise interference in the original data, then the smoothed data containing noise is multiplied by the preset estimation coefficient α to generate the adjusted data X t = α × (X + U + N).
[0120] Step 303: Quantize the adjusted data X t to the fine lattice Lambda_F to generate the quantization value quantized_value.
[0121] The modulo operation and rounding operation can be performed using the quantization function mod(round(...), Lambda_F) to obtain the quantization value quantized_value, that is, quantized_value = mod(round(X t ), Lambda_F), where X t can be substituted into the corresponding formula according to whether there is noise interference in the original data.
[0122] Step 304: Map the quantization value quantized_value to the coarse lattice Lambda_C to obtain the encoded information S_Y.
[0123] Perform the modulo operation of the quantization value quantized_value with the coarse lattice Lambda_C to map the quantization value quantized_value to the coarse lattice Lambda_C, thereby obtaining the encoded information S_Y, that is, S_Y = mod(quantized_value, Lambda_C).
[0124] Step 305: Decompress the encoded information S_Y according to the preset estimation coefficient α and the preset smoothing vector U to generate the reconstructed data Y_hat;
[0125] According to the preset estimation coefficient α and the preset smoothing vector U, the encoded information S_Y can be decompressed to generate the reconstructed data Y_hat, that is, Y_hat = X + α × mod(S_Y - U - α × X, Lambda_C).
[0126] Step 306: Calculate the distortion measure D based on the preset target data Y and the reconstructed data Y_hat.
[0127] The preset target data Y is known or obtained in some way. The distortion metric D is calculated using the Euclidean distance function based on the preset target data Y and the reconstructed data Y_hat, i.e., D = norm(Y_hat - Y). The distortion metric can evaluate the reconstruction quality of the reconstructed data. If the distortion metric is within the preset range, it can be considered that the reconstruction quality of the reconstructed data is high. If the distortion metric is not within the preset range, the quality of the reconstructed data is not high, and it can be considered that the decompression performance does not meet the requirements. The decompression performance is adjusted by adjusting the preset estimation coefficient.
[0128] Step 307, determine the compression distance distance from the input data X to the target data Y according to the coding information S_Y.
[0129] Take the length of the coding information or a certain complexity metric as the compression distance distance from the input data to the target data. In one example, distance = length(S_Y).
[0130] Step 308, input the compression distance distance into the preset discriminant network and the preset generation network respectively;
[0131] Input the compression distance into the preset discriminant network and the preset generation network simultaneously. The preset discriminant network and the preset generation network can process the compression distance in parallel to improve the efficiency of data processing.
[0132] Step 309, the preset discriminant network performs clustering discrimination based on the compression distance and generates a clustering result;
[0133] The preset discriminant network uses the distance information as input features for clustering discrimination, generates a clustering result, and outputs data such as clustering labels, cluster centers, or clustering quality metrics.
[0134] Step 3010, the preset generation network generates new data samples based on the clustering result and the compression distance.
[0135] The preset generation network then generates data based on the clustering result and the compression distance, generating new data samples similar to the cluster centers. At the same time, the compression distance can be used as an indicator to evaluate the clustering effect and the data compression efficiency of the generated data. A lower compression distance usually means a better clustering effect and a higher data compression efficiency.
[0136] See Figure 4 , which is another flowchart of the data clustering and generation method of the simulation service platform shown in the embodiments of the present application. When the input data is the observed data Y, the method includes:
[0137] Step 401, obtain the input data Y, add the input data Y to the preset smoothing vector U to generate the smoothed data Y_smooth;
[0138] Obtain the observation data Y, which can be obtained through sensors, databases, user input, etc. Use a random number generator to generate a preset smoothing vector U with the same dimension as Y. The element distribution of the preset smoothing vector U can be a Gaussian distribution or other suitable distributions. Add the preset smoothing vector U to the observation data Y to obtain the smoothed data Y_smooth = Y + U.
[0139] Step 402: Subtract the noise vector N from the smoothed data Y_smooth.
[0140] After generating the smoothed data, it is also possible to subtract the noise vector N from the smoothed data to obtain the smoothed data Y without noise 1 = Y + U – N, reducing the impact of noise on subsequent processing.
[0141] Step 403: Adjust the smoothed data without noise based on the preset estimation coefficient α to generate the adjusted data Y t .
[0142] Multiply the preset estimation coefficient α by the smoothed data to scale and adjust the smoothed data Y without noise 1 so as to generate the adjusted data Y t = α×(Y + U – N).
[0143] Step 404: Quantize the adjusted data Y t to the fine lattice Lambda_F to generate the quantization vector quantized_vector.
[0144] Divide the adjusted data Y t by the fine lattice quantization step Lambda_F, then perform a rounding operation, and finally multiply the result by Lambda_F to obtain the quantization vector quantized_vector. The purpose of this step is to map continuous data values to discrete lattice points.
[0145] Step 405: Map the quantization vector quantized_vector to the coarse lattice Lambda_C to obtain the encoded information S_X.
[0146] Perform a modulo operation of the quantization vector quantized_vector with respect to the coarse lattice Lambda_C to map the quantization vector quantized_vector to the coarse lattice Lambda_C, thereby obtaining the encoded information S_X, that is, S_X = mod(quantized_vector, Lambda_C). Through the modulo operation, the quantization vector is further compressed into the coarse lattice for easy storage and transmission.
[0147] Step 406, decompress the encoded information S_X according to the preset estimation coefficient α and the preset smoothing vector U to generate the reconstructed data X_hat;
[0148] According to the preset estimation coefficient α and the preset smoothing vector U, the encoded information S_X can be decompressed to generate the reconstructed data X_hat, that is, X_hat = Y + α × mod(S_X - U - α × Y, Lambda_C).
[0149] Step 407, calculate the distortion measure D based on the preset target data X and the reconstructed data X_hat.
[0150] The preset target data X is known or obtained in some way. According to the preset target data X and the reconstructed data X_hat, the distortion measure D is calculated using the Euclidean distance function, that is, D = norm(X_hat - X). The distortion measure can evaluate the reconstruction quality of the reconstructed data. If the distortion measure is within the preset range, it can be considered that the reconstruction quality of the reconstructed data is high. If the distortion measure is not within the preset range, the quality of the reconstructed data is not high, and it can be considered that the decompression performance does not meet the requirements. The decompression performance can be adjusted by adjusting the preset estimation coefficient.
[0151] Step 408, determine the compression distance distance from the input data Y to the target data X according to the encoded information S_X.
[0152] Take the length of the encoded information or a certain complexity measure as the compression distance distance from the input data to the target data. In one example, distance = length(S_X).
[0153] Step 409, input the compression distance distance into the preset discriminant network and the preset generation network respectively;
[0154] Input the compression distance into the preset discriminant network and the preset generation network simultaneously. The preset discriminant network and the preset generation network can process the compression distance in parallel to improve the efficiency of data processing.
[0155] Step 4010, the preset discriminant network performs clustering discrimination based on the compression distance to generate a clustering result;
[0156] The preset discriminant network uses the distance information as input features for clustering discrimination, generates a clustering result, and outputs data such as clustering labels, cluster centers, or clustering quality indicators.
[0157] Step 4011, the preset generation network generates new data samples based on the clustering result and the compression distance.
[0158] The preset generation network generates data based on the clustering result and the compression distance, generating new data samples similar to the clustering center. The compression distance can be used as an indicator to evaluate the clustering effect and the data generation compression efficiency. A lower compression distance usually means a better clustering effect and a higher data compression efficiency.
[0159] The embodiments of the present application disclose a method for data clustering and generation of a simulation service platform, including obtaining input data, adding the input data to a preset smoothing vector to generate smoothed data, compressing the smoothed data based on a preset nested lattice system to obtain encoded information, determining the compression distance from the input data to the target data according to the encoded information, inputting the compression distance into a preset discriminant network and a preset generation network respectively, the preset discriminant network performing clustering discrimination based on the compression distance to generate a clustering result, and the preset generation network generating new data samples based on the clustering result and the compression distance. By processing data in parallel based on the nested lattice and the neural network, optimizing the clustering discrimination process of the neural network through nested lattice encoding, and at the same time using the learning ability of the neural network for data generation, the efficiency and accuracy of simulation data processing are improved. Using the nested lattice encoding technology to calculate the compression distance between data and taking it as the input feature for neural network discrimination and generation realizes the close combination of the clustering and generation processes, improves the pertinence and effectiveness of data generation, introduces a random smoothing technology in the preprocessing stage, adding random noise to the input data to reduce errors caused by noise or outliers, and improving the adaptability and robustness of the system to the input data.
[0160] Corresponding to the foregoing embodiments of the application function implementation method, the present application also provides a system for data clustering and generation of a simulation service platform, an electronic device, and corresponding embodiments.
[0161] Figure 5 It is a schematic structural diagram of a data clustering and generation system 500 shown in the embodiments of the present application.
[0162] See Figure 5 and the system includes:
[0163] A preprocessing module 501, configured to obtain input data and add the input data to a preset smoothing vector to generate smoothed data;
[0164] A compression module 502, configured to compress the smoothed data based on a preset nested lattice system to obtain encoded information;
[0165] An input module 503, configured to input the encoded information into a preset discriminant network and a preset generation network respectively;
[0166] A clustering module 504, configured to perform clustering discrimination on the encoded information by a preset discriminant network to generate a clustering result;
[0167] A generation module 505, configured to generate new data samples based on a clustering result and coding information by a preset generation network.
[0168] In an optional embodiment of the present application, the preset nested lattice system includes a fine lattice and a coarse lattice, and the compression module 502 includes:
[0169] An adjustment sub-module, configured to adjust smooth data based on a preset estimation coefficient to generate adjusted data;
[0170] A quantization sub-module, configured to quantize the adjusted data to the fine lattice to generate quantization information;
[0171] A mapping sub-module, configured to map the quantization information to the coarse lattice to obtain coding information.
[0172] In an optional embodiment of the present application, the system further includes:
[0173] A decompression module, configured to decompress the coding information according to a preset estimation coefficient and a preset smooth vector to generate reconstructed data;
[0174] A distortion module, configured to calculate a distortion metric based on a preset target data and the reconstructed data; the distortion metric is used to evaluate the reconstruction quality of the reconstructed data.
[0175] In an optional embodiment of the present application, the input module 503 includes:
[0176] A compression distance sub-module, configured to determine a compression distance from input data to a preset target data according to the coding information;
[0177] An input sub-module, configured to input the compression distance into a preset discriminant network and a preset generation network respectively.
[0178] In an optional embodiment of the present application, the clustering module 504 includes:
[0179] A clustering sub-module, configured to perform clustering discrimination based on the compression distance by a preset discriminant network to generate a clustering result;
[0180] The generation module 505 includes:
[0181] A generation sub-module, configured to generate new data samples based on the clustering result and the compression distance by a preset generation network.
[0182] In an optional embodiment of the present application, the system further further includes:
[0183] A noise module, configured to add or subtract a noise vector to / from the smooth data.
[0184] In an optional embodiment of the present application, the system further includes:
[0185] An adjustment module, configured to adjust a preset generation network according to a clustering result; or,
[0186] Adjust a preset discrimination network according to new data samples; or,
[0187] Adjust a preset smoothing vector or a preset estimation coefficient according to new data samples.
[0188] The embodiment of the present application discloses a data clustering and generation system for a simulation service platform, including obtaining input data, adding the input data to a preset smoothing vector to generate smoothed data, compressing the smoothed data based on a preset nested lattice system to obtain encoded information, and inputting the encoded information into a preset discrimination network and a preset generation network respectively. The preset discrimination network performs clustering discrimination on the encoded information to generate a clustering result, and the preset generation network generates new data samples based on the clustering result and the encoded information. It can process data in parallel based on the nested lattice and the neural network, optimize the clustering discrimination process of the neural network through nested lattice encoding, and at the same time utilize the learning ability of the neural network to generate data, improving the efficiency and accuracy of simulation data processing. Combining with the random smoothing technology, it improves the resistance to noise and outliers.
[0189] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0190] Figure 6 It is a schematic structural diagram of an electronic device shown in the embodiment of the present application.
[0191] See Figure 6 , the electronic device 600 includes a memory 610 and a processor 620.
[0192] The processor 620 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0193] The memory 610 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM may store static data or instructions required by the processor 620 or other modules of the computer. The permanent storage device may be a read-write storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. The system memory may store some or all of the instructions and data required by the processor during operation. In addition, the memory 610 may include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 610 may include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, ultra density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.
[0194] Executable code is stored on the memory 610, and when the executable code is processed by the processor 620, it can cause the processor 620 to execute some or all of the methods described above.
[0195] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, and the computer program or the computer program product includes computer program code instructions for executing some or all of the above steps of the method according to the present application.
[0196] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or the computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the above method according to the present application.
[0197] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A simulation service platform data clustering and generation method, characterized in that: The method comprises: Obtain input data, and add the input data to a preset smoothing vector to generate smoothed data; the input data includes at least one of a training sample for model training, raw data to be encrypted, and data to be detected for anomalies, wherein the input data includes a picture, the preset smoothing vector has the same dimension as the input data, and the element distribution of the preset smoothing vector is a Gaussian distribution; Compressing the smoothed data based on a preset nested lattice system to obtain encoded information, including: the preset nested lattice system includes a fine lattice and a coarse lattice, adjusting the smoothed data based on preset estimation coefficients to generate adjusted data, quantizing the adjusted data to the fine lattice to generate quantized information, mapping the quantized information to the coarse lattice to obtain the encoded information, wherein the preset estimation coefficient is multiplied by the smoothed data to scale the smoothed data to generate the adjusted data; Inputting the encoded information into a preset discriminant network and a preset generative network respectively; The preset discriminant network performs cluster discrimination based on the encoding information to generate a clustering result; which includes: the preset discriminant network performs cluster discrimination based on a compression distance to generate the clustering result; the compression distance is the distance from the input data to the preset target data determined according to the encoding information; The preset generation network generates a new data sample based on the clustering result and the encoding information; this includes: the preset generation network generates the new data sample based on the clustering result and the compression distance.
2. The method according to claim 1, characterized in that The step of compressing the smooth data based on the preset nested lattice system to obtain the coded information comprises: Decompressing the coded information according to the preset estimation coefficient and the preset smoothing vector to generate reconstructed data; A distortion metric is calculated based on preset target data and the reconstructed data; the distortion metric is used to evaluate the reconstruction quality of the reconstructed data.
3. The method according to claim 2, characterized in that The step of inputting the coded information into a preset discrimination network and a preset generation network respectively comprises: Determining a compression distance from the input data to preset target data according to the encoding information; The compression distance is input into the preset discrimination network and the preset generation network respectively.
4. The method according to claim 1, characterized in that: After adding the input data to the preset smoothing vector to generate smoothed data, the method further includes: A noise vector is added or subtracted based on the smoothed data.
5. The method according to claim 1, characterized in that The method further comprises: adjusting the preset generation network according to the clustering result; or, Adjust the preset discriminant network according to the new data sample; or, The preset smoothing vector or the preset estimation coefficient is adjusted according to the new data sample.
6. A simulation service platform data clustering and generation system, characterized in that: The system comprises: A preprocessing module, used for acquiring input data, adding the input data to a preset smoothing vector, and generating smoothed data; the input data includes at least one of a training sample for model training, raw data to be encrypted, and data to be detected for anomalies, wherein the input data includes a picture, the preset smoothing vector has the same dimension as the input data, and the element distribution of the preset smoothing vector is a Gaussian distribution; A compression module, configured to compress the smoothed data based on a preset nested lattice system to obtain coded information, including: the preset nested lattice system includes a fine lattice and a coarse lattice, adjusting the smoothed data based on preset estimation coefficients to generate adjustment data, quantizing the adjustment data to the fine lattice to generate quantization information, mapping the quantization information to the coarse lattice to obtain the coded information, wherein the preset estimation coefficient is multiplied by the smoothed data to scale the smoothed data to generate the adjustment data; An input module, used to input the encoding information into a preset discrimination network and a preset generation network respectively; A clustering module, used for the preset discriminant network to perform clustering discrimination on the coded information and generate a clustering result; including: a clustering submodule, used for the preset discriminant network to perform clustering discrimination based on a compression distance and generate the clustering result; the compression distance is the distance from the input data to the preset target data determined according to the coded information; A generation module, used for the preset generation network to generate new data samples based on the clustering results and the encoding information; including: a generation submodule, used for the preset generation network to generate the new data samples based on the clustering results and the compression distance.
7. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: An executable code is stored thereon, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method as claimed in any one of claims 1 to 5.
Citation Information
Patent Citations
Information compression clustering distance calculation device
CN114529743A
Tone generation method based on voice conversion
CN118197329A