Simulation service platform data classification method based on dynamic distance measurement
By adopting a data classification method based on dynamic distance metrics on the simulation service platform, using the lattice system and plane factor dynamic adjustment distance metric strategy, the problem of complex, high-dimensional, and low classification accuracy of large-scale data sets is solved, and more efficient and accurate data clustering is achieved.
Patent Information
- Application Number
- CN202510037993.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing simulation service platform data classification methods are difficult to adapt to complex, high-dimensional, and large-scale data sets, and the data classification accuracy is low.
The simulation service platform data classification method based on dynamic distance metric is adopted. By constructing a lattice system, the data set is mapped into the lattice system, and in each iterative optimization process, the plane factor of the lattice system is calculated based on the standard deviation of each cluster, and the distance measurement strategy of the corresponding cluster is dynamically adjusted to achieve data classification.
It improves the accuracy and efficiency of data clustering, and is especially suitable for processing high-dimensional, large-scale or complex data sets, solving the problem of low data classification accuracy in the prior art.
Smart Images

Figure CN119989041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of simulation service platform data processing, and in particular to a simulation service platform data classification method based on dynamic distance measurement. Background Art
[0002] In industrial design and virtual simulation services, it is often necessary to process high-dimensional, large-scale data sets. These data sets often contain complex structural information and a large number of details, making it difficult for traditional data processing and clustering methods to meet actual needs in terms of computational efficiency and accuracy. At the same time, high computing costs and long processing cycles not only increase the operating costs of enterprises, but also extend the product development and listing cycles, limiting the market competitiveness of enterprises.
[0003] In industrial design and virtual simulation services, accurate identification of differences between data points is the key to ensuring product design quality and the accuracy of simulation results. However, traditional distance measurement methods, such as Euclidean distance and Manhattan distance, often cannot effectively distinguish the differences between data points when processing data sets with complex structures. This inadaptability leads to inaccurate classification results and makes it difficult to meet the high-precision requirements of industrial design and simulation.
[0004] Therefore, the field of industrial design and virtual simulation services urgently needs a data processing and classification method that can efficiently and accurately process high-dimensional, large-scale data sets and adapt to the characteristics of complex data structures. Summary of the invention
[0005] In view of the above analysis, an embodiment of the present invention aims to provide a simulation service platform data classification method based on dynamic distance measurement, so as to solve the problem that the existing simulation service platform data classification method is difficult to adapt to complex, high-dimensional, large-scale data sets and has low data classification accuracy.
[0006] The present invention discloses a simulation service platform data classification method based on dynamic distance measurement, the method comprising:
[0007] Acquire simulation data output by the simulation service platform to form a first data set to be classified; construct a lattice system matching the first data set to be classified;
[0008] mapping a first data set to be classified into the lattice system to form a second data set to be classified;
[0009] Data clustering is performed on the second data set to be classified, and in each round of iterative optimization, the plane factor corresponding to the lattice system is calculated according to the standard deviation of each cluster, and the distance measurement strategy of the corresponding cluster is dynamically adjusted according to the calculated plane factor to achieve data classification.
[0010] On the basis of the above scheme, the present invention also makes the following improvements:
[0011] Furthermore, the plane factor corresponding to the lattice system is calculated based on the standard deviation of each cluster, and the execution is:
[0012] Calculating the volume noise ratio of the lattice system according to the standard deviation;
[0013] According to the standard deviation, all lattice points in the lattice system are traversed to obtain the Theta series;
[0014] According to the volume noise ratio and the Theta series corresponding to the standard deviation, the plane factor corresponding to the corresponding standard deviation is obtained.
[0015] Furthermore, the plane factor ∈_Λ(σ) corresponding to the standard deviation σ is expressed as:
[0016] ∈_Λ(σ)=((VNR / (2π))^(n / 2))*Θ_Λ(t)-1 (1)
[0017] Wherein, VNR represents the volume noise ratio of the lattice system Λ, which is the ratio of the function of the volume V of the lattice system to σ^2; n represents the data dimension of the lattice system; Θ_Λ(t) represents the Theta series, t=1 / (2π*σ^2); the Theta series is obtained by traversing all lattice points λ in the lattice system Λ, calculating the value of exp(-λ^2*t) and accumulating them.
[0018] Furthermore, the distance measurement strategy of the corresponding cluster is dynamically adjusted according to the calculated plane factor, and the following is executed:
[0019] According to the calculated plane factor, it is determined whether the distance measurement strategy of the corresponding cluster needs to be dynamically adjusted; if necessary, the distance measurement strategy of the corresponding cluster after dynamic adjustment is used to update the cluster center and standard deviation of the corresponding cluster.
[0020] Further, according to the calculated plane factor, it is determined whether the distance measurement strategy of the corresponding cluster needs to be dynamically adjusted, and the following is executed:
[0021] Determine whether ∈_Λ(σ) is within the range of [∈_Λmin,∈_Λmax]. If so, keep the distance measurement strategy unchanged; if ∈_Λ(σ) is greater than ∈_Λmax, tighten the distance measurement strategy; if ∈_Λ(σ) is less than ∈_Λmin, relax the distance measurement strategy;
[0022] Among them, ∈_Λmin and ∈_Λmax represent the lower and upper threshold limits without adjusting the plane factor, respectively.
[0023] Further, a lattice system matching the first data set to be classified is constructed by executing:
[0024] Analyzing characteristics of a first data set to be classified and determining matching basis vectors;
[0025] A generator matrix is constructed according to the selected basis vectors; and a lattice system is constructed using the generator matrix.
[0026] Further, the characteristics of the first data set to be classified include statistical characteristics and distribution patterns; the characteristics of the first data set to be classified are analyzed, matching basis vectors are determined, and the following are performed:
[0027] Determining a clustering direction of the first data set by analyzing a distribution pattern of the first data set to be classified;
[0028] Determining the degree of aggregation of the first data set by analyzing statistical characteristics of the first data set to be classified;
[0029] Designing a plurality of different candidate basis vectors according to the aggregation direction and aggregation degree of the first data set, wherein the data dimension of each candidate basis vector is equal to the dimension of the simulation data;
[0030] A plurality of mutually orthogonal basis vectors matching the first data set to be classified are selected from the plurality of candidate basis vectors.
[0031] Further, the characteristics of the first data set to be classified are analyzed to determine the matching basis vectors, and the following steps are also performed:
[0032] The principal component analysis method or the independent component analysis method is used to extract features of the first data set to be classified, and the extracted feature vectors are also used as candidate basis vectors to participate in the selection of basis vectors.
[0033] Further, a generator matrix is constructed according to the selected basis vectors; a lattice system is constructed using the generator matrix, and the following steps are performed:
[0034] Arrange the selected basis vectors in columns to construct an nxn generator matrix M = [m_1, m_2, ..., m_n];
[0035] Define a lattice system Λ={λ=M·i,i∈Z^n} to generate an n-dimensional lattice system Λ; wherein λ represents a lattice point, n represents a data dimension of the lattice system, i represents a scaling factor, and Z^n represents an n-dimensional integer space.
[0036] Further, the first data set to be classified is mapped into the lattice system to form a second data set to be classified, and the following steps are performed:
[0037] Each simulation data in the first data set to be classified is respectively mapped to a corresponding lattice point in the lattice system to form corresponding second simulation data, and all the second simulation data are aggregated to form the second data set to be classified.
[0038] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0039] The simulation service platform data classification method based on dynamic distance measurement provided by the present invention realizes dynamic adjustment of the distance between data points by introducing the concepts of lattice Gaussian distribution and plane factor ∈_Λ(σ), thereby improving the accuracy and efficiency of data clustering. The method is particularly suitable for processing high-dimensional, large-scale or complex-structured data sets, and provides a new technical means for the field of data analysis and mining, which well solves the problem that the simulation service platform data classification method in the prior art is difficult to adapt to complex, high-dimensional, large-scale data sets and has low data classification accuracy.
[0040] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. In the entire drawings, the same reference symbols represent the same components;
[0042] Figure 1 A flow chart of a data classification method based on dynamic distance measurement provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0043] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0044] A specific embodiment 1 of the present invention discloses a data classification method based on dynamic distance measurement, the flow chart of which is as follows: Figure 1 As shown, the method specifically includes the following steps.
[0045] Step S1: Acquire simulation data output by a simulation service platform to form a first data set to be classified; and construct a lattice system matching the first data set to be classified.
[0046] The simulation service platform in this embodiment may involve multiple fields such as general machinery, aerospace, automotive industry, and electronic appliances, and may also involve various forms of simulation such as fluid simulation and electromagnetic simulation. In simulation service platforms in different fields and in different forms, as long as there is a need for data classification for the first data set output by the simulation, this solution is applicable. By classifying the data, a large amount of simulation data can be effectively managed and analyzed, improving the accessibility and analysis efficiency of the data. At the same time, data classification can help discover potential patterns in the data, so as to better understand the distribution patterns of the data. By classifying the simulation data, data features of different categories can also be identified, and then resources can be reasonably allocated to improve the efficiency and effect of the simulation. For example, in the field of electromagnetic simulation, it is necessary to classify the electromagnetic field data at each simulated spatial position to determine how many different electromagnetic field distributions exist in the space.
[0047] Preferably, step S1 can be implemented in the following manner.
[0048] Step S11: Analyze the characteristics of the first data set to be classified and determine matching basis vectors.
[0049] Specifically, the characteristics of the first data set to be classified include statistical characteristics and distribution patterns. The aggregation direction of the first data set is determined by analyzing the distribution pattern of the first data set to be classified; the aggregation degree of the first data set is determined by analyzing the statistical characteristics of the first data set to be classified; according to the aggregation direction and aggregation degree of the first data set, a plurality of different candidate basis vectors are designed, and the data dimension of each candidate basis vector is equal to the dimension of the simulation data; and a plurality of mutually orthogonal basis vectors matching the first data set to be classified are selected from the plurality of candidate basis vectors. In addition, methods such as principal component analysis (PCA) or independent component analysis (ICA) can also be used to assist in selecting basis vectors. That is, the first data set to be classified is subjected to feature extraction using principal component analysis or independent component analysis, and the extracted feature vectors are also used as candidate basis vectors to participate in the selection of basis vectors.
[0050] In the specific implementation process, the statistical characteristics of the first data set to be classified, such as mean, median, variance, standard deviation, skewness and kurtosis, and the distribution mode, such as normal distribution, uniform distribution or other distribution, can be used to analyze the central tendency and dispersion of the data. In addition, the statistical characteristics also include the covariance matrix of the data, which can be used to analyze the linear relationship between different features, which is crucial for selecting basis vectors and constructing lattice systems. The characteristics of the first data set to be classified can guide the selection of basis vectors to ensure that the basis vectors can better reflect the intrinsic structure of the data.
[0051] According to the above method, based on the characteristics of the first data set to be classified, combined with the principal component analysis method or the independent component analysis method, a set of appropriate basis vectors {m_1, m_2, ..., m_n} can be selected, wherein the data dimension of each basis vector is n, which is the same as the data dimension of each simulation data in the first data set. Since these basis vectors are orthogonal to each other and have the same data dimension, the redundancy and error of data representation can be effectively reduced.
[0052] Step S12: construct a generator matrix according to the selected basis vectors; and construct a lattice system using the generator matrix.
[0053] Arrange the selected basis vectors in columns to construct an nxn generator matrix M = [m_1, m_2, ..., m_n]. Ensure that the column vectors of M are linearly independent to form a valid n-dimensional spatial basis. Using the generator matrix M, define the lattice system Λ = {λ = M·i, i∈Z^n} to generate an n-dimensional lattice system Λ, where λ represents the lattice point, i represents the scaling factor, and Z^n represents the n-dimensional integer space. The appropriate scaling factor i can be selected according to the actual situation of the first data set. The lattice system Λ can cover the entire n-dimensional data space, and its distribution and density in space can be adjusted by selecting the basis vectors.
[0054] It should be noted that the lattice, as a discrete set of points, is usually used to represent potential positions in the data space. The lattice system Λ is defined on Z^n, where Z represents an integer set and n represents the data dimension of the lattice system, which can be understood as an n-dimensional integer point set. Each point on the lattice has a unique coordinate representation, and these coordinates are all integers. For the two-dimensional case, the lattice system Λ is a 2x2 matrix defined by a set of linearly independent vector bases. For example, the lattice system can be represented as the unit matrix [10; 01], whose vector bases are (1,0) and (0,1). Therefore, each point on the lattice system Λ represents a potential position in the data space.
[0055] Step S2: Mapping the first data set to be classified into the lattice system to form a second data set to be classified.
[0056] In a specific implementation process, each simulation data in the first data set to be classified is mapped to a corresponding lattice point in the lattice system to form corresponding second simulation data, and all second simulation data are aggregated to form the second data set to be classified.
[0057] During the specific implementation, the lattice system is traversed, and the lattice point with the smallest Euclidean distance to each simulation data in the first data set to be classified is found in the lattice system as the lattice point corresponding to the corresponding simulation data; an index corresponding to each simulation data in the first data set to be classified and the lattice point is established, and each simulation data is mapped to the corresponding lattice point to form the corresponding second simulation data. It should be noted that mapping the simulation data to the lattice system is equivalent to performing a coordinate system conversion, and the coordinate points (corresponding to the lattice points) corresponding to the simulation data after the coordinate system conversion in the lattice system are described as the second simulation data. All second simulation data are aggregated to form the second data set to be classified.
[0058] Step S3: clustering the second data set to be classified, and in each round of iterative optimization, calculating the plane factor corresponding to the lattice system according to the standard deviation of each cluster, and dynamically adjusting the distance measurement strategy of the corresponding cluster according to the calculated plane factor; so as to achieve data classification.
[0059] Specifically, this embodiment can select a suitable clustering algorithm according to the characteristics of the second data set to be classified and the analysis requirements, such as K-means, hierarchical clustering, Gaussian mixture model clustering, etc. The specific implementation process of step S3 is described as follows.
[0060] Step S31: Initialize the number of clusters to be classified, the cluster center of each cluster and the initial distance measurement strategy. In the specific implementation process, the initial cluster center can be selected randomly or according to a certain strategy. At the same time, the initial distance measurement strategy can be implemented by using Euclidean distance for distance measurement.
[0061] Step S32: Calculate the distance between each second simulation data in the second data set and each cluster center according to the current distance measurement strategy, assign each data to the cluster with the shortest distance, and update the cluster center of each cluster according to the data in each cluster.
[0062] Step S33: Calculate the standard deviation of the data in each cluster, and calculate the plane factor corresponding to the lattice system according to the standard deviation of each cluster to dynamically adjust the distance measurement strategy of the corresponding cluster; and use the dynamically adjusted distance measurement strategy as the distance measurement strategy for the next iteration of the corresponding cluster.
[0063] Step S34: Return to step S32 until the clustering termination condition is met (such as the intra-cluster variance reaches the minimum, the number of iterations reaches a preset upper limit, etc.), and the data clustering result is obtained.
[0064] In addition, the quality and effect of clustering results can be evaluated by calculating clustering evaluation indicators (such as silhouette coefficient, Davies-Bouldin index, etc.).
[0065] That is, in this embodiment, a distance measurement strategy is set for each cluster respectively, and the same initial distance measurement strategy can be used during initialization. In subsequent iterations, the distance measurement strategy of the corresponding cluster is updated based on the plane factor calculated based on the standard deviation of the data in each cluster.
[0066] Below, the specific implementation process of step S32 is explained as follows: During the plane factor calculation process, for a given standard deviation σ, a Gaussian distribution is defined on the lattice system Λ. The Gaussian distribution can describe the probability distribution of data points in the lattice space, and can be used as a key parameter to measure the width of the data point distribution, where σ is used to control the width of the distribution. Analyze the characteristics of the Gaussian distribution, especially its performance under different VNRs (volume noise ratios). VNR is defined as the ratio of a specific function of the volume V of the lattice system to σ^2, that is, VNR=f(V) / σ^2. Among them, f(V) is a set function about V, such as f(V)=V^(2 / n).
[0067] In this embodiment, the plane factor ∈_Λ(σ) can be calculated based on the nested lattice theory and mathematical derivation, and the plane factor is used as a key parameter to evaluate the sensitivity of the lattice system Λ to external noise or disturbance under a given σ. The specific calculation process is described as follows: First, the volume V of the lattice system Λ is calculated, and then the volume noise ratio VNR is calculated using V and σ. VNR is used to evaluate the diffusion degree and shape of the Gaussian distribution in the lattice space, providing a basis for subsequent plane factor calculations. Then, based on the VNR and the geometric characteristics of the lattice Λ, the exact value of ∈_Λ(σ) is obtained through complex mathematical operations (such as determinant calculations, series summation, etc.). The specific steps may include:
[0068] Calculate the volume V of the lattice system Λ=|det(Λ)| (the absolute value of the determinant of the lattice system Λ).
[0069] The volume noise ratio VNR is calculated as (V^(2 / n)) / σ^2. VNR reflects the relative relationship between the lattice volume and the Gaussian distribution width.
[0070] Calculate the Theta series Θ_Λ(t), where t=1 / (2π*σ^2), by traversing all lattice points λ in the lattice system Λ, calculating the value of exp(-λ^2*t) and accumulating it to obtain the Theta series.
[0071] Finally, the plane factor is calculated according to ∈_Λ(σ)=((VNR / (2π))^(n / 2))*Θ_Λ(t)-1.
[0072] In addition, in actual calculations, due to the complexity of the Theta series, it can be simplified and approximated as follows. Specifically, the value of the Theta series is approximated by defining a function related to VNR and σ (such as theta_approximation = 1 / (1+exp(-VNR)) in the example). At this time, according to the approximate value of VNR, the Theta series and the mathematical formula, the plane factor ∈_Λ(σ) = ((VNR / (2π))^(n / 2))*theta_approximation-1 is calculated.
[0073] After calculating the plane factor, the distance measurement strategy of the corresponding cluster can be dynamically adjusted according to the calculated plane factor. The specific implementation method is described as follows: according to the calculated plane factor, determine whether it is necessary to dynamically adjust the distance measurement strategy of the corresponding cluster; if necessary, use the dynamically adjusted distance measurement strategy of the corresponding cluster to update the cluster center and standard deviation of the corresponding cluster, that is, use the dynamically adjusted distance measurement strategy to perform data clustering.
[0074] In this embodiment, a set of strategies for dynamically adjusting the distance measurement between data points can be designed based on the calculated ∈_Λ(σ) value to dynamically adjust the distance measurement between data points. Specifically: determine whether ∈_Λ(σ) is within the range of [∈_Λmin,∈_Λmax]. If so, keep the current distance measurement strategy unchanged, such as using Euclidean distance for distance measurement; if ∈_Λ(σ) is greater than ∈_Λmax, tighten the distance measurement strategy, such as using the square of Euclidean distance or higher powers for distance measurement; if ∈_Λ(σ) is less than ∈_Λmin, relax the distance measurement strategy, and use the wth power of Euclidean distance for distance measurement, where w is less than 1. Among them, ∈_Λmin and ∈_Λmax represent the lower threshold and upper threshold of the plane factor without adjusting, respectively. That is, under high VNR, ∈_Λ(σ) is large, indicating that Gaussian distribution can separate data points well, and a more stringent distance metric can be used, such as using the square of the Euclidean distance or higher powers as the distance metric to promote the effective separation of data points; under low VNR, ∈_Λ(σ) is small, the distribution tends to be uniform, and a looser distance metric is used, such as directly using a smaller power of the Euclidean distance to avoid over-segmentation. Therefore, the distance metric in the clustering algorithm can be dynamically adjusted according to the value of ∈_Λ(σ). For example, in the clustering process, different clustering strategies or parameters can be selected according to the size of ∈_Λ(σ) to improve the accuracy and efficiency of clustering.
[0075] In addition, the initialization method of the cluster center or the distance metric in the iteration process can be dynamically adjusted according to the size of ∈_Λ(σ). If ∈_Λ(σ) is large, a farther initial cluster center may be selected, and a stricter distance threshold may be used to update the cluster members in the iteration process. Exemplarily, in each iteration step of the clustering algorithm, the function is called to obtain the value of ∈_Λ(σ), and the relevant parameters in the clustering process are adjusted according to the value, such as the update rule of the cluster center, the distance metric between data points, etc.
[0076] In summary, the simulation service platform data classification method based on dynamic distance metric provided in the present embodiment realizes dynamic adjustment of the distance between data points by introducing the concepts of lattice Gaussian distribution and plane factor ∈_Λ(σ), thereby improving the accuracy and efficiency of data clustering. The method is particularly suitable for processing high-dimensional, large-scale or complex data sets, and provides a new technical means for data analysis and mining. In addition, the simulation service platform data classification method based on dynamic distance metric provided in the present embodiment applies the above-mentioned compression distance metric method based on nested lattice theory to data clustering, performs cluster analysis on the data set through iterative optimization algorithms (such as K-means, hierarchical clustering, etc.), and uses the σ of each cluster to obtain the corresponding ∈_Λ(σ) value, and adjusts the distance metric in the clustering process according to the value, and according to the dynamic distance metric result of the data point, the data point is assigned to the nearest cluster until the cluster termination condition (such as the minimum variance within the cluster, the number of iterations reaches the upper limit, etc.) is met.
[0077] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0078] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A simulation service platform data classification method based on dynamic distance measurement, characterized in that: The method comprises: Acquire simulation data output by the simulation service platform to form a first data set to be classified; construct a lattice system matching the first data set to be classified; mapping a first data set to be classified into the lattice system to form a second data set to be classified; Data clustering is performed on the second data set to be classified, and in each round of iterative optimization, the plane factor corresponding to the lattice system is calculated according to the standard deviation of each cluster, and the distance measurement strategy of the corresponding cluster is dynamically adjusted according to the calculated plane factor to achieve data classification.
2. The method for classifying data of a simulation service platform based on dynamic distance metric according to claim 1, characterized in that: Calculate the plane factor corresponding to the lattice system based on the standard deviation of each cluster, and execute: Calculating the volume noise ratio of the lattice system according to the standard deviation; According to the standard deviation, all lattice points in the lattice system are traversed to obtain the Theta series; According to the volume noise ratio and the Theta series corresponding to the standard deviation, the plane factor corresponding to the corresponding standard deviation is obtained.
3. The method for classifying data of a simulation service platform based on dynamic distance metric according to claim 2, characterized in that: The plane factor ∈_Λ(σ) corresponding to the standard deviation σ is expressed as: ∈_Λ(σ)=((VNR / (2π))^(n / 2))*Θ_Λ(t)-1 (1) Wherein, VNR represents the volume noise ratio of the lattice system Λ, which is the ratio of the function of the volume V of the lattice system to σ^2; n represents the data dimension of the lattice system; Θ_Λ(t) represents the Theta series, t=1 / (2π*σ^2); the Theta series is obtained by traversing all lattice points λ in the lattice system Λ, calculating the value of exp(-λ^2*t) and accumulating them.
4. The method for classifying data of a simulation service platform based on dynamic distance metric according to claim 3 is characterized in that: According to the calculated plane factor, dynamically adjust the distance measurement strategy of the corresponding cluster and execute: According to the calculated plane factor, it is determined whether the distance measurement strategy of the corresponding cluster needs to be dynamically adjusted; if necessary, the distance measurement strategy of the corresponding cluster after dynamic adjustment is used to update the cluster center and standard deviation of the corresponding cluster.
5. The method for classifying data of a simulation service platform based on dynamic distance metric according to claim 4, characterized in that: According to the calculated plane factor, determine whether it is necessary to dynamically adjust the distance measurement strategy of the corresponding cluster and execute: Determine whether ∈_Λ(σ) is within the range of [∈_Λmin,∈_Λmax]. If so, keep the distance measurement strategy unchanged; if ∈_Λ(σ) is greater than ∈_Λmax, tighten the distance measurement strategy; if ∈_Λ(σ) is less than ∈_Λmin, relax the distance measurement strategy; Among them, ∈_Λmin and ∈_Λmax represent the lower and upper threshold limits without adjusting the plane factor, respectively.
6. The method for classifying data of a simulation service platform based on dynamic distance metric according to any one of claims 1 to 5, characterized in that: To construct a lattice system that matches the first data set to be classified, execute: Analyzing characteristics of a first data set to be classified and determining matching basis vectors; A generator matrix is constructed according to the selected basis vectors; and a lattice system is constructed using the generator matrix.
7. The method for classifying data of a simulation service platform based on dynamic distance metric according to claim 6, characterized in that: The characteristics of the first data set to be classified include statistical characteristics and distribution patterns; the characteristics of the first data set to be classified are analyzed, matching basis vectors are determined, and the following are performed: Determining a clustering direction of the first data set by analyzing a distribution pattern of the first data set to be classified; Determining the degree of aggregation of the first data set by analyzing statistical characteristics of the first data set to be classified; Designing a plurality of different candidate basis vectors according to the aggregation direction and aggregation degree of the first data set, wherein the data dimension of each candidate basis vector is equal to the dimension of the simulation data; A plurality of mutually orthogonal basis vectors matching the first data set to be classified are selected from the plurality of candidate basis vectors.
8. The method for classifying data of a simulation service platform based on dynamic distance measurement according to claim 7, characterized in that: Analyze the characteristics of the first data set to be classified, determine the matching basis vectors, and also perform: The principal component analysis method or the independent component analysis method is used to extract features of the first data set to be classified, and the extracted feature vectors are also used as candidate basis vectors to participate in the selection of basis vectors.
9. The method for classifying data of a simulation service platform based on dynamic distance measurement according to claim 8, characterized in that: According to the selected basis vectors, a generator matrix is constructed; using the generator matrix, a lattice system is constructed, and the following is performed: Arrange the selected basis vectors in columns to construct an nxn generator matrix M = [m_1, m_2, ..., m_n]; Define a lattice system Λ={λ=M·i,i∈Z^n} to generate an n-dimensional lattice system Λ; wherein λ represents a lattice point, n represents a data dimension of the lattice system, i represents a scaling factor, and Z^n represents an n-dimensional integer space.
10. The method for classifying data of a simulation service platform based on dynamic distance measurement according to claim 9, characterized in that: Mapping the first data set to be classified into the lattice system to form a second data set to be classified, executing: Each simulation data in the first data set to be classified is respectively mapped to a corresponding lattice point in the lattice system to form corresponding second simulation data, and all the second simulation data are aggregated to form the second data set to be classified.