Virtual winning neuron-based SOM model clustering method
By using PCA to initialize the weight matrix in the SOM algorithm and introducing the virtual winning neuron mechanism, the problems of traditional SOM algorithms in complex data distribution and sensitivity are solved, and a more efficient and stable clustering effect is achieved.
Patent Information
- Application Number
- CN202411830964.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional SOM algorithms are difficult to produce good clustering results when processing complex data distributions, and are highly sensitive to input data, are susceptible to noise interference, and random initialization of the weight matrix leads to slow convergence speed and local optimal solution problems.
The weight matrix is initialized by principal component analysis (PCA) and a virtual winning neuron mechanism is introduced. By calculating the weighted average Euclidean distances of multiple neuronal nodes with higher similarity, the weight matrix is updated to more comprehensively characterize the characteristics of the input data.
It improves the accuracy and stability of clustering results, reduces the dependence on the initial weight matrix, reduces the consumption of training time and computing resources, and enhances the robustness of the algorithm.
Smart Images

Figure CN120011833A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of machine learning in artificial intelligence, and in particular relates to a clustering method based on a virtual winning neuron SOM model, which is used for anomaly detection problems of sample data. Background Art
[0002] Machine learning (ML) is the process of automatically learning general rules or knowledge from limited observation samples through computers, and generalizing and generalizing them to unknown observation samples. In recent years, ML technology has developed rapidly and played an important role in agronomy, industry, economics, medicine and other fields. It is often used for natural language processing, image processing, and data analysis. When using ML technology to analyze data, cluster analysis, anomaly detection, association rule mining, and predictive modeling are tasks that often need to be completed.
[0003] In recent years, data anomaly detection has become an indispensable part of modern hydropower station management. Hydropower station equipment such as turbines, generators, transformers, pumps, etc. may experience mechanical failures, electrical failures or other problems during long-term operation. These problems are often gradual and not easy to appear in a short period of time. Through data anomaly detection, the operating status of the equipment can be monitored in real time and the potential failure risks of the equipment can be identified in advance. For example, excessive temperature, abnormal vibration, unstable flow, etc. may be precursors to failures. By analyzing the equipment operation data, estimating the health status of the equipment, and discovering anomalies in time, it is helpful to carry out preventive maintenance and prevent sudden equipment failures. Moreover, hydropower stations, as important infrastructure, are usually located in remote areas. Equipment maintenance and repair may be relatively difficult and costly. Data anomaly detection can issue early warnings in time when equipment anomalies occur, thereby avoiding serious equipment failures, reducing the risk of unexpected downtime or equipment damage, reducing maintenance costs, and ensuring the long-term stable operation of equipment. In addition, modern hydropower stations are usually equipped with remote monitoring systems. Through sensors and data acquisition equipment, equipment operation data is collected in real time. Combined with anomaly detection algorithms, intelligent operation and maintenance can be realized. Data anomaly detection provides hydropower stations with a large amount of equipment operation data and fault information, which can provide a scientific basis for subsequent equipment management, optimized scheduling, maintenance plans, etc. Operators can also grasp the operating status of equipment at any time through the remote monitoring platform and respond quickly when an abnormality occurs. This automated and intelligent monitoring method not only improves operation and maintenance efficiency, but also reduces labor costs and improves management level.
[0004] Anomaly detection refers to the discovery of data patterns that are much smaller than data clusters, that is, objects in the data set that are significantly different from other data. Most traditional data anomaly detection methods use a certain statistical distribution for modeling, and then use this model to determine whether the distribution of data points is abnormal; or set a threshold for a single amount of measurement data, and then determine the singularity based on the threshold to find the "dirty data". However, since real data is more complex and does not necessarily conform to any statistical distribution model under ideal conditions, the results obtained using traditional anomaly detection methods are relatively one-sided, it is difficult to distinguish the type of outliers, and it is impossible to make joint judgments. In order to achieve high-quality data mining and efficiently and accurately monitor abnormal data, researchers at home and abroad have combined ML technology with data anomaly detection in recent years to improve the simple modeling problems of traditional algorithms.
[0005] According to the information provided by the training samples and the different feedback methods, the ML techniques used for anomaly detection can be divided into two categories: supervised learning and unsupervised learning. Among them, the supervised anomaly detection method uses label information to establish the mapping relationship between the input data and its corresponding label by building a classification or regression model. Different mapping relationships reflect the "normal" and "abnormal" states of the data; then the model can judge whether it is normal or abnormal data based on the mapping relationship between the test data and its label. The classic supervised learning algorithms mainly include decision trees (DTs), linear regression (linear regression), K-nearest neighbor (KNN), support vector machine (linear support vector machine, SVM), etc. Using supervised learning for anomaly detection can make full use of data label information and improve the accuracy of the model; however, when the amount of abnormal data is small, the classification category imbalance problem will occur, and when the data set label is missing, the performance of the supervised learning algorithm will decrease, which will affect the detection results.
[0006] In real life, the labels of most data sets are not completely known, and due to the lack of prior knowledge, it is difficult to label all data through manual operations, and it is impossible to train supervised models. Therefore, unsupervised learning is more widely used in practice. Unsupervised learning can analyze and process unlabeled data sets, that is, without any prior knowledge, it can classify or distinguish them by learning the characteristic rules between data. Commonly used unsupervised algorithms are clustering and dimensionality reduction. Clustering algorithms include k-means clustering, hierarchical clustering, DBSCAN clustering, etc.; Dimensionality reduction algorithms include principal component analysis (PCA), singular value decomposition (SVD), etc. When most samples in the data set are normal values and there is a large difference between the abnormal values and the normal values, the unsupervised anomaly detection method can be used. There is no need to use labels for correction during the algorithm training process. By finding the characteristic relationship between the data, the samples that are most mismatched with other data are detected as abnormal test data. The methods for judging anomalies based on clustering ideas include: data that does not belong to any cluster is anomaly, data that is far from the cluster center is anomaly, and data that belongs to a cluster with few or sparse data points is anomaly; and the dimension reduction algorithm is often used in the data preprocessing stage to perform feature analysis on the data and optimize the training set. Data anomalies are usually reflected in the distance between data, distribution density, degree of deviation, etc., and suitable unsupervised learning models can be selected according to different anomalies.
[0007] As a classic clustering model in unsupervised learning, the Self-Organizing Map (SOM) can be used to achieve dimensionality reduction, data visualization, clustering, classification, feature extraction and other purposes. By training the SOM network, the input data is mapped to the output layer nodes, and the output layer nodes have a certain association according to the distance, so as to achieve the learning of the characteristics of the data itself and the topological relationship between the data. The traditional SOM algorithm contains two layers: the input layer and the output layer, where the output layer is also called the competition layer, so the self-organizing map network is also called the self-organizing competition network. The dimension of the input layer is the same as the dimension of the input feature; the output layer contains multiple neuron nodes to form a two-dimensional matrix. The input nodes are fully connected to the output layer neurons, and the connecting lines are the weight edges. After training and learning, the nodes of the output layer have a certain association according to the distance; by learning a set of weights, the input data can be mapped to the output layer nodes, and the output layer nodes represent the characteristics of the input data and retain the topological structure between the input data, so that the data with associations in the input space can also be gathered in one place in the output space to form a set of associated neurons. This is because the weight matrix will form a clustering area near each neuron after multiple learning. After multiple iterations, the neuron weights in the clustering area remain consistent with the input data or approach the input data, so that input data with similar characteristics will be concentrated on several adjacent neurons, thus achieving clustering.
[0008] The traditional SOM algorithm is easy to implement and has low computational complexity, but it has strict requirements on data distribution. For data distribution that does not conform to Euclidean distance or linear relationship, the algorithm may not produce good clustering results. Using a single winning neuron to train the model is not conducive to interpreting the input data from multiple angles and cannot fully represent the characteristics of all aspects of the data. Changes in a single winning neuron will cause the algorithm to be more sensitive to the input data and easily affected by the noise of the input data. Randomly initializing the weight matrix can easily slow down the convergence of the SOM network and cause it to fall into a local optimal solution when exploring the parameter space. Summary of the invention
[0009] In order to solve the deficiencies in the prior art, the present invention provides a clustering method based on a virtual-winner SOM model (virtual-winner SOM, vwSOM). Firstly, the principal component analysis (PCA) method is used to obtain an initial weight matrix, so that the initial weights can better capture the main features of the data, thereby improving the clustering effect. Subsequently, when the new input sample data is mapped to the output layer, multiple neuron nodes with high similarity are selected in the weight matrix, and the virtual winning neurons are calculated and updated accordingly, so as to comprehensively characterize the features of the input data within the smallest possible error range. Finally, the optimization goal of improving the accuracy and stability of the clustering results is achieved.
[0010] The present invention adopts the following technical solution:
[0011] A clustering method based on a virtual winning neuron SOM model specifically comprises the following steps:
[0012] Assume that the input layer of each SOM model consists of D nodes, where D is the same as the input feature dimension; the output layer contains M neuron nodes, forming a two-dimensional matrix, where M = X*Y. According to the empirical formula, Where N is the number of training samples.
[0013] Step 1: Sample data preprocessing: Input training data datas and perform normalization and regularization preprocessing to make the sample data features have the same measurement scale and avoid data overfitting.
[0014]
[0015] Step 2: Use PCA to initialize the weight matrix. By performing PCA analysis on the preprocessed data, the eigenvectors of the data set are calculated. These eigenvectors are the directions with the largest variance in the data set. According to the number of principal components selected, each principal component is used as an initial weight vector to initialize the SOM weight matrix. Data matrix X
[0016]
[0017] Calculate the covariance matrix S of X,
[0018]
[0019] Perform eigenvalue decomposition on the covariance matrix S,
[0020] S=PΛP T (5)
[0021] Find the eigenvalues and their eigenvectors and arrange them in descending order. Take the eigenvectors corresponding to the non-zero eigenvalues to form the matrix P. There exists
[0022]
[0023] Where Λ is a diagonal matrix containing decreasing non-negative real eigenvalues λ1≥λ2≥...≥λ d ≥0. Select the eigenvectors corresponding to the first k eigenvalues as the principal components, where k is the dimension after dimensionality reduction. The new data matrix T (k-dimensional) after principal component analysis can be obtained.
[0024] T=PX (7)
[0025] The extracted principal components are used to initialize the weight matrix of the SOM model, and each principal component is used as an initial weight vector.
[0026] Step 3: Calculate the virtual winning neuron to update the weight matrix. Select the three nodes closest to the input sample as the winning neurons, calculate the distance weights of these three neurons, and find the weighted average Euclidean distance d between the input sample and the three neuron nodes. s1 , will differ from the eigenvalue of x1 by d s1 Set a virtual winning neuron s1 at , and update the weights of other nodes through the distance relationship between the virtual node and other nodes in the output layer. Select the three winning neurons closest to the input sample x1 in Euclidean distance, and calculate the distances between the current neuron and these three winning neurons, which are d1, d2, and d3 respectively; use the exponential decay function to calculate the weights of these three distances, so that the relatively far winning neurons contribute less to the final weighted distance, and the relatively close winning neurons contribute more to the final weighted distance. The degree of attenuation can be controlled by adjusting the parameters in the exponential decay function. For example:
[0027]
[0028] According to the distance weights obtained above, the Euclidean distances corresponding to the three winning neurons are weighted averaged to obtain the distance between the virtual neuron node s1 and the current node.
[0029]
[0030] Traverse each neuron in the output layer and update its weight vector w ij ,
[0031]
[0032] Among them, δ is the neighborhood attenuation function, d ijRepresents the distance between the input sample x1 and the neuron node at position (i, j) in the output layer.
[0033] Compared with the prior art, the present invention has the following advantages and effects:
[0034] 1. The present invention adopts the principal component analysis (PCA) initial weight matrix, so that the initial weight can better capture the main features of the data, thereby improving the clustering effect; then, when the new input sample data is mapped to the output layer, multiple neuron nodes with high similarity are selected in the weight matrix, and the virtual winning neurons are calculated and updated based on this, so as to fully characterize the characteristics of the input data within the smallest possible error range.
[0035] 2. The present invention initializes the weight matrix by PCA, thus overcoming the disadvantages of many model training iterations and low clustering accuracy when the weight matrix is randomly initialized. According to the number of required principal components, the corresponding eigenvector is selected as the initial weight vector, providing a good initial state for the training of the model, reducing the influence of the initial weight matrix on the clustering effect; it helps to accelerate the convergence process and reduce the consumption of training time and computing resources.
[0036] 3. The present invention changes the number of winning neurons based on the single winning neuron mechanism of the traditional SOM model and proposes the concept of virtual neurons. By observing the Euclidean distance array between the input sample and each point in the output layer, it is found that the difference between the Euclidean distance of the winning neuron node and the Euclidean distance of the two surrounding neurons is very small, which indicates that the three neurons have similar representation degrees for the input sample. Inspired by the virtual node, the SOM algorithm is improved, and the three nodes closest to the input sample are selected as the winning neurons. By calculating the distance weights of these three neurons, the weighted average Euclidean distance d between the input sample and the three neuron nodes is calculated. s1 , will differ from the eigenvalue of x1 by d s1 A virtual winning neuron s1 is set at , and the weights of other nodes are updated through the distance relationship between the virtual node and other nodes in the output layer. By changing the number of winning neurons, it is expected to grasp more information with a smaller error, more comprehensively characterize the characteristics of the input data, and improve the clustering accuracy of the algorithm. In addition, compared with the idea of a single winning neuron, the multiple winning neuron algorithm has a good comprehensive feature extraction of the input data and strong algorithm stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is the SOM structure diagram.
[0038] Figure 2 It is the flow chart of vwSOM model training of the present invention.
[0039] Figure 3 It is a visualization of Iris clustering results using different algorithms.
[0040] Figure 4 It is a visualization of Wine clustering results using different algorithms. Specific implementation methods
[0042] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0043] Figure 2 This is a flow chart of model training of the present invention.
[0044] Next, we will use the Iris dataset and the Wine dataset as the data to be analyzed to analyze the specific implementation method of the vwSOM clustering method, and combine it with Figure 3 , Figure 4 The effect comparison of clustering analysis of data by vwSOM clustering method, DBSCAN algorithm and hybrid SOM algorithm is described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.
[0045] The real labels of Iris and Wine are introduced to calculate the accuracy (ACC), precision (P), recall (R), F1 and other indicators of each algorithm to evaluate the performance of the algorithm. Among them, ACC refers to the ratio of all correctly predicted samples to the total number of samples. It measures the overall prediction accuracy of the model, but in the case of imbalanced sample categories, high accuracy will lose its meaning; P refers to the ratio of the number of positive samples correctly predicted by the model to the number of all predicted positive samples. It measures how many of the samples predicted as positive by the model are truly positive samples. The higher the accuracy, the lower the misjudgment rate of the model; R refers to the ratio of the number of positive samples correctly predicted by the model to the number of actual positive samples. It measures the model's ability to identify positive samples. The higher the recall, the higher the proportion of positive samples correctly identified by the model; Accuracy and recall are often contradictory. Improving one indicator may lead to a decrease in the other indicator, so the F1 indicator is introduced to measure model performance. F1 is the weighted harmonic mean of accuracy and recall, which comprehensively considers both accuracy and recall. Since ACC measures the overall prediction accuracy of the model, the experiment was repeated 1000 times under the same conditions, and the variance (VA) of ACC was calculated to measure the stability of the algorithm's data clustering effect; the average operation time (Average Operation Time, AOT) refers to the average time used by the model to process each test data. This indicator is used to measure the complexity of the algorithm.
[0046] (1) Performance analysis of the embodiment:
[0047] Figure 3 The clustering results of different methods on the Iris training set and test set are shown; Figure 4The clustering results of different methods on the Wine training set and test set are shown; by calculating the Pearson correlation coefficient that characterizes the linear correlation between the four features in the Iris dataset and the class, it is found that the petal length and petal width have a high degree of class correlation and can better characterize the class, so these two features are used as the horizontal and vertical axes in the cluster diagram to determine the data position; by calculating the Pearson correlation coefficient of the linear correlation between the 13 features in Wine and the class, it is found that the alcohol and malic acid components in the wine can better characterize the class, so these two features are selected as the horizontal and vertical axes of the scatter plot; points of different colors in the cluster diagram represent different clusters. From observing the scatter plot, we can see that when the data are represented by features with high class correlation, they present different distribution characteristics, that is, different features locate the position of the data on the scatter plot; when the model is trained with the training set, different colors represent different types of data, and then the test data is input to test its clustering effect. It can be found that the test data is clustered into the same class of the training data with the same characteristics, which is basically consistent with the evaluation index value obtained from the experimental results; the number of points presented in the scatter plot is slightly less than the number of training data and test data set by the experimental code. This is because when the features of different data are extremely close or completely consistent, they will appear in the same position on the scatter plot, and because of the consistency of features, they will be clustered into the same category, and the colors presented in the scatter plot are also consistent.
[0048] Table 1 shows the results of various indicators obtained by clustering the Iris dataset using the DBSCAN algorithm, the hybrid SOM algorithm, and the vwSOM algorithm; Table 2 shows the results of indicators obtained by clustering the Wine dataset using each algorithm. When the structure of the processed dataset is complex, if only distance similarity is used for clustering, the effect is not good, so the algorithm using density similarity for clustering came into being. DBSCAN is a classic density-based clustering algorithm. Its basic idea is: divide the sample data into core points, boundary points, and outliers according to their distribution density; define clustering core points and boundary points by density direct access, density reachability, and density connection to form different clusters; however, when it processes larger datasets, its convergence speed is slow and the clustering effect is not good when the density distribution of the dataset is uneven. The structural diagram of the traditional SOM model is shown in the figure. Figure 1 As shown in the figure, the hybrid SOM algorithm combines the traditional SOM model with the KNN model. The specific content is: the SOM algorithm is trained on a healthy data set with a small amount of noise. After fitting the healthy data, the nodes that are too sparse or below the minimum BMU threshold are removed to avoid BMU being contaminated by noise; a layer of KNN model is added after the SOM model to classify the data based on the Euclidean distance between the centroid and the observed data points; the test data is input into the trained KNN, and the outliers are judged according to the Euclidean distance between the test data and the nodes in the cluster.
[0049] Table 1 Comparison of clustering effects of different algorithms on the Iris dataset
[0050] P R F1 ACC VA AOT / ms DBSCAN 0.85 0.84 0.84 0.8800 0.0025 0.0244 Hybrid SOM 0.93 0.90 0.91 0.9395 0.0013 0.0675 wxya 0.94 0.93 0.93 0.9412 0.0012 0.2221
[0051] Table 2 Comparison of clustering effects of different algorithms on the Wine dataset
[0052] P R F1 ACC VA AOT / ms DBSCAN 0.85 0.84 0.84 0.8700 0.0038 0.0251 Hybrid SOM 0.92 0.91 0.91 0.9268 0.0019 0.0679 wxya 0.94 0.92 0.93 0.9401 0.0014 0.2263
[0053] Since the Iris data is distributed according to the characteristics, the boundaries between the categories are not strong, the hyperparameter selection of the DBSCAN algorithm is difficult, and it is difficult to determine the size of the cluster, so it is not easy to determine the boundary points. Its F1 index and ACC are 0.84 and 0.88 respectively, both lower than the vwSOM algorithm; cluster analysis of the Wine data set, affected by the uneven distribution of samples, the ACC value is reduced and the VA value is increased, indicating that the clustering stability is reduced. However, this algorithm has low complexity and is easy to implement. The processing time for a single data in the Iris data set is only 0.0244ms, which is suitable for real-time data processing scenarios with low accuracy requirements.
[0054] The hybrid SOM algorithm is used to cluster the Iris dataset, and the ACC is 0.9395; the F1 index is 0.91, which is lower than the result of the vwSOM algorithm. This is because its R value is low, that is, the number of incorrectly predicted negative samples is large; when it processes the Wine dataset, the ACC value is slightly reduced. The hybrid SOM algorithm has good scalability. It reduces the number of algorithm inputs and outputs through threshold judgment, reduces the calculation time, and reduces the complexity of the algorithm, but it will also lose some effective information.
[0055] The F1 index of the vwSOM algorithm for clustering Iris is 0.93, the ACC is 0.9412, and the VA value is relatively low, indicating that its clustering accuracy for data is high and its performance is relatively stable; the F1 value obtained for clustering Wine is still 0.93, and the ACC and VA values are basically unchanged, indicating that the uneven distribution of samples of different classes in the data set has little impact on the vwSOM algorithm; the algorithm takes 0.2221ms to process a single Iris data and 0.2263ms to process a single Wine data. The AOT value increases with the increase in the number of data features. Although it is slightly larger than the AOT of the above two algorithms, it meets the requirements of real-time data processing and achieves a good balance between algorithm complexity and calculation accuracy, and can be applied to data anomaly processing in real life.
[0056] The vwSOM model uses the PCA algorithm to initialize the weight matrix and add virtual winning neurons on the basis of the traditional SOM model to change the traditional single winning neuron mechanism into a multi-win mechanism. The ablation experiment was used to observe the role of each module of the vwSOM in terms of algorithm accuracy and stability. The results are shown in Table 3. Taking the Iris data set as an example, the ACC of the traditional SOM algorithm is 0.8986, the VA is 0.0026, and the R value is as low as 0.85; the PCA algorithm is used to obtain the direction with the largest variance in the data set, and each principal component is used as an initial weight vector to obtain the weight matrix of the initialized SOM model, which provides a good initial state for the training of subsequent models and accelerates the convergence of the algorithm. The P value is high, the ACC is 0.8999, and the VA is significantly reduced to 0.0017, indicating that the stability of the algorithm has been improved; the single winning neuron mechanism in the traditional SOM algorithm is innovated, and the idea of virtual winning neurons is introduced. The algorithm F1 value is higher, the ACC is significantly improved, and its value is 0.9385, and the VA value is 0.0020, which is lower than the traditional algorithm. This is because by changing the number of winning neurons, it is possible to grasp more information with a smaller error, so that the algorithm can more comprehensively characterize the characteristics of the input data. By analyzing the experimental data, it can be observed that PCA can improve the stability of the algorithm; the idea of virtual winning neurons obtained from multiple winning neurons contributes to the improvement of algorithm accuracy; combining the two and applying them to traditional SOM, we get the vwSOM algorithm, whose F1 value is 0.93, ACC value is 0.9412, and VA is 0.0012, indicating that the algorithm has high accuracy and good stability.
[0057] Table 3 Ablation experiment
[0058] P R F1 ACC VA SOM 0.88 0.85 0.86 0.8986 0.0026 SOM+PCA 0.91 0.88 0.89 0.8999 0.0017 SOM+virtualwinner 0.94 0.92 0.93 0.9385 0.0020 wxya 0.94 0.93 0.93 0.9412 0.0012
[0059] (2) Conclusion:
[0060] A clustering method based on virtual winning neuron SOM model is proposed. The weight matrix is initialized by PCA algorithm, and the single winning neuron mechanism in traditional SOM is improved to multiple winning neurons, and the idea of virtual neuron node is proposed. The training set data is preprocessed, and the weight matrix is initialized by PCA algorithm. The corresponding feature vector is selected as the initial weight vector according to the number of principal components required. The weight matrix constructed in this way is used for subsequent model training; the sample data is traversed, and the Euclidean distance between the sample and each neuron node is calculated, that is, the Euclidean distance between the sample feature and each weight vector is calculated; the first three neurons with the smallest distance are selected as winning neurons, and the value obtained by weighted average of the three distances is used as the distance between the sample data and the virtual neuron node mapped on the output layer; each weight vector is updated according to the distance relationship between the virtual neuron node and other nodes; when the training set data is traversed, the weight matrix is updated; the model is iterated multiple times to make the weight matrix infinitely close to the characteristics of the sample data. The implementation plan results show that the method proposed by the present invention achieves better clustering effect and high clustering accuracy; the model is less sensitive to input data and has better stability.
[0061] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A clustering method based on a virtual winning neuron SOM model, characterized in that: The following steps are included: Step 1: Sample data preprocessing: input training data and perform normalization and regularization preprocessing to make the sample data features have the same measurement scale and avoid data overfitting. Step 2: Use PCA to initialize the weight matrix. Perform PCA analysis on the preprocessed data datas to calculate the eigenvectors of the data set. These eigenvectors are the directions with the largest variance in the data set. According to the number of principal components selected, each principal component is used as an initial weight vector to initialize the SOM weight matrix. The data matrix X Calculate the covariance matrix S of X, Perform eigenvalue decomposition on the covariance matrix S, S=PΛP T (5) Find the eigenvalues and their eigenvectors and arrange them in descending order. Take the eigenvectors corresponding to the non-zero eigenvalues to form the matrix P. There exists Where Λ is a diagonal matrix containing decreasing non-negative real eigenvalues λ1≥λ2≥...≥λ d ≥0, select the eigenvectors corresponding to the first k eigenvalues as the principal components, where k is the dimension after dimensionality reduction, and obtain the new data matrix T (k-dimensional) after principal component analysis T=PX (7) Use the extracted principal components to initialize the weight matrix of the SOM model, and use each principal component as an initial weight vector; Step 3: Calculate the virtual winning neuron to update the weight matrix, select the three nodes closest to the input sample as the winning neurons, and calculate the distance weights of these three neurons to find the weighted average Euclidean distance between the input sample and the three neuron nodes. The difference between the eigenvalue of x1 and A virtual winning neuron s1 is set at , and the weights of other nodes are updated through the distance relationship between this virtual node and other nodes in the output layer; Select the three winning neurons with the closest Euclidean distance to the input sample x1, and calculate the distances between the current neuron and the three winning neurons, which are d1, d2, and d3 respectively; use the exponential decay function to calculate the weights of these three distances, so that the relatively far winning neurons contribute less to the final weighted distance, and the relatively close winning neurons contribute more to the final weighted distance. The degree of attenuation is controlled by adjusting the parameters in the exponential decay function, as follows: According to the distance weights obtained above, the Euclidean distances corresponding to the three winning neurons are weighted averaged to obtain the distance between the virtual neuron node s1 and the current node. Traverse each neuron in the output layer and update its weight vector w ij , Among them, δ is the neighborhood attenuation function, d ij Represents the distance between the input sample x1 and the neuron node at position (i, j) in the output layer.