A High-Dimensional Sparse Image Representation Method Based on a Deep Semi-Supervised Learning Framework
Through the semi-DRRBM model of the deep semi-supervised learning framework, the problem of large occupancy of computing resources for high-dimensional sparse data and long training time is solved, and efficient low-dimensional feature extraction and clustering performance improvement is achieved.
Patent Information
- Application Number
- CN202310422895.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-04-19
AI Technical Summary
The prior art uses a large amount of computing resources when processing high-dimensional sparse data and the classification results are affected by the scarcity of labels. The traditional method is not suitable for sparse data with independent features, and the training model takes a long time. The pcGRBM model is not suitable for high-dimensional data modeling and has limited ability to learn hidden features.
Using a semi-DRRBM model based on a deep semi-supervised learning framework, high-dimensional vectors are reconstructed through semi-DRRBM objective function and Gibbs sampling, combining supervised information and reconstruction constraints, model parameters are updated to extract low-dimensional and dense hidden features.
Effectively extract low-dimensional and dense hidden features in high-dimensional sparse data, improving the accuracy and efficiency of data representation, enhancing the learning ability and clustering performance of the model, and shortening the training time.
Smart Images

Figure CN116452824B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning images, and particularly to a high-dimensional sparse image representation method based on a deep semi-supervised learning framework. Background Art
[0002] High-dimensional sparse data often appears in industrial applications and recommendation systems. However, the representational learning of high-dimensional sparse data has always been a difficult problem that requires huge computing resources in terms of its calculation and memory. At the same time, the scarcity of labels will have a negative impact on the classification results. For sparse high-dimensional data, dimensionality reduction methods are usually used, that is, mapping the high-dimensional data space to a low-dimensional feature space. However, these methods are sometimes not applicable to sparse data with independent features. The restricted Boltzmann machine (RBM) is an energy-based binary representation learning model. Representational learning is a method for extracting useful knowledge and internal structure of high-dimensional sparse data. That is, the concept factorization (CF) with locality constraints is called local coordinate concept factorization (LCF), which is an effective method for learning the sparse representation of high-dimensional data. Deep learning is very popular in the field of single-atom level determination. Taking the recently popular carbon allotrope graphene as an example, a graphene atomic-level structure image dataset is proposed. These datasets have high-dimensional sparse features, and machine learning is combined to obtain their excellent low-dimensional representations.
[0003] To learn sparse representations, two approximate local linear representation (ALLR) algorithms with probability simplex constraints and symmetric constraints are proposed respectively. These algorithms perform better when learning high-dimensional data on public datasets. They perform somewhat inadequately on high-dimensional data of atomic structures. The latent factor (LF) model is a successful method for accurately representing high-dimensional and sparse (HiDS) matrices. However, there are usually some abnormal data in these matrices, and the training model takes a long time. The pcGRBM model based on GRBM is used to study the semi-supervised representation learning of continuous data. The pcGRBM model is not suitable for modeling high-dimensional data and has limited ability to learn hidden features. Summary of the Invention
[0004] In order to extract low-dimensional and dense hidden features from high-dimensional and sparse images, the present invention provides a high-dimensional sparse image representation method based on a deep semi-supervised learning framework, including the following steps:
[0005] A high-dimensional sparse image representation method based on a deep semi-supervised learning framework, characterized in that it includes the following steps:
[0006] The input of the semi-DRRBM objective function is:
[0007]
[0008] is the supervision information of the visible high-dimensional layer, where is the visible layer vector of the semi-DRRBM model, N represents the number of all instances, and each visible layer vector has a low-dimensional hidden feature , and then a high-dimensional vector is reconstructed using one-step Gibbs sampling during the training process , is 's low-dimensional hidden feature, is the set of reconstructed high-dimensional vectors;
[0009] 1) Initialization ; The objective function of the semi-DRRBM is defined as follows:
[0010] (Equation 1)
[0011] , which is used to update its value according to this algorithm, where W=(w j,i )∈RN×N represents the connection weight vector between the hidden layer and the visible layer, b is the visible layer bias vector, a is the hidden layer bias vector, i, j represent an arbitrarily chosen or specified natural number, and when used as a subscript, it represents the position in the array. ξ∈ (0, 1) represents the adjustment parameter, , is the energy function of the RBM model, , σ is a logistic sigmoid function , N represents the number of all instances, and D represents the original data set;
[0012] represents the visible layer vector. Each visible layer vector has a low-dimensional hidden feature , and then a high-dimensional vector is reconstructed using one-step Gibbs sampling during the training process , is 's low-dimensional hidden feature, where the subscript k represents the size of the array included in the vector. When using other subscripts, the meanings of v, v’, h, h’ are the same. The relationships between the subscripts p, q and the subscripts m, n are as described below:
[0013] Let ∀ and ∀ both belong to the k-th cluster (k = 1,2,...,K, where K represents the total number of clusters). All these unordered pairs ( , ) forms an internal cluster set I, and |I| represents the size of I;
[0014] ∀ and ∀ belong to the i-th and j-th clusters (i ≠ j) respectively. All unordered pairs ([[]] , ) form an internal cluster set O, and |O| represents the size of O;
[0015] Meanwhile, ∀ and ∀ have a one-to-one mapping and , ( , ) form a set I’, and |I’| represents the size of I’;
[0016] ∀ and ∀ have a one-to-one mapping and , ( , ) form a set O’, and |O’| represents the size of O’;
[0017] 2) Obtain the sample hidden layer features through ;
[0018] 3) Use to reconstruct the visible layer data;
[0019] 4) Use to obtain the hidden layer features of the sample reconstruction;
[0020] 5) Calculate the partial derivative of ;
[0021] 6) For the previous hidden layer data h, visible layer data v, reconstructed hidden layer data h’, visible layer data v’, and the parameter obtained during the i-th calculation, calculate the for this calculation;
[0022] 7) When the new parameter is obtained through step 6, use it as the new input and repeat steps 2 - 6 until the maximum value is reached, which can be stopped by repeating to the specified number of times or when it is less than a certain error specified by humans;
[0023] 8) Return: The parameters of the Semi-DRRBM model , and substituting its value into the Semi-DRRBM model can complete dimensionality reduction.
[0024] Preferably, in step 2) ,
[0025] Define the function
[0026] (Formula 2)
[0027] (Formula 3)
[0028] ;
[0029] For the partial derivative is:
[0030] (Formula 4)
[0031] where v mi is the element at the i-th position of vector v m , v ni is the element at the i-th position of vector v n , h mj , h nj are respectively the sum of the elements at the j-th position of vector h m and vector h n ; v pi is the element at the i-th position of vector v p , v qi is the element at the i-th position of vector v q , h pj , h qj are respectively the elements at the j-th position of vector h p and vector h q ;
[0032] For the partial derivative:[[]]END]]
[0033] (Formula 5)
[0034] v mi ’ is the element at the i-th position of vector v m ’, v ni ’ is the element at the i-th position of vector v n ’, h mj ’, h nj ’ are respectively the sum of the elements at the j-th position of vector h m ’ and vector h n ’; v pi ’ is the element at the i-th position of vector v p ’, v qi ’ is the element at the i-th position of vector vq The element at the i-th position of pj ’, h qj ’ and h p ’ are the elements at the j-th position of the vector h q ’ and the vector h
[0035] For Partial derivative with respect to:
[0036] (Equation 6)
[0037] (Equation 7)
[0038] Because is independent of a, so their partial derivatives with respect to a j are 0;
[0039] (Equation 8).
[0040] Preferably, in step 6), the previous hidden layer data h and visible layer data v, and the reconstructed hidden layer data h' and visible layer data v', and the parameters obtained in the i-th calculation are used to calculate the for this calculation; where the update rule of the semi-DRRBM model parameters takes the following form:
[0041] (Equation 9)
[0042] W ij (τ) The connection weight at the -th iteration, W ij (τ+1) The connection weight at the -th iteration;
[0043] ε represents the learning rate, ξ ∈ (0, 1) represents the adjustment parameter, and each visible layer vector has a low-dimensional hidden feature ;
[0044] (Equation 10)
[0045] b j (τ) The bias of the visible layer units at the -th iteration, b j (τ+1) The bias of the visible layer units at the -th iteration.
[0046] (Formula 11)
[0047] where < >0 and < >1 represent the expected values of the subscript data and the reconstructed data respectively;
[0048] a j (τ) At the th calculation, the bias of the hidden layer unit, a j (τ+1) At the th calculation, the bias of the hidden layer unit.
[0049] Advantages of the present invention: Based on the RBM model, the present invention proposes a new variant model of semi-DRRBM (semi-supervised dimensionality reduction RBM model), which can extract low-dimensional and dense hidden features from high-dimensional and sparse data sets. During the representation learning process of semi-DRRBM, tiny constraints of positive and negative samples are constructed in the high-dimensional visible layer. The positive sample constraint increases the similarity of the same graphene atomic structure in the low-dimensional dense space of the hidden layer. At the same time, the negative sample constraint expands the difference of different graphene atomic structures in the low-dimensional hidden layer. Description of the Drawings
[0050] Figure 1 It is the representation learning process of the Semi-DRRBM model according to the embodiment of the present invention;
[0051] Figure 2 It is the semi-supervised dimensionality reduction process of the deep semi-RL structure according to the embodiment of the present invention;
[0052] Figure 3 It is the T-SNE visualization of the original data with true labels according to the embodiment of the present invention;
[0053] Figure 4 It is the T-SNE visualization of the hidden representation of the semi-DRRBM model with clustering labels according to the embodiment of the present invention (semi-DRRBM-AP algorithm);
[0054] Figure 5 It is the T-SNE visualization of the hidden representation of the semi-DRRBM model with clustering labels according to the embodiment of the present invention (semi-DRRBM-KM algorithm);
[0055] Figure 6The average performance of the semi-DRRBM-KM and semi-DRRBM-AP algorithms based on the semi-DRRBM model in this embodiment of the invention on six high-dimensional graphene atomic structure image datasets (from 15 to 60 labels);
[0056] Figure 7 The convergence analysis of the semi-DRRBM model in this embodiment of the invention. Detailed implementation manners
[0057] To make the objectives, technical solutions and advantages of this application clearer and more understandable, the following provides examples with reference to the accompanying drawings and further elaborates on this application in detail.
[0058] High-dimensional sparse data widely appears in industries and recommendation systems. However, the current research on dimensionality reduction of high-dimensional sparse data is still insufficient. When conducting research on its dimensionality reduction, the ability to learn its hidden features is limited, and often a satisfactory low-dimensional representation of high-dimensional data cannot be obtained, and it often takes a long time.
[0059] As Figure 1 shown, the present invention provides a high-dimensional sparse image characterization method based on a deep semi-supervised learning framework, including the following steps:
[0060] The input of the semi-DRRBM objective function is:
[0061]
[0062] is the supervision information of the visible high-dimensional layer, where is the visible layer vector of the semi-DRRBM model, N represents the number of all instances, and each visible layer vector has a low-dimensional hidden feature , and then the high-dimensional vector is reconstructed using one-step Gibbs sampling during the training process, is 's low-dimensional hidden feature, is the reconstructed high-dimensional vector set;
[0063] 1) Initialize , where the objective function of semi-DRRBM is defined as follows:
[0064] (Equation 12)
[0065] , which is used to update its value according to this algorithm, where W = (w j,i) ∈ RN×N represents the connection weight vector between the hidden layer and the visible layer, b is the visible layer bias vector, a is the hidden layer bias vector, i, j represent an arbitrarily chosen or specified natural number, and when used as subscripts, they indicate positions in the array. ξ ∈ (0, 1) represents the adjustment parameter. , is the energy function of the RBM model. , σ is a logistic sigmoid function , N represents the number of all instances, and D represents the original data set;
[0066] represents the visible layer vector, and each visible layer vector has a low-dimensional hidden feature , and then a high-dimensional vector is reconstructed using one-step Gibbs sampling during the training process. , is the low-dimensional hidden feature of, where the subscript k represents the size of the array contained in the vector. When using other subscripts, the meanings of v, v’, h, h’ are the same. The relationships between the subscripts p, q and the subscripts m, n are as described below:
[0067] Let ∀ and ∀ simultaneously belong to the k-th cluster (k = 1, 2,..., K, K represents the total number of clusters). All these unordered pairs ( , ) form an internal cluster set I, and |I| represents the size of I;
[0068] ∀ and ∀ belong to the i-th and j-th clusters respectively (i ≠ j). All unordered pairs ( , ) form an inner cluster set O, and |O| represents the size of O;
[0069] Meanwhile, ∀ and ∀ have a one-to-one mapping and , ( , ) form a set I’, and |I’| represents the size of I’;
[0070] ∀ and ∀ have a one-to-one mapping and , ( , ) form a set O', and |O'| represents the size of O';
[0071] 2) Obtain the sample hidden layer features through ;
[0072] 3) Use to reconstruct the visible layer data;
[0073] 4) Use to obtain the hidden layer features of the sample reconstruction;
[0074] 5)
[0075] Define the function
[0076] (Formula 13)
[0077] (Formula 14)
[0078] ;
[0079] Because,
[0080] So,
[0081] (Formula 15)
[0082] where the vector v represents the visible layer vector, h represents the low-dimensional hidden feature (also a vector), and in the training process, a high-dimensional vector is reconstructed using one-step Gibbs sampling , is 's low-dimensional hidden feature. (Steps 2 - 4 have been completed).
[0083] At the same time, let ∀ and ∀ both belong to the k-th cluster (k = 1, 2,..., K, K represents the total number of clusters) ∀ and ∀ belong to the i-th and j-th clusters (i ≠ j) respectively, and at the same time ∀ and ∀ have a one-to-one mapping and , ∀ and ∀ have a one-to-one mapping and Each visual layer vector vp, vq, vm, vn, and their mappings vp’, vq’, vm’, vn’ have hidden vectors hp, hq, hm, hn, and their mappings hp’, hq’, hm’, hn’.
[0084] (Formula 16)
[0085] For the partial derivative is:
[0086] (Formula 17)
[0087] where, v mi is the element at the i-th position of vector v m , v ni is the element at the i-th position of vector v n , h mj , h nj are the sum of the elements at the j-th position of vectors h m and h n respectively; v pi is the element at the i-th position of vector v p , v qi is the element at the i-th position of vector v q , h pj , h qj are the elements at the j-th position of vectors h p and h q respectively;
[0088] For the partial derivative:
[0089] (Formula 18)
[0090] v mi ’ is the element at the i-th position of vector v m ’, v ni ’ is the element at the i-th position of vector v n ’, h mj ’, h nj ’ are the sum of the elements at the j-th position of vectors h m ’ and h n ’ respectively; v pi ’ is the element at the i-th position of vector v p ’, v qi ’ is the element at the i-th position of vector v q ’, h pj ’, h qj ’ are the elements at the j-th position of vectors hp ’ and the vector h q ’s element at the j-th position;
[0091] and The partial derivative with respect to is:
[0092] (Equation 19)
[0093] (Equation 20)
[0094] Using the above method to obtain their partial derivatives with respect to :
[0095] The partial derivative with respect to :
[0096] (Equation 21)
[0097] (Equation 22)
[0098] Because is independent of a, so their partial derivatives with respect to ai are 0
[0099] (Equation 23)
[0100] 6) Substitute the previous hidden layer data h, visible layer data v, reconstructed hidden layer data h’, visible layer data v’, and the parameters obtained during the i-th calculation into Equations 14 - 16 for calculation to obtain the of this calculation. The update principle is to update the parameters using the formulas calculated in Steps 1 - 5.
[0101] The update rule of the semi-DRRBM model parameters takes the following form: (Equation 24)
[0102] W ij (τ) The connection weight at the -th iteration, W ij (τ+1) The connection weight at the -th iteration;
[0103] ε represents the learning rate, ξ ∈ (0, 1) represents the adjustment parameter, and the visible layer vector all have a low-dimensional hidden feature ;
[0104] Let ∀ and ∀ simultaneously belong to the k-th cluster (k = 1, 2, …, K, where K represents the total number of clusters). All these unordered pairs ( , ) form an internal cluster set I, and |I| represents the size of I; ∀ and ∀ belong to the i-th and j-th clusters respectively (i ≠ j). All unordered pairs ( , ) form an inter-cluster set O, and |O| represents the size of O; simultaneously, ∀ and ∀ have a one-to-one mapping and , ( , ) form a set I’, and |I’| represents the size of I’; ∀ and ∀ have a one-to-one mapping and , ( , ) form a set O’, and |O’| represents the size of O’;
[0105] (Formula 25)
[0106] b j (τ) The bias of the visible layer units when iterating to the -th calculation, b j (τ+1) The bias of the visible layer units when iterating to the -th calculation.
[0107] (Formula 26)
[0108] Where < >0 and < >1 represent the expected values of the subscript data and the reconstructed data respectively;
[0109] a j (τ) The bias of the hidden layer units when iterating to the -th calculation, a j (τ+1) The bias of the hidden layer units when iterating to the -th calculation;
[0110] 7) When obtaining new parameters through step 6 After that, use it as the new input and repeat steps 2-6 again, iterate until the maximum value, or stop when repeating to the specified number of times, or stop when less than a certain error specified by humans;
[0111] 8) Return: The parameters of the Semi-DRRBM model , substitute it into the Semi-DRRBM model to complete dimensionality reduction.
[0112] Because each vector v in the visible layer and the vector h in the hidden layer have a one-to-one correspondence, the supervision information in the high-dimensional space can be used in the hidden low-dimensional space. In addition, there is also a one-to-one correspondence between the visible layer vector v and the reconstructed visible layer vector v'. Therefore, the present invention fuses the supervision information of the visible layer G and the reconstructed visible layer G' into the hidden low-dimensional space and the reconstructed hidden layer . Thus, the representation ability of the semi-DRRBM of the present invention is enhanced.
[0113] Such as Figure 2 shown, a deep semi-supervised representation learning (deep semi-RL) framework is constructed using six -DRRBM models. The deep semi-RL framework has a high-dimensional sparse visible layer ( ), Figure 2 Each dashed box in represents a semi-DRRBM model. Each shallow model is independently trained. The hidden layer features of the upper shallow model are used as the input of the lower shallow model. In the hidden layer, the dimensions of the high-dimensional sparse visible layer are reduced to 30000, 10000, 3000, 1000, 500, and 250 respectively. During the dimensionality reduction process of semi-RL, some supervision information, for example, is fused into the low-dimensional hidden layer units. The final low-dimensional dense hidden features ( ) are directly used to verify the representation ability of the proposed deep semi-RL framework without fine-tuning. In the following part, it is used as the input for the downstream clustering task.
[0114] The names of the six created high-dimensional datasets are: GNFC-L, GNFC-SBS, GNIC-L, GNIC-SBS, GNFC-CPK, and GNIC-S. To demonstrate the feature learning and generalization capabilities of the shallow semi-DRRBM model of the present invention for high-dimensional and sparse atomic-level structure images, the low-dimensional hidden and dense features are used as the inputs to the K-Means (KM) and Affinity Propagation (AP) algorithms to evaluate the performance. To distinguish the method of the present invention from the traditional KM and AP algorithms, in the following experiments, the proposed methods are respectively referred to as semi-DRRBM-KM and semi-DRRBM AP. The dimension of the visible layer of the semi-DRRBM model is 62,500, and then it is reduced to 625 in the hidden layer. The learning rates of the semi-DRRBM and the traditional RBM model proposed by the present invention are the same, and in the experiment, they are both set to 0.1. The comparative algorithms based on the baseline model are referred to as RBM-KM and RBM-AP in the experiment. In the semi-RL framework, the dimension of the visible layer is 62,500. Then it is gradually reduced to 30,000, 10,000, 3,000, 1,000, 500, and 250 in the subsequent hidden layers (see Figure 2 ). In the comparative experiment, the KM, AP, RBM-KM, and RBM-AP algorithms are repeated ten times, and the results are their averages.
[0115] The semi-DRRBM-KM and semi-DRRBM-AP algorithms based on the layer semi-DRRBM model proposed by the present invention are compared with KM, AP, RBM-KM, RBM-AP, and Semi-EAGR. In addition, the proposed deep semi-reinforcement learning framework is compared with the shallow semi-DRRBM model to show the deep representation learning ability. Three external evaluation metrics, namely the clustering accuracy, Jaccard index (Jac index), and Fowlkes and Mallows index (FM index), are used to demonstrate the performance of the proposed shallow model and deep framework. To explore the influence of different numbers of labels on the performance of the semi-DRRBM model, 15, 30, 45, and 60 labels are respectively used during the training process.
[0116] 1. Performance Evaluation of the Shallow Semi-DRRBM Model
[0117] 1) Clustering: Comparison among KM, AP, RBMKM, RBM-AP, Semi-EAGR, semi-DRRBM-KM and semi-DRRBM-AP algorithms. From the experimental results, the top 2 algorithms for the precision evaluation metric are semi-DRRBM-KM and semi-DRRBM-AP. Their average values are 0.7495 and 0.7693 respectively. In addition, for each dataset, almost all the top 2 algorithms are semi-DRRBM-KM and semi-DRRBM-AP.
[0118] For the comparison results of another evaluation metric (Jac index), the top 2 algorithms for each dataset belong to the semi-DRRBM-KM and semi-DRRBM-AP of the present invention, and their average Jac indices are 0.6002 and 0.6217 respectively.
[0119] For the comparison results of the FM index external evaluation metric, for each dataset, the top 2 algorithms still belong to the semi-DRRBM-KM and semi-DRRBM-AP of the present invention, and their average FM indices are 0.7560 and 0.7770 respectively. The algorithms of the present invention show an absolute advantage over the comparative methods.
[0120] Figure 3 shows the performance trends of the accuracy, Jac metric and FM index metric with the number of labels (from 15 to 60). As the supervised information increases, most of the evaluation metrics indicate that the performance of the semi-DRRBM-KM and semi-DRRBM-AP algorithms of the present invention has been improved.
[0121] 2) Visualization of hidden representations: Figure 3 The TSNE visualization results of the original data with true labels are shown. Obviously, all the instances of each original dataset are entangled together. Their separability is poor. In the hidden layer and the reconstructed hidden layer of the shallow semi-DRRBM model of the present invention, some semi-supervised dimensionality reduction strategies with partial labels can improve the separability of the low-dimensional and dense features in the hidden layer. Figure 4 The T-SNE visualization results of the hidden features of the semi-DRRBM model with clustering labels of the present invention, which is derived from the semi-DRRBMAP algorithm, are shown. Figure 5 The visualization results of another semi-DRRBM-KM algorithm are shown. From all the visualization results, the separability of the hidden layer of semiDRRBM has been improved.
[0122] 3) Convergence analysis: Figure 6 shows the convergence of the proposed semi-DRRBM model. For each dataset, the loss value of the semi-DRRBM model drops rapidly in the first 40 iterations. Then, for all high-dimensional and sparse atomic structure image datasets, it gradually converges in the following iteration steps. Therefore, the semi-supervised dimensionality reduction strategy of the present invention can guarantee the convergence of the proposed shallow semi-DRRBM model.
[0123] 4) Generalization ability: We use the same low-dimensional representation of the proposed semi-DRRBM model as the input for two different clustering algorithms to demonstrate the generalization ability in the experiment. The same representation of the semi-DRRBM model can exhibit high performance for different algorithms, and the performance evaluation shows that the two different algorithms based on the semi-DRRBM model have better performance than the most relevant comparison algorithms.
[0124] 2. Deep representation ability of semi-RL: Performance comparison between the shallow semi-DRRBM model and the deep semi-RL framework with 15 labels. The results of all three evaluation metrics (accuracy, Jax index, and FM index) simultaneously show that the deep semi-reinforcement learning framework has better performance than the shallow semi-DRRBM model. The average performance of the former is improved to 0.7667, 0.6202, and 0.7663 respectively. The same low-dimensional hidden features of the deep semi-RL framework can also be applied to the AP algorithm. The average performance of the deep framework can be further improved to 0.7709, 0.6242, and 0.7809 respectively in the future.
[0125] The present invention proposes a deep semi-RL (semi-supervised representation learning) framework with six hidden layers based on the semi-DRRBM model. Different from the pcGRBM model (pairwise constraints restricted Boltzmann machine with Gaussian visible units (pcGRBM) model), the proposed shallow semi-DRRBM variant model and deep semi-RL framework are applicable to binary data modeling. Most importantly, the constraints are more reasonably applied to the encoding process of the hidden layer rather than the reconstruction process of the visible layer. To prove the rationality of the model, six novel high-dimensional and sparse pixel-level image datasets with 2D / 3D graphene atomic structures are used. Experimental results show that the proposed semi-DRRBM model can converge quickly on all sparse datasets. T-SNE three-dimensional visualization shows that the semi-DRRBM model has a better hidden representation distribution than the original visible layer space. External evaluation metrics all prove that the proposed semi-DRRBM model has better performance than the baseline and comparison models in clustering tasks and also shows a certain generalization ability.
[0126] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A high-dimensional sparse image representation method based on a deep semi-supervised learning framework, characterized in that, It includes the following steps: The input of the semi-DRRBM objective function is: is the supervision information of the visible high-dimensional layer, where is the visible layer vector of the semi-DRRBM model, N represents the number of all instances, and each visible layer vector has a low-dimensional hidden feature , and then a high-dimensional vector is reconstructed using one-step Gibbs sampling during the training process , is 's low-dimensional hidden feature, is the set of reconstructed high-dimensional vectors; 1) Initialization ; The objective function of semi-DRRBM is defined as follows: (Formula 1) , which is used to update its value according to this algorithm, where W = (w j,i ) ∈ RN×N represents the connection weight vector between the hidden layer and the visible layer, b is the visible layer bias vector, a is the hidden layer bias vector, i, j represent an arbitrarily selected or specified natural number, and when used as subscripts, they indicate positions in the array, ξ ∈ (0, 1) represents the adjustment parameter, , is the energy function of the RBM model, , σ is a logistic sigmoid function , N represents the number of all instances, and D represents the original data set; represents the visible layer vector, and each visible layer vector has a low-dimensional hidden feature , and then a high-dimensional vector is reconstructed by one-step Gibbs sampling during the training process , is the low-dimensional hidden feature of, where the subscript k represents the size of the array included in the vector. When using other subscripts, the meanings of v, v’, h, and h’ are the same. The relationships between the subscripts p, q and the subscripts m, n are described as follows: Let ∀ and ∀ both belong to the k-th cluster, where k = 1, 2,..., K and K represents the total number of clusters. All these unordered pairs ( , ) form an intra-cluster set I, and |I| represents the size of I; ∀ and ∀ belong to the \(i\)-th and \(j\)-th clusters respectively (\(i\neq j\)), and all unordered pairs ( , ) form an inner-cluster set \(O\), and \(|O|\) represents the size of \(O\); Meanwhile, ∀ and ∀ has a one-to-one mapping and , ( , ) constitute a set I’, and |I’| represents the size of I’; ∀ and ∀ has a one-to-one mapping and , ( , ) constitute a set O', and |O'| represents the size of O'; 2) By obtaining the sample hidden layer features; 3) Use Reconstruct the visible layer data; 4) Use to obtain the hidden layer features of the sample reconstruction; 5) Calculate partial derivatives; 6) Use the previous hidden layer data h, visible layer data v, reconstructed hidden layer data h', visible layer data v', and the parameters obtained during the i-th calculation to calculate the for this calculation; 7) After obtaining the new parameters through step 6 take them as the new input and perform steps 2 - 6 again, iterate until the maximum value is reached, or repeat until the specified number of times is reached and stop, or stop when it is less than a certain error specified by humans; 8) Return: The parameters of the Semi-DRRBM model , and substituting them into the Semi-DRRBM model can complete dimensionality reduction.
2. The high-dimensional sparse image representation method based on a deep semi-supervised learning framework according to claim 1, characterized in that In step 2), , Define a function (Formula 2) (Formula 3) ; For the partial derivative is: (Formula 4) Among them, v mi is the element at the i-th position of vector v m , v ni is the element at the i-th position of vector v n , h mj and h nj are respectively the sum of the elements at the j-th position of vector h m and vector h n ; v pi is the element at the i-th position of vector v p , v qi is the element at the i-th position of vector v q , h pj and h qj are respectively the elements at the j-th position of vector h p and vector h q . For partial derivative with respect to: (Formula 5) v mi ’ is the element at the i-th position of the vector v m ’; v ni ’ is the element at the i-th position of the vector v n ’; h mj ’ and h nj ’ are respectively the sum of the elements at the j-th position of the vector h m ’ and the vector h n ’; v pi ’ is the element at the i-th position of the vector v p ’; v qi ’ is the element at the i-th position of the vector v q ’; h pj ’ and h qj ’ are respectively the elements at the j-th position of the vector h p ’ and the vector h q ’; The partial derivative of with respect to: (Formula 6) (Formula 7) Because is independent of a, so their partial derivative with respect to a j is 0; (Formula 8).
3. A high-dimensional sparse image representation method based on a deep semi-supervised learning framework according to claim 1, characterized in that Step 6) Use the previous hidden layer data h, visible layer data v, reconstructed hidden layer data h', visible layer data v', and the parameters obtained during the i-th calculation , and calculate the for this calculation; where the update rule for the semi-DRRBM model parameters takes the following form: (Formula 9) W ij (τ) The connection weight at the -th calculation iteration, W ij (τ+1) The connection weight at the -th calculation iteration; ε represents the learning rate, ξ ∈ (0, 1) represents the adjustment parameter, and the visible layer vectors all have a low-dimensional hidden feature ; (Formula 10) b j (τ) Bias of the visible layer units at the -th calculation, b j (τ+1) Bias of the visible layer units at the -th calculation (Formula 11) where <0> and <1> respectively represent the expected values of the subscript data and the reconstructed data; a j (τ) The bias of the hidden layer unit when iterating to the -th calculation, a j (τ+1) The bias of the hidden layer unit when iterating to the -th calculation.