A Cross-Media Data Dimensionality Reduction Method Based on Learning and Reasoning
Through the cross-media data dimensionality reduction method based on learning and reasoning, the problem of difficulty in dealing with time-series features and multimedia data in the prior art is solved, and flexible data dimensionality reduction and information retention are achieved, which is suitable for feature extraction and visualization of multimedia data.
Patent Information
- Application Number
- CN202211675814.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-12-26
AI Technical Summary
Existing data dimensionality reduction methods are difficult to effectively process data containing timing characteristics and multimedia data, especially to effectively combine data between media to obtain useful information.
A cross-media data dimensionality reduction method based on learning and reasoning is adopted. By classifying multimedia data by media, using an encoder to extract feature vectors, splicing and dimensionality reduction are performed again, and finally using a decoder to reconstruct the original data, retaining timing information and removing noise.
It realizes flexible data dimensionality reduction, retains timing characteristics and key information, and effectively integrates multimedia data to facilitate the calculation and visualization of downstream tasks.
Smart Images

Figure CN116028807B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and data dimensionality reduction, and more specifically, to a cross-media data dimensionality reduction method based on learning and reasoning. Background Art
[0002] In recent years, with the increasing convenience of data acquisition, it has become easier for people to obtain high-dimensional data. However, a large amount of high-dimensional data needs to be effectively processed before it can be used in practice. High-dimensional data has the characteristics of large data volume, difficult to calculate, information redundancy, containing noise information, and being unintuitive. In order to effectively extract the information contained in these data and discard redundant and noise information, an effective dimensionality reduction method is required.
[0003] Most of the existing data dimensionality reduction methods utilize the inherent characteristics of data to provide a general means. For example, data is regarded as a matrix and matrix decomposition (i.e., linear transformation) is performed. Another example is to start from the variance that can represent the irrelevance between data in multi-dimensional data. The greater the correlation, the less information the data gives. Therefore, it is still hoped to retain a larger variance after dimensionality reduction. These methods are widely used and effective, but for a certain type of special data, such as data containing time series information, they lack pertinence and thus it is difficult to achieve the best results. In particular, for multi-media data, traditional dimensionality reduction methods are difficult to combine the data between media to obtain useful information.
[0004] At the same time, the dimensionality reduction method based on neural network has a very high applicability to this type of data because a special type of data has been learned through a specially designed network. For example, a dimensionality reduction method for high-dimensional measurement data of dynamic systems based on deep learning is proposed in the prior art. First, data is collected, and a neural network and an objective function of a deep autoencoder are constructed for the obtained data. Then, the data is input into the neural network for training. Finally, the decoder is discarded and only the encoder is used for dimensionality reduction. This dimensionality reduction algorithm based on deep learning can handle both linear data and non-linear data. The running speed of the model is relatively fast under the premise of being trained, and it has an explicit dimensionality reduction function. The deep autoencoder achieves better results in classification tasks after dimensionality reduction of data than linear dimensionality reduction, but it does not have pertinence to the distribution of data. Especially for some specific large multi-media data with time series characteristics, effective dimensionality reduction cannot be performed. In addition, this solution first obtains data and then constructs a neural network according to the data, which is not flexible enough and slow. Summary of the Invention
[0005] To solve the problem that the current data dimensionality reduction methods cannot effectively process data with time series characteristics and multi-media data dimensionality reduction, the present invention proposes a cross-media data dimensionality reduction method based on learning and reasoning, which is flexible in dimensionality reduction, and the data after dimensionality reduction retains the characteristics of time series information and has good effects.
[0006] To achieve the above technical effects, the technical solution of the present invention is as follows:
[0007] A cross-media data dimensionality reduction method based on learning and reasoning, characterized in that the method comprises the following steps:
[0008] S1. Obtain a multi-media data set and classify the multi-media data set by medium;
[0009] S2. Input each medium data into the encoder corresponding to the medium data to extract the feature vector of the medium data, where the feature vector is a vector obtained by multiplying the time dimension by the feature vector dimension;
[0010] S3. Concatenate the feature vectors of several media output by all encoders into a vector;
[0011] S4. Perform dimensionality reduction on the concatenated vector again to obtain an intermediate vector with reduced dimensions;
[0012] S5. Use the decoder to dimensionally up-sample the intermediate vector into several matrices with the same format as the original medium data to reconstruct the original multi-media data;
[0013] S6. Construct a loss function, substitute the up-sampled matrix data and the original multi-media data into the loss function to calculate the loss, update the encoder-decoder parameter data, and repeat steps S1 to S5 until the loss function converges;
[0014] S7. After the loss function converges, discard the decoder in S5 and retain the network components mentioned in S2 to S4.
[0015] Preferably, in step S1, at regular intervals, use several media to sample several attributes of the research subject data, and straighten the sampling result of a certain attribute at a certain moment into a vector. After classifying the multi-media data set by medium, the format of each medium data is a matrix with the same time dimension for convenient training.
[0016] Preferably, in step S2, the number of encoders is the same as the number of media, and the structure of each encoder is the same.
[0017] Preferably, the encoder uses a long short-term memory recurrent neural network LSTM or a gated recurrent unit GRU.
[0018] Preferably, the encoder uses a recurrent neural network. When extracting the feature vector of the medium data, in chronological order, the feature vector at a certain moment is obtained in sequence, where each feature vector is determined by the current input value and the value of the hidden layer retained from the previous moment; the process satisfies:
[0019] S t= f(U·x y + W·S t-1 )
[0020] y t = g(V·S t )
[0021] where x t is the input vector of media data at the current moment, and S t and S t-1 represent the values of the hidden layer retained at the current moment and the previous moment respectively; U and W represent two different weight matrices, f and g are the ReLU activation function and the softmax activation function respectively, and y t is the feature vector of the media data output at the current moment. The dimension of y t is much smaller than that of x t ; Connect y1, y2, …, y t head to tail in chronological order to form a new vector y, whose dimension is the sum of the dimensions of all y t vectors.
[0022] Here, using the time series analysis ability of the recurrent neural network, a series of matrix linear transformation operations and activation operations are performed, so that the obtained result retains the features including time series information, and the obtained vector is not directly obtained, but the feature vector at a certain moment is obtained in sequence according to the time order.
[0023] Preferably, in step S3, let the feature vector output by a certain medium be q m , m refers to the m-th medium. Concatenate the feature vectors q1, q2, …, q m , …, q M of several media output by all encoders head to tail in chronological order to form a vector, and the dimension of this vector is the sum of the dimensions of the feature vectors q m of the output media.
[0024] Preferably, in step S4, use the deep neural network DNN to reduce the dimension of the concatenated vector again. The deep neural network DNN consists of an input layer, two hidden layers and an output layer. The number of neurons in the latter layer is 0.25 times that of the previous layer. The neurons in the previous layer and the neurons in the latter layer are in a fully connected form, that is, each neuron in the previous layer is connected to each neuron in the latter layer; Let the outputs of all neurons in the previous layer be a1, a2, …, a n , then the output of a certain neuron in the latter layer is obtained as follows:
[0025] S41. Obtain the weighted linear sum through the output of the previous layer and the weight parameters of this neuron:
[0026]
[0027] S42. Use the ReLU function to activate and get the output:
[0028] z=ReLU(h)
[0029] Among them, w i is the weight parameter of the i-th neuron in the previous layer, and z is the output of this neuron.
[0030] Preferably, in step S4, the feature vectors of the output media before splicing are adjusted so that the dimensions of the feature vectors of the output media are the same, the feature vectors of the media with the same dimension are spliced into a matrix, and then the convolutional neural network CNN is used for dimensionality reduction.
[0031] Here, a deep neural network is used, and the forward and backward propagation ideas are utilized, combined with the offset vector and the ReLU activation function to perform a series of linear calculations and activation operations, so that the data can be effectively reduced in dimension again.
[0032] Preferably, in step S5, the process of using a decoder to upgrade the intermediate vector into a plurality of matrices having the same format as the original media data to reconstruct the original multimedia data is as follows:
[0033] S51. First, a deep neural network DNN is used. The structure of the deep neural network DNN is the same as that of the deep neural network DNN used in step S4, which is composed of an input layer, two hidden layers and an output layer. The number of neurons in the latter layer is 4 times the number of neurons in the previous layer, and a vector is obtained, and the dimension is the same as the vector concatenated in S3;
[0034] S52. Input the vectors into multiple recurrent neural networks respectively. The number of recurrent neural networks is the same as the number of media, and corresponds to each recurrent neural network in S2 with the same structure. For a recurrent neural network, concatenate the vectors output at different times into one vector, which corresponds to the original data of this medium in S2 with the same format.
[0035] Preferably, in step S6, the loss function is mean square error or cross entropy; when the loss function is mean square error, the expression is:
[0036]
[0037] Among them, M is the number of media, T is the number of sampling moments, and n i is the vector dimension of the i-th media data, x i,t represents the output vector of the decoder of the i-th media data at time t, is the original vector of the i-th media data at time t.
[0038] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0039] The present invention proposes a cross-media data dimensionality reduction method based on learning and reasoning. Multimodal data is collected and classified according to the type of medium. Then, each piece of medium data is input into the encoder corresponding to the medium data to extract the feature vector of the medium data. The feature vector is a vector obtained by multiplying the time dimension by the feature vector dimension, realizing the extraction of temporal features, so that the dimensionality-reduced vector retains the temporal features. By using dimensionality reduction after vector concatenation, effective fusion of feature vectors between media is achieved. Finally, the decoder is used to up-dimension the intermediate vector into several matrices with the same format as the original medium data to reconstruct the original multimodal data. It can well process data with temporal information and multiple media, effectively remove noise, retain key features and information, and facilitate the calculation and visualization of downstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic flowchart showing the cross-media data dimensionality reduction method based on learning and reasoning proposed in Embodiment 1 of the present invention;
[0041] Figure 2 It is a schematic diagram showing the extraction of the feature vector of medium data in chronological order using a recurrent neural network proposed in Embodiment 3 of the present invention;
[0042] Figure 3 It is a schematic diagram showing the structure of the deep neural network proposed in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0044] For better illustration of this embodiment, some parts of the drawings are omitted, enlarged or reduced, and do not represent the actual size;
[0045] For those skilled in the art, it is understandable that some well-known content descriptions in the drawings may be omitted.
[0046] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0047] The description of the positional relationship in the drawings is only for illustrative purposes and should not be construed as a limitation of this patent;
[0048] Embodiment 1
[0049] As Figure 1 shown, this embodiment proposes a cross-media data dimensionality reduction method based on learning and reasoning. Referring to Figure 1 , the method includes the following steps:
[0050] S1. Obtain a multi-media data set and classify the multi-media data set by media type;
[0051] In step S1, at regular intervals, use several media to sample several attributes of the research subject data. For the convenience of training, straighten the sampling result of a certain attribute at a certain moment into a vector. After classifying the multi-media data set by media type, the format of each media data is a matrix with the same time dimension.
[0052] S2. Input each media data into the encoder corresponding to the media data to extract the feature vector of the media data. The feature vector is a vector obtained by multiplying the time dimension by the feature vector dimension; in step S2, the number of encoders is the same as the number of media, and the structure of each encoder is the same.
[0053] S3. Concatenate the feature vectors of several media output by all encoders into a vector;
[0054] S4. Perform dimensionality reduction on the concatenated vector again to obtain an intermediate vector with reduced dimensions;
[0055] S5. Use the decoder to increase the dimension of the intermediate vector into several matrices with the same format as the original media data to reconstruct the original multi-media data;
[0056] S6. Construct a loss function, substitute the matrix data after dimensionality increase and the original multi-media data into the loss function to calculate the loss, update the encoder-decoder parameter data, and repeat steps S1 to S5 until the loss function converges;
[0057] S7. After the loss function converges, discard the decoder in S5 and retain the network components mentioned in S2 to S4.
[0058] Embodiment 2
[0059] In this embodiment, the encoder uses a long short-term memory recurrent neural network LSTM or a gated recurrent unit GRU.
[0060] Embodiment 3
[0061] In this embodiment, the encoder uses a recurrent neural network. The recurrent neural network extracts information from the previous moment node to obtain the information of the previous moment, thereby extracting the neural network of temporal features. Using the time series analysis ability of the recurrent neural network, a series of matrix linear transformation operations and activation operations are performed, so that the obtained result retains the features including temporal information. When extracting the feature vector of the media data, in chronological order, the feature vector of a certain moment is obtained in sequence, where each feature vector is determined by the current input value and the value of the hidden layer retained from the previous moment; asFigure 2 As shown, the process satisfies:
[0062] S t = f(U·x t + W·S t-1 )
[0063] y t = g(V·S t )
[0064] where x t is the media data input vector at the current moment, and S t and S t-1 represent the values of the hidden layers retained at the current moment and the previous moment respectively; U and W represent two different weight matrices, f and g are the ReLU activation function and the softmax activation function respectively, and y t is the feature vector of the media data output at the current moment, and the dimension of y t is much smaller than that of x t ; Connect y1, y2,..., y t head to tail in chronological order to form a new vector y, whose dimension is the sum of the dimensions of all y t vectors.
[0065] Example 4
[0066] In step S3, let the feature vector output by a certain medium be q m , where m refers to the m-th medium. Concatenate the feature vectors q1, q2,..., q m ,..., q M of several media output by all encoders head to tail in chronological order to form a vector, and the dimension of this vector is the sum of the dimensions of the output feature vectors q m of several media.
[0067] In step S4, use the deep neural network DNN to reduce the dimension of the concatenated vector again. The basic structure of the deep neural network DNN is as Figure 3 shown. Use the deep neural network, and use the forward and backward propagation ideas, combined with the offset vector and the ReLU activation function to perform a series of linear calculations and activation operations, so that the data is effectively reduced in dimension again.
[0068] In this example, the deep neural network DNN is composed of one input layer, two hidden layers and one output layer. The number of neurons in the latter layer is 0.25 times that of the previous layer. The neurons in the previous layer and the neurons in the latter layer are in a fully connected form, that is, each neuron in the previous layer is connected to each neuron in the latter layer; Let the outputs of all neurons in the previous layer be a1, a2,..., a n, the output of a certain neuron in the subsequent layer is obtained as follows:
[0069] S41. Obtain the weighted linear sum through the output of the previous layer and the weight parameters of this neuron:
[0070]
[0071] S42. Activate using the ReLU function (a commonly used activation function in neural networks. In the case where the result is negative, the neuron is not activated, making the network sparser) to obtain the output:
[0072] z = ReLU(h)
[0073] where, w i is the weight parameter of the i-th neuron in the previous layer, and z is the output of this neuron.
[0074] In step S5, the process of using the decoder to elevate the intermediate vector to several matrices with the same format as the original media data to reconstruct the original multi-media data is as follows:
[0075] S51. First, use the deep neural network DNN. The structure of the deep neural network DNN is the same as the deep neural network DNN used in step S4, consisting of one input layer, two hidden layers, and one output layer. The number of neurons in the subsequent layer is 4 times that of the previous layer, obtaining a vector with the same dimension as the vector concatenated in S3;
[0076] S52. Input the vector into multiple recurrent neural networks respectively. The number of recurrent neural networks is the same as the number of media, and corresponding to each recurrent neural network in S2, the structure is the same; for a recurrent neural network, concatenate the vectors output at different times into a vector, and this vector corresponds to the original data of this media in S2 with the same format.
[0077] In step S6, the loss function is the mean squared error or cross-entropy; when the loss function is the mean squared error, the expression is:
[0078]
[0079] where, M is the number of media, T is the number of sampling times, n i is the vector dimension of the i-th media data, x i,t represents the output vector of the decoder of the i-th media data at time t, is the original vector of the i-th media data at time t.
[0080] Example 5
[0081] In step S4, the dimensionality reduction method can also be: adjusting the feature vectors of several media in the output before splicing to make the dimensionalities of the feature vectors of the several media in the output the same, splicing the feature vectors of the several media with the same dimensionality into a matrix, and then using a convolutional neural network (CNN) for dimensionality reduction.
[0082] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A cross-media data dimensionality reduction method based on learning and reasoning, characterized in that, The method comprises the following steps: S1. Obtain a multimedia data set and classify the multimedia data set by media; S2. Input each media data into the encoder corresponding to the media data, and extract the feature vector of the media data, where the feature vector is a vector multiplied by the time dimension and the feature vector dimension; S3. concatenate the feature vectors of several media output by all encoders into one vector; S4. Reduce the dimension of the concatenated vectors again to obtain an intermediate vector with reduced dimension; S5. Using the decoder to upgrade the intermediate vector into several matrices with the same format as the original media data to reconstruct the original multimedia data; S6. Construct a loss function, substitute the dimension-upgraded matrix data and the original multimedia data into the loss function to calculate the loss, update the encoder-decoder parameter data, and repeat steps S1 to S5 until the loss function converges; S7. After the loss function converges, the decoder in S5 is discarded and the network components mentioned in S2 to S4 are retained.
2. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 1, characterized in that In step S1, at regular intervals, several attributes of the research subject data are sampled using several media, and the result of sampling a certain attribute at a certain moment is straightened into a vector. After the multimedia data set is classified by media, the format of each media data is a matrix with the same time dimension.
3. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 1, wherein In step S2, the number of encoders is the same as the number of media, and the structure of each encoder is the same.
4. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 3, wherein The encoder uses a long short-term memory recurrent neural network LSTM or a gated recurrent unit GRU.
5. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 3, characterized in that The encoder uses a recurrent neural network to extract the feature vector of the media data, and obtains the feature vectors at a certain moment in time order, where each feature vector is determined by the current input value and the value of the hidden layer retained at the previous moment; the process satisfies: S t = f(U·x t + W·S t-1 ) y t = g(V·S t ) Among them, x t is the media data input vector at the current moment, S t and S t-1 respectively represent the values of the hidden layer retained at the current moment and the previous moment; U and W represent two different weight matrices, f and g are the ReLU activation function and the softmax activation function respectively, y t is the feature vector of the media data output at the current moment, and the dimension of y t is much smaller than that of x t ; Connect y1, y2, ……, y t head to tail in chronological order to form a new vector y, and the dimension is the sum of the dimensions of all y t vectors.
6. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 1, characterized in that In step S3, let the feature vector output by a certain medium be q m , where m refers to the m-th medium. The feature vectors q1, q2,..., q m ,..., q M of several media output by all encoders are concatenated end to end in chronological order to form a vector, and the dimension of this vector is the sum of the dimensions of the feature vectors q m of the several media output.
7. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 6, characterized in that In step S4, the deep neural network DNN is used to reduce the dimension of the spliced vector again. The deep neural network DNN consists of an input layer, two hidden layers, and an output layer. The number of neurons in the latter layer is 0.25 times that of the previous layer. The neurons in the previous layer and the neurons in the latter layer are in a fully connected form, that is, each neuron in the previous layer is connected to each neuron in the latter layer. Let the outputs of all neurons in the previous layer be a1, a2, … n , then the output of a certain neuron in the latter layer is obtained as follows: S41. The weighted linear sum is obtained by the output of the previous layer and the weight parameter of this neuron: S42. Use the ReLU function to activate and get the output: z=ReLU(h) where, w i is the weight parameter of the i-th neuron in the previous layer, and z is the output of this neuron.
8. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 6, wherein In step S4, the feature vectors of the output media before splicing are adjusted so that the dimensions of the feature vectors of the output media are the same, the feature vectors of the media with the same dimension are spliced into a matrix, and then the convolutional neural network CNN is used for dimensionality reduction.
9. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 7, wherein In step S5, the process of using the decoder to upgrade the intermediate vector into a number of matrices with the same format as the original media data to reconstruct the original multimedia data is as follows: S51. First, a deep neural network DNN is used. The structure of the deep neural network DNN is the same as that of the deep neural network DNN used in step S4, which is composed of an input layer, two hidden layers and an output layer. The number of neurons in the latter layer is 4 times the number of neurons in the previous layer, and a vector is obtained, and the dimension is the same as the vector concatenated in S3; S52. Input the vectors into multiple recurrent neural networks respectively. The number of recurrent neural networks is the same as the number of media, and for each recurrent neural network corresponding to S2, the structure is the same. For a recurrent neural network, concatenate the vectors output at different times into a vector, which corresponds to the original data of this medium in S2 and has the same format.
10. The cross-media data dimensionality reduction method based on learning and reasoning according to claim 9, wherein In step S6, the loss function is mean squared error or cross entropy. When the loss function is mean squared error, the expression is: Among them, M is the number of media, T is the number of sampling moments, and n i is the vector dimension of the i-th media data, and x i,t represents the output vector of the decoder for the i-th media data at time t, is the original vector of the i-th media data at time t.
Citation Information
Patent Citations
Image description generating method based on neural network and image attention focuses
CN106777125A
Machine Learning Model-Based Video Compression
US20220329876A1