A Multi-Dimensional Feature Selection Method Based on Deep Learning

Through a multi-dimensional feature selection method based on deep learning, the Inception network and feature dimension selection module are used to dynamically select feature dimensions to improve image retrieval efficiency, solving the problem of low retrieval efficiency in the prior art.

CN113987232BActive Publication Date: 2025-06-13SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111198581.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-14
Publication Date
2025-06-13
Estimated Expiration
2041-10-14

AI Technical Summary

Technical Problem

The prior art is inefficient in processing huge image data on the Internet and cannot meet the needs of fast retrieval.

Method used

A multi-dimensional feature selection method based on deep learning is adopted to extract common features through the Inception network, and a feature dimension selection module is designed, including actor network, critic network and reward function, and feature dimensions are dynamically selected to reduce search time.

Benefits of technology

By dynamically selecting feature dimensions, the query time is significantly reduced and the efficiency of image retrieval is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113987232B_ABST
    Figure CN113987232B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-dimensional feature selection method based on deep learning. This method learns general features of images through an Inception deep learning model and a feature dimension random truncation model, stores the general features of database images, and selects the feature dimensions of a query image through a carefully designed feature dimension selection model, and queries in the database based on this dimension, so as to reduce the time required for querying.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technologies and computer vision, and more specifically, to a multi-dimensional feature selection method based on deep learning. Background Art

[0002] In recent years, retrieval methods based on deep networks have made remarkable progress. More extensive research efforts have been devoted to learning accurate image retrieval models. However, for the huge amount of image data on the Internet, mere accuracy cannot meet the actual needs. Therefore, academic researchers have shown great attraction to faster image retrieval technologies.

[0003] Most of the current retrieval technologies and existing metric learning methods convert all input samples into fixed-length feature vectors. These existing methods ignore those simple examples that can be represented by shorter feature dimensions, so the retrieval efficiency is relatively low.

[0004] For the above problems, it is natural to think of reducing the search time by dynamically selecting the dimension of features. In order to select features, first, a general feature extraction model is needed, specifically the Inception network. After testing, it is found that the general features extracted by the Inception network have little loss in accuracy compared with the features trained separately. Then, a feature dimension selection module is designed, which specifically includes an Actor Network, a Critic Network, and a Reward Function. Summary of the Invention

[0005] The present invention provides a multi-dimensional feature selection method based on deep learning, which reduces the time required for queries.

[0006] In order to achieve the above technical effects, the technical solution of the present invention is as follows:

[0007] A multi-dimensional feature selection method based on deep learning, comprising the following steps:

[0008] S1: Establish a deep learning network model G for extracting general features of images;

[0009] S2: Add a feature dimension random truncation model after the network model G;

[0010] S3: Obtain the general features of the training set and the test set by training on the training set;

[0011] S4: After obtaining the general features of the images, establish a feature dimension selection model;

[0012] S5: Train and test the feature dimension selection model;

[0013] S6: Establish a process for providing a background interface, provide a retrieval entry, and return retrieval results.

[0014] Further, the specific process of step S1 is as follows:

[0015] S11: Establish a feature extraction layer for the G network, represent each frame of the preprocessed pictures in each video as a low-dimensional real vector, and import the pre-trained model on a large-scale labeled photo into the Inception network;

[0016] S12: Extract a set of feature vectors X of a set length for the images by training the Inception network.

[0017] Further, the specific design of the dimension truncation module in step S2 is as follows:

[0018] S21: Map the feature vector X of a set length into a K-dimensional real vector using a fully connected layer, where K is the maximum allowable feature dimension size

[0019] S22: After encoding each vector into a real vector in S21, establish a dimension truncation module for the G network. Through this module, randomly select a dimension from the minimum dimension (set to 16) to the maximum dimension (set to 128), and sequentially truncate to obtain a feature of a random length. Use the same feature dimension in a small batch, and train the network with the features of these different dimensions each time to obtain a general feature with a maximum length of K. Whenever a feature of a random length is needed, only need to sequentially truncate this general feature.

[0020] Further, the specific process of step S3 is as follows:

[0021] S31: Divide the data set into training data and test data;

[0022] S32: The overall model needs to be trained. The training steps of the G network are as follows: Each small batch of image samples is extracted by the G network to obtain image features with a length of the maximum dimension K. After passing through the feature dimension random truncation model, randomly extract an integer from the minimum dimension to the maximum dimension K as the dimension, and then sequentially truncate the feature of the maximum dimension to obtain a feature matrix of this dimension. Use the minimization of the loss function to train the G network model and train the parameters of the G network;

[0023] S33: The testing steps of the model are as follows: Train the model in a single dimension to obtain feature extraction models for several fixed dimensions. Then, perform the following operations on each fixed-dimension model: First, pass through the training dataset. Input the test data into the G network, and then the G network generates features and stores the features in the database. Then, pass through the test dataset. Use the test dataset as the query set, and calculate the R@K by calculating the distance between the features of each image and the data in the database. The specific calculation method is as follows: Calculate the distances between all image features, then sort them in ascending order of distance. Next, determine whether they belong to the same class of videos. If there are images of the same class among the top K images, it is 1; otherwise, it is 0. Take the average of all the results in the test set to obtain the final result R@K.

[0024] For the general model, the first k dimensions are intercepted and compared with the corresponding fixed-dimension network.

[0025] Further, the specific process of step S4 is as follows:

[0026] S41: Build an Actor network composed of three fully connected layers. The function of this network is to take the general features of an image as the state input and output the predicted appropriate dimension as the action output.

[0027] S42: Build a Critic network composed of several fully connected layers. The function of this network is to take the general features of an image as the state and the action output by the Actor network as the input, and output the score for the Actor network to optimize the Actor network.

[0028] S43: Build a Reward function that returns a score for the dimension output by the Actor network, combined with the length penalty of the dimension, determined by the output of the Actor network and the accuracy penalty of the actual evaluation criterion (R@K) as the supervision information for the Critic network.

[0029] Further, the specific process of step S5 is as follows:

[0030] S51: Divide the dataset into training data and test data.

[0031] S52: The overall model needs to be trained. The training steps for the feature dimension selection model are as follows: The general features of the image are extracted by the G network, and the Actor network and the Critic network are updated alternately, and the slow update method is used. In the first step, the Critic network is fixed, and the dimension selected by the Actor network is obtained through the Actor network. The Actor network is optimized using the score obtained by the Critic network. In the second step, the Actor network is fixed, and the score output by the Critic network for the Actor network is compared with the score of the Reward function to supervise the training of the Critic network. The learning rates and frequencies of the updates of the two networks are different;

[0032] S53: During testing, the dimension d selected is obtained using the Actor network, and the distance is compared with the first d dimensions of the general features of the training set in the database to obtain a ranking. This ranking is evaluated using R@1, R@2, R@4, etc. The specific calculation method is: Calculate the distances between all image features, then sort them from smallest to largest, and then determine whether they belong to the same type of video. If there are images of the same type among the first K images, it is 1, otherwise it is 0. Take the average of all the results in the test set to obtain the final result R@K.

[0033] Further, the specific process of step S6 is as follows:

[0034] S61: Save the trained Inception model and the feature dimension selection model;

[0035] S62: Create a background service process and reserve an interface for image input;

[0036] S63: Through accessing the interface created in S62, the image is input. After that, the background service process in S62 will first preprocess the image into the input format required by the Inception model in S61. Next, the Inception model saved in S61 is called, and the processed image is input into the model to obtain the generality of the image. Then, through the feature dimension selection model in S61, the appropriate dimension size d is obtained, the feature is sequentially intercepted, and the distance is calculated with the first d dimensions of the image general feature data stored in the database, and after sorting from small to large, the first k images are returned. The first k images are the retrieval results of the k most similar images.

[0037] Further, in step S12, the feature extraction process is as follows: First, the Inception model is pre-trained using the imagenet image dataset, and then fine-tuned. After each image passes through the pre-trained Inception model, a set of feature vectors of length k is generated, where k refers to the maximum feature length of the image.

[0038] Further, in step S53, the Reward function combines the penalties of length and precision, such that the selected length is as short as possible with little precision loss. The evaluation criterion for precision loss is R@K, while the length loss is determined by the length output by the Actor network. The specific implementation of the Reward function is as follows:

[0039] Reward = Rc × Ra = recalli / recallall × c × (2 - c).

[0040] Where c equals 1 - di / dall, representing the penalty of length, recalli represents the R@K of the selected length, recallall represents the R@K of the ancestor length, and the ratio of the two represents the precision loss. During the training process, SGD is used for optimization.

[0041] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:

[0042] The present invention can learn the general features of images through the Inception deep learning model and the feature dimension random truncation model, store the general features of the database images, and select the feature dimensions of the query images through the carefully designed feature dimension selection model, and query in the database with this dimension, so as to reduce the query time. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a complete graph of the algorithm model of the present invention;

[0044] Figure 2 It is a schematic diagram of the feature selection module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0046] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;

[0047] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0048] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0049] As Figure 1-2 shown, a multi-dimensional feature selection method based on deep learning includes the following steps:

[0050] S1: Establish a deep learning network model G for extracting general features of images;

[0051] S2: Add a feature dimension random truncation model after the network model G;

[0052] S3: Obtain the general features of the training set and the test set by training on the training set;

[0053] S4: After obtaining the general image features, establish a feature dimension selection model;

[0054] S5: Train and test the feature dimension selection model;

[0055] S6: Establish a process for providing a background interface, provide a retrieval entry and return the retrieval results.

[0056] The specific process of step S1 is as follows:

[0057] S11: Establish the feature extraction layer of the G network, represent each frame of the picture in each preprocessed video as a low-dimensional real vector, and import the model pre-trained on the large-scale labeled photos into the Inception network;

[0058] S12: Extract a set of feature vectors X of a set length for the images by training this Inception network.

[0059] The specific design of the dimension truncation module in step S2 is as follows:

[0060] S21: Map the feature vector X of the set length into a K-dimensional real vector using a fully connected layer, where K is the maximum allowable feature dimension size

[0061] S22: After encoding each vector into a real vector in S21, establish the dimension truncation module of the G network. Through this module, randomly select a dimension from the minimum dimension (set to 16) to the maximum dimension (set to 128), and sequentially truncate to obtain a feature of a random length. Use the same feature dimension in a small batch, and train the network with the features of these different dimensions each time to obtain a general feature with a maximum length of K. Whenever a feature of a random length is needed, only need to sequentially truncate this general feature.

[0062] The specific process of step S3 is as follows:

[0063] S31: Divide the data set into training data and test data;

[0064] S32: The overall model needs to be trained. The training steps of the G network are as follows: For each mini-batch of image samples, the G network extracts image features with a length of the maximum dimension K. After passing through the feature dimension random truncation model, a random integer is drawn from the minimum dimension to the maximum dimension K as the dimension, and then the features of the maximum dimension are sequentially truncated to obtain a feature matrix of this dimension. The G network model is trained using the minimization of the loss function to train the parameters of the G network.

[0065] S33: The testing steps of the model are as follows: The model is trained separately for a certain dimension to obtain several feature extraction models with fixed dimensions. Then, the following operations are performed on each fixed dimension model respectively: First, pass through the training dataset, input the test data into the G network, and then the G network generates features, which are stored in the database. Then, pass through the test dataset. Using the test dataset as the query set, the distance between the features of each image and the data in the database is calculated to calculate R@K. The specific calculation method is as follows: Calculate the distances between all image features, then sort them in ascending order of distance, and then determine whether they belong to the same type of video. If there are the same type among the first K images, it is 1, otherwise it is 0. The average of all the results in the test set is taken to obtain the final result R@K.

[0066] For the general model, the first k dimensions are intercepted and compared with the corresponding fixed-dimension network.

[0067] The specific process of step S4 is as follows:

[0068] S41: An Actor network composed of three fully connected layers is established. The function of this network is to take the general features of an image as the state input and output the predicted appropriate dimension as the action output.

[0069] S42: A Critic network composed of several fully connected layers is established. The function of this network is to take the general features of an image as the state and the action output by the Actor network as the input, and output the score for the Actor network to optimize the Actor network.

[0070] S43: A Reward function is established. This function returns a score for the dimension output by the Actor network, combined with the length penalty of the dimension, determined by the output of the Actor network and the accuracy penalty of the actual evaluation criterion (R@K) as the supervision information for the Critic network.

[0071] The specific process of step S5 is as follows:

[0072] S51: The dataset is divided into training data and test data.

[0073] S52: The overall model needs to be trained. The training steps for the feature dimension selection model are as follows: The G network extracts the general image features, and the Actor network and the Critic network are updated alternately, and the slow update method is used. In the first step, the Critic network is fixed, and the dimension selected by the Actor network is obtained through the Actor network, and the Critic network is used to obtain the score to optimize the Actor network. In the second step, the Actor network is fixed, and the score output by the Critic network for the Actor network is compared with the score of the Reward function to supervise the training of the Critic network. The learning rates and frequencies of the updates of the two networks are different;

[0074] S53: During testing, the Actor network is used to obtain the selected dimension d, and the distance is compared with the first d dimensions of the general features of the training set in the database to obtain a ranking. This ranking is evaluated using R@1, R@2, R@4, etc. The specific calculation method is: Calculate the distances between all image features, then sort them from smallest to largest, and then determine whether they belong to the same type of video. If there are the same type among the first K images, it is 1, otherwise it is 0. Take the average of all the results in the test set to obtain the final result R@K.

[0075] The specific process of step S6 is as follows:

[0076] S61: Save the trained Inception model and the feature dimension selection model;

[0077] S62: Create a background service process and reserve an interface for image input;

[0078] S63: By accessing the interface created in S62, the image is input. Then the background service process of S62 will first preprocess the image into the input format required by the Inception model in S61. Next, the Inception model saved in S61 is called, and the processed image is input into the model to obtain the generality of the image. Then, through the feature dimension selection model in S61, the appropriate dimension size d is obtained, the feature is sequentially intercepted, and the distance is calculated with the first d dimensions of the general image feature data stored in the database, and the first k images are returned after sorting from small to large. The first k images are the retrieval results of the k most similar images.

[0079] In step S12, the feature extraction process is as follows: First, the Inception model is pre-trained with the imagenet image dataset and then fine-tuned. After each image passes through the pre-trained Inception model, a set of feature vectors of length k is generated, where k refers to the maximum feature length of the image.

[0080] In step S53, the Reward function combines the penalties of length and accuracy, such that the selected length is as short as possible with little loss of accuracy. The evaluation criterion for accuracy loss is R@K, while the length loss is determined by the length output by the Actor network. The specific implementation of the Reward function is as follows:

[0081] Reward = Rc × Ra = recalli / recallall × c × (2 - c).

[0082] Wherein, c is equal to 1 - di / dall, representing the penalty of length, recalli represents the R@K of the selected length, recallall represents the R@K of the ancestor length, and the ratio of the two represents the accuracy loss. Stochastic Gradient Descent (SGD) is used for optimization during the training process.

[0083] Identical or similar reference numerals correspond to identical or similar components;

[0084] The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limitations of this patent;

[0085] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A multi-dimensional feature selection method based on deep learning, characterized in that, it includes the following steps: S1: Establish a deep learning network model G for general feature extraction of images; S2: Add a feature dimension random truncation model after the network model G; in step S2, the specific design process of the feature dimension random truncation model is: S21: Use a fully connected layer to map a feature vector X of a set length into a K-dimensional real vector, where K is the maximum allowable feature dimension size; S22: After encoding each vector into a real vector in S21, establish a dimension truncation module of the G network. Randomly select a dimension from the minimum dimension to the maximum dimension through the dimension truncation module and sequentially truncate it to obtain a feature of a random length. Use the same feature dimension in a small batch, and each time train the network with these features of different dimensions to obtain a general feature of the maximum length. Whenever a feature of a random length is needed, only need to sequentially truncate this general feature; S3: Through training on the training set, obtain the general features of the training set and the test set; S4: After obtaining the general features of the images, establish a feature dimension selection model; the specific process of step S4 is: S41: Establish an Actor network composed of several fully connected layers. The function of this network is to take the general feature of an image as the state input and output the predicted appropriate dimension; S42: Establish a Critic network composed of several fully connected layers. The function of this network is to take the general feature of an image as the state and the dimension output by the Actor network as the input, and output the score for the Actor network to optimize the Actor network; S43: Establish a Reward function, which returns a score for the dimension output by the Actor network, combining the length penalty of the dimension and the accuracy penalty of the actual evaluation standard as the supervision information of the Critic network; S5: Train and test the feature dimension selection model; S6: Establish a process for providing a background interface, provide a retrieval entry and return the retrieval result.

2. The multi-dimensional feature selection method based on deep learning according to claim 1, characterized in that, the specific process of step S1 is: S11: Establish a feature extraction layer of the G network, represent each frame of the picture in each preprocessed video as a low-dimensional real vector, and import the model pre-trained on a large-scale labeled photo into the Inception network; S12: Extract a set of feature vectors X of a set length for the images by training this Inception network.

3. The multi-dimensional feature selection method based on deep learning according to claim 2, characterized in that, the specific process of step S3 is: S31: Divide the data set into training data and test data; S32: The overall model needs to be trained. The training steps of the G network are as follows: Extract the image features by the G network, after randomly truncating the features through the feature dimension random truncation model, use the minimization of the loss function to train the G network model and train the parameters of the G network; S33: Pass the training set data through the feature extraction model G to obtain the general features of the maximum length, and store them in the database. For the test set data, after obtaining the full-length general features, evaluate the effectiveness of the general features.

4. The multi-dimensional feature selection method based on deep learning according to claim 3, wherein, the specific process of step S5 is: S51: Divide the data set into training data and test data; S52: The overall model needs to be trained. The training steps of the feature dimension selection model are as follows: Extract the general features of the image by the G network. Each time the model is updated, it is divided into two steps. The first step is to fix the Critic network, obtain the selected dimension through the Actor network, and use the score obtained by the Critic network to optimize the Actor network; The second step is to fix the Actor network, compare the score output by the Critic network for the Actor network with the score of the Reward function to supervise the training of the Critic network; S53: During testing, use the Actor network to obtain the selected dimension d, compare the distance with the first d dimensions of the general features of the training set in the database, obtain a ranking, and evaluate this ranking using the evaluation criteria.

5. The multi-dimensional feature selection method based on deep learning according to claim 4, wherein, the specific process of step S6 is: S61: Save the trained Inception model and the feature dimension selection model; S62: Create a background service process and reserve an interface for image input; S63: Through accessing the interface created in S62, input the image. Then the background service process in S62 will first preprocess the image into the input format required by the Inception model in S61; Next, call the Inception model saved in S61, input the processed image into the model, and obtain the generality of the image; Then pass through the feature dimension selection model in S61 to obtain the appropriate dimension size d, sequentially intercept this feature, calculate the distance with the first d dimensions of the image general feature data stored in the database, and sort them from small to large and return the first k images. The first k images are the retrieval results of the k most similar images.

6. The multi-dimensional feature selection method based on deep learning according to claim 5, wherein, In step S12, the feature extraction process is as follows: First, pre-train the Inception model with the imagenet image data set, and then perform fine-tuning; After each image passes through the pre-trained Inception model, a set of feature vectors of length k will be generated. This k refers to the maximum feature length of the image.

7. The multi-dimensional feature selection method based on deep learning according to claim 6, wherein, in step S22, the minimum dimension of the dimension interception module is set to 16, and the maximum dimension is set to 128.

8. The multi-dimensional feature selection method based on deep learning according to claim 7, wherein, In step S53, the Reward function combines the penalties of length and accuracy, such that the selected length is as short as possible with little loss of accuracy. The evaluation criterion for accuracy loss is R@K, while the length loss is determined by the length output by the Actor network. Stochastic Gradient Descent (SGD) is used for optimization during the training process.

Citation Information

Patent Citations

  • Offline meta-reinforcement learning model training method and device, equipment and storage medium

    CN112348113A

  • Grammatical error correction method, apparatus, computer system, and readable storage medium

    WO2021174823A1