A fish classification method based on Resnet and random forest fusion
By combining Resnet and random forest methods, the features of fish images are extracted and classified, and the problems of manual extraction of features and low classification accuracy in noise environments in the prior art are solved, thereby achieving higher fish classification accuracy.
Patent Information
- Application Number
- CN202111295872.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-03
AI Technical Summary
The existing fish image classification methods have problems with difficulty in manually extracting features and low classification accuracy in noise environments.
The fusion method based on Resnet and random forests is adopted to extract the surface and deep features of fish images through the Resnet network, and the random forest classifier is used for classification.
It improves the accuracy of fish classification, can effectively avoid interference in a noisy environment, and ensures the full extraction of feature information and the accuracy of classification.
Smart Images

Figure CN114219957B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing, and in particular relates to a fish classification method based on the fusion of Resnet and random forest. Background Art
[0002] In fishery and marine fish research, fish classification plays an important role in fish food processing and fish resource protection. There are many species of marine fish, and they are similar in color and appearance, which increases the difficulty of fish classification. Traditional fish classification methods mostly manually extract fish feature information, and extract features based on surface information such as fish outline, color, texture, etc. This method is very blind and cannot extract deep information of fish images.
[0003] In recent years, there have also been studies that use neural network models to automatically extract feature information for classification. Most of them use foreign publicly available data sets. Some data sets were collected a long time ago and the image clarity is not high. In addition, most of the fish that make up the data sets were collected in foreign waters, which does not conform to the actual distribution of fish species in the waters around my country.
[0004] Therefore, a marine fish classification method based on the combination of Resnet and random forest is proposed. A self-built dataset collected from the sea is used. The first few layers of multiple convolutional layers of the neural network are used to extract the surface information of the fish images, and the latter layers are used to extract the deep information of the images. Finally, the extracted features are classified using random forest. Summary of the invention
[0005] The purpose of the present invention is to solve the technical problems of difficulty in manually extracting features and low classification accuracy in a noisy environment in the existing extraction and processing of fish images.
[0006] A fish classification method based on Resnet and random forest fusion, which includes the following steps:
[0007] 1. Collect and select data sets, and collect several types of marine fish as the initial data sets;
[0008] 2. Preprocess the data set;
[0009] 3. Enhance the data set;
[0010] 4. Construct a feature extraction network and extract image features to extract feature information of several types of fish;
[0011] 5. Construct a random forest and put the characteristic information of several types of fish into the random forest classifier for training;
[0012] 6. Construct feature extraction network and random forest fish classification model;
[0013] The constructed fish classification model is used to classify fish.
[0014] In step 1, collect and take pictures of marine fish, and organize the corresponding labels and quantities of fish species.
[0015] In step 2, the following steps are included:
[0016] 1) Use the pre-trained multi-category target detection network model and the fish dataset publicly available on the Internet to train fish target detection;
[0017] 2) Carry out the target positioning of the fish picture, retain a small amount of background information, and crop out a rectangular area containing only the fish target.
[0018] In step 3, the following steps are included:
[0019] 1) Image flipping: flip the fish image horizontally or vertically;
[0020] 2) Image rotation: rotate the fish image by a certain angle;
[0021] 3) Perform image translation: translate the fish image horizontally or vertically;
[0022] 4) Perform contrast transformation: change the S and V brightness components of the fish image;
[0023] 5) Noise disturbance: add salt and pepper noise or Gaussian noise to the fish image;
[0024] Through the above steps, the number of images of each type of fish is expanded and randomly divided into several fish training sets and several fish verification sets.
[0025] In step 4, the following steps are included:
[0026] 1) Input the fish picture obtained in step 3 into the Resnet network. The input image is multi-channel and has a size of H×W. After a convolution kernel of F, the number of convolution kernels is K, and the step length is S, when using CNN to process image information, most of the edge pixels in the input image are only operated by the convolution kernel once, while the pixels in the middle of the image will be scanned multiple times; this reduces the reference degree of boundary information to a certain extent; on the other hand, after using zero padding, the new boundary has an impact on a certain part of the actual processing; to a certain extent, this problem can be solved; at the same time, input images of different sizes can be padded to make them of the same size; assuming the input size is (H, W), the post-size is (FH, FW), the output size is (OPH, OPW), the padding length is P, and the stride is S, then the output size formula is as follows: For a convolution operation with zero padding size P, the output size is H1×W1×K. The calculation formula is shown in formula (1) and formula (2):
[0027]
[0028]
[0029] 2) The result obtained in step 1) is passed through the Batch Normalization layer, which can accelerate network convergence and control overfitting. In the BN layer, the mean is μ, the variance is δ, the offset is β, the learnable parameter is γ, a small number (to prevent the denominator from being 0): ∈, the input is X, the output is Y, and the calculation formula is shown in formula (3) and formula (4):
[0030]
[0031] Y=γX1+β (4)
[0032] 3) After the modified linear unit ReLU activation function, the bias is: b, the weight is: W, the transposed matrix is T, the input is X, and the output is Y. The calculation formula is shown in formula (5):
[0033] Y=max(0,W T ×X+b) (5)
[0034] 4) After passing through the maximum pooling layer MAXPOOL, the feature map of the first stage is obtained;
[0035] 5) The output feature map of step 4) is input into the second stage. The second stage consists of 3 bottlenecks. The first bottleneck increases the image channel by 4 times, and the second and third bottlenecks keep the channel and image size unchanged to obtain the feature map;
[0036] 6) Input the output feature map of step 5) into the third stage, which consists of 4 bottlenecks. The first bottleneck increases the image channel by 2 times and reduces the image size by half. The second, third and fourth bottlenecks keep the channel and image size unchanged to obtain the feature map.
[0037] 7) Input the output feature map of step 6) into the fourth stage. The fourth stage consists of 6 bottlenecks. The first bottleneck increases the image channel by 2 times and reduces the image size by half. The second, third, fourth, fifth, and sixth bottlenecks keep the channel and image size unchanged to obtain the feature map.
[0038] 8) Input the output feature map of step 7) into the fifth stage, which consists of three bottlenecks. The first bottleneck increases the image channel by 2 times and reduces the image size by half. The second and third bottlenecks keep the channel and image size unchanged to obtain the feature map.
[0039] 9) The feature map obtained in step 8) is passed through an average pooling layer with an output size of 1×1 to reduce the number of calculated parameters and prevent overfitting, and then flattened to obtain an image feature vector with a dimension of D;
[0040] 10) The Resnet network uses the NLLLoss loss function. In the NLLLoss loss function, the input is x, the target is y, the weight is w, the batch size is N (starting from n = 1), and the output is l(x, y). The default is restored to "mean". Its calculation formula is shown in formula (6):
[0041]
[0042] 11) Update network weights, calculate losses of predicted results and actual results to update network nodes;
[0043] 12) Obtain the feature set of the training set, and pass the dataset through the feature extraction network updated by the above steps 1) to 11) to obtain an image feature set of size N and length D.
[0044] In step 5, the following steps are specifically included:
[0045] 1) Select a feature sample set. Use the image feature set obtained in step 4 to extract N samples from the image feature set by Bootstraping (sampling with replacement) in each round to obtain a sample subset of size N, which is used to train a decision tree. Perform X rounds of extraction and generate X decision trees in total.
[0046] 2) Generate a decision tree. The feature vector of each image in the image feature set is D-dimensional. In each round of decision tree generation, several features are randomly selected from the D features to form a new feature subset. The decision tree is generated by using the new feature subset. In the process of generating the decision tree, since the X decision trees are random in the selection of training sets and features, the decision trees are independent of each other. The feature selection of the decision tree nodes uses the "gini" quantitative evaluation standard, which indicates the probability of a randomly selected sample in the sample set being misclassified. The smaller the "gini" index is, the smaller the probability that the set and the selected sample are misclassified, that is, the higher the purity of the set is. The number of classification categories is K (starting from k = 1), and the probability that the sample belongs to the kth category is p k , the probability distribution is calculated according to formula (7):
[0047]
[0048] 3) Combine decision trees. Since the generated X decision trees are independent of each other and the importance of each decision tree is equal, there is no need to consider their weights when combining them, or they can be considered to have the same weights. For classification problems, the final classification result is determined by voting of all decision trees.
[0049] In step 6, the Resnet feature extractor constructed in step 4 and the random forest classifier constructed in step 5 are integrated to construct a Resnet+random forest fish classification model.
[0050] Compared with the prior art, the present invention has the following technical effects:
[0051] The present invention uses the YOLOv3 target detection network to locate fish images and cut off the redundant background information of the image, which can reduce the impact on the accuracy of subsequent fish classification; the Resnet50 convolutional neural network based on residual units can automatically extract the deep features of fish graphics; the random forest classifier based on decision trees can effectively avoid noise interference and accurately classify the extracted fish features. The model that uses the fusion of Resnet50 and random forest can ensure the accuracy of fish classification under the premise of fully extracting feature information. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:
[0053] Figure 1 is a flow chart of the present invention;
[0054] Figure 2 It is a schematic diagram of the structure of the Resnet50 residual unit in the embodiment;
[0055] Figure 3 It is a structural diagram of the Resnet50 feature extractor in the embodiment;
[0056] Figure 4 A schematic diagram of the structure of a random forest classifier in an embodiment; DETAILED DESCRIPTION
[0057] Embodiment: A fish classification method based on Resnet and random forest fusion. In the specific embodiment, Resnet50 is taken as an example. The specific implementation method flow chart is as follows: Figure 1 As shown, the following steps are included:
[0058] 1. Collect 40 types of marine fish as the initial data set.
[0059] 2. Fish dataset preprocessing.
[0060] 2.1 Use the fish dataset to train the pre-trained YOLOv3 target detection network for fish target detection.
[0061] 2.2 Use the trained YOLOv3 network to detect and locate fish in the initial dataset, and crop out the image dataset containing the fish target area.
[0062] 3. Dataset enhancement.
[0063] 3.1 Use a variety of data enhancement methods to enhance the fish dataset processed by YOLOv3.
[0064] 3.2 The enhanced dataset is randomly divided into 8,000 training sets and 2,000 validation sets.
[0065] 4. Construct the Resnet50 feature extraction network. The structure of the feature extraction network is as follows Figure 3 As shown in Figure 4, the characteristic information of 40 types of fish is extracted.
[0066] 5. Construct a random forest classifier. The structure of the classifier is as follows Figure 4 As shown in Figure 2, the characteristic information of 40 types of fish is put into the random forest classifier for training.
[0067] 6. Obtain a fish classification model that integrates Resnet50 and random forest.
[0068] Wherein, step 1 specifically includes:
[0069] 1.1 Manually collect and photograph marine fish images, each image contains only one fish;
[0070] 1.2 Initially collect 40 species of marine fish and organize the fish species labels and quantities accordingly;
[0071] Step 2 specifically includes:
[0072] 2.1 Use the pre-trained multi-category target detection network YOLOv3 model, and then use the large fish dataset publicly available on the Internet for fish target detection training.
[0073] 2.2 The fish dataset obtained in step 1 is passed through the YOLOv3 network to locate the target of the fish image, retain a small amount of background information, and crop out a rectangular area containing only the fish target.
[0074] Step 3 specifically includes:
[0075] 3.1 Image flip: Flip the fish image horizontally or vertically.
[0076] 3.2 Image rotation: Rotate the fish image by 90, -90 or 180 degrees.
[0077] 3.3 Image translation: translate the fish image horizontally or vertically.
[0078] 3.4 Contrast transformation: changing the S and V brightness components of the fish image.
[0079] 3.5 Noise perturbation: Add salt and pepper noise or Gaussian noise to the fish image.
[0080] 3.6 After the above five steps, the number of fish of each type is expanded to 250, and the total number of images of forty types of fish is 10,000, which are randomly divided into 8,000 fish training sets and 2,000 fish verification sets.
[0081] Step 4 specifically includes:
[0082] 4.1 Input the fish image obtained in step 3 into the Resnet50 network. The residual unit structure is as follows: Figure 2 As shown, the Resnet50 network structure is as follows Figure 3As shown in the figure, the input image is 3 channels and the size is 224×224. After a convolution kernel of F=7×7, the number of convolution kernels is K=64, and the stride is S=2, when using CNN to process image information, most of the edge pixels in the input image are only operated once by the convolution kernel, while the pixels in the middle of the image will be scanned multiple times. This reduces the reference degree of boundary information to a certain extent. On the other hand, after using zero padding, the new boundary has an impact on a certain part of the actual processing. To some extent, this problem can be solved. At the same time, input images of different sizes can be padded to make them of the same size. Assume that the input size is (H, W), the post-size is (FH, FW), the output size is (OPH, OPW), the padding length is P, and the stride is S. The output size formula is as follows: For a convolution operation with zero padding size P, the output size is H1×W1×K. The calculation formula is shown in equations (1) and (2).
[0083]
[0084]
[0085] 4.2 The result obtained in step 4.1 is passed through the Batch Normalization layer to accelerate network convergence and control overfitting. In the BN layer, the mean is μ, the variance is δ, the offset is β, the learnable parameter is γ, a small number (to prevent the denominator from being 0): ∈, the input is X, the output is Y, and the calculation formula is shown in Equation (3) and Equation (4).
[0086]
[0087] Y=γX1+β (4)
[0088] 4.3 After the modified linear unit ReLU activation function, the bias is: b, the weight is: W, the transposed matrix is T, the input is X, the output is Y, and the calculation formula is shown in formula (5).
[0089] Y=max(0,W T ×X+b) (5)
[0090] 4.4 After the maximum pooling layer MAXPOOL, the feature map size of the first stage is 64×56×56.
[0091] 4.5 Input the output feature map of step 4.4 into the second stage. The second stage consists of 3 bottlenecks. The first bottleneck increases the image channel by 4 times, and the second and third bottlenecks keep the channel and image size unchanged to obtain a 256×56×56 feature map.
[0092] 4.6 Input the output feature map of step 4.5 into the third stage. The third stage consists of 4 bottlenecks. The first bottleneck increases the image channels by 2 times and reduces the image size by half. The second, third and fourth bottlenecks keep the channel and image size unchanged to obtain a 512×28×28 feature map.
[0093] 4.7 Input the output feature map of step 4.6 into the fourth stage. The fourth stage consists of 6 bottlenecks. The first bottleneck increases the image channels by 2 times and reduces the image size by half. The second, third, fourth, fifth, and sixth bottlenecks keep the channel and image size unchanged to obtain a 1024×14×14 feature map.
[0094] 4.8 Input the output feature map of step 4.7 into the fifth stage. The fifth stage consists of 3 bottlenecks. The first bottleneck increases the image channels by 2 times and reduces the image size by half. The second and third bottlenecks keep the channel and image size unchanged to obtain a 2048×7×7 feature map.
[0095] 4.9 The feature map obtained in step 4.8 is passed through an average pooling layer with an output size of 1×1 to reduce the number of calculated parameters and prevent overfitting. After a flattening operation, an image feature vector with a dimension of 2048 is obtained.
[0096] 4.10 The Resnet50 network uses the NLLLoss loss function. In the NLLLoss loss function, the input is x, the target is y, the weight is w, the batch size is N (starting from n = 1), the output is l(x, y), and the default is restored to "mean". Its calculation formula is shown in formula (6).
[0097]
[0098] 4.11 Update of network weights. Update network nodes by calculating the loss of predicted results and actual results.
[0099] 4.12 Obtain the feature set of the training set. The dataset is passed through the feature extraction network updated by the above steps 4.1 to 4.11 to obtain an image feature set of size 10000 and length 2048.
[0100] Step 5 specifically includes:
[0101] 5.1 Selection of feature sample set. The image feature set obtained in step 4.12 has a total of 10,000 samples. In each round, 10,000 samples are extracted from the image feature set by Bootstraping (sampling with replacement) to obtain a sample subset of 10,000 to train a decision tree. A total of 90 rounds of extraction were performed, and a total of 90 decision trees were generated.
[0102] 5.2 Generation of decision trees. The feature vector of each image in the image feature set is 2048-dimensional. In each round of decision tree generation, 50 features are randomly selected from the 2048 features to form a new feature subset, and the decision tree is generated by using the new feature subset. In the process of generating decision trees, since these 90 decision trees are random in the selection of training sets and features, each decision tree is independent of each other. The feature selection of decision tree nodes uses the "gini" quantitative evaluation standard, which indicates the probability of a randomly selected sample in the sample set being misclassified. The smaller the "gini" index, the smaller the probability that the selected sample in the set is misclassified, that is, the higher the purity of the set. The number of classification categories is K (starting from k=1), and the probability that the sample belongs to the kth category is p. k , the probability distribution is calculated according to formula (7).
[0103]
[0104] 5.3 Combination of decision trees. Since the 90 decision trees generated are independent of each other and the importance of each decision tree is equal, there is no need to consider their weights when combining them, or they can be considered to have the same weights. For classification problems, the final classification result is determined by voting of all decision trees.
[0105] Step 6 specifically includes:
[0106] 6.1 Integrate the Resnet50 feature extractor constructed in step 4 and the random forest classifier constructed in step 5 to construct a fish classification model of Resnet50+random forest.
[0107] The present invention proposes a marine fish classification method based on the combination of Resnet and random forest, and uses a self-built data set collected from the sea area. The first few layers of multiple convolutional layers of the neural network are used to extract surface information of fish images, and then the latter layers are used to extract deep information of the images. Finally, the extracted features are classified using random forest; the accuracy of fish classification can be guaranteed under the premise of sufficient extraction of feature information.
Claims
1. A fish classification method based on Resnet and random forest fusion, characterized in that: It includes the following steps: Step 1: Collect and select data sets, and collect several types of marine fish as the initial data set; Step 2: Preprocess the data set; Step 3: Enhance the data set; Step 4: construct a feature extraction network and extract image features to extract feature information of several types of fish; Step 5: Construct a random forest and put the characteristic information of several types of fish into the random forest classifier for training; Step 6: Construct feature extraction network and random forest fish classification model; The constructed fish classification model is used to classify fish; In step 3, the following steps are included: Step 3-1: Flip the image: flip the fish image horizontally or vertically; Step 3-2: Perform image rotation: rotate the fish image by a certain angle; Step 3-3: Perform image translation: translate the fish image horizontally or vertically; Step 3-4: Perform contrast transformation: change the S and V brightness components of the fish image; Step 3-5: Noise disturbance: add salt and pepper noise or Gaussian noise to the fish image; Through the above steps, the number of images of each type of fish is expanded, and they are randomly divided into several fish training sets and several fish verification sets; In step 4, the following steps are included: Step 4-1: Input the fish picture obtained in step 3 into the Resnet network. The input image is multi-channel and has a size of H×W. It passes through a convolution kernel of F, the number of convolution kernels is K, and the step length is S. When using CNN to process image information, most of the edge pixels in the input image are only operated once by the convolution kernel, while the pixels in the middle of the image will be scanned multiple times; this reduces the reference degree of boundary information to a certain extent; on the other hand, after using zero padding, the new boundary has an impact on a certain part of the actual processing; at the same time, input images of different sizes can be padded to make them consistent in size; assuming the input size is (H, W), the post-size is (FH, FW), the output size is (OPH, OPW), the padding length is P, and the stride is s, then the output size formula is as follows: For a convolution operation with a zero padding size of P, the output size is H1×W1×K. The calculation formula is shown in formula (1) and formula (2): Step 4-2: Pass the result of step 4-1 through the Batch Normalization layer to accelerate network convergence and control overfitting. In the BN layer, the mean is μ, the variance is δ, the offset is β, and the learnable parameter is γ; a small number is ∈, which is used to prevent the denominator from being 0; the input is X, the output is Y, and the calculation formula is shown in formula (3) and formula (4): Y=γX1+β(4) Step 4-3: After the modified linear unit ReLU activation function, the bias is b, the weight is W, the transposed matrix is T, the input is X, and the output is Y. The calculation formula is shown in formula (5): Y=max(0,W T ×X+b) (5) Step 4-4: After passing through the maximum pooling layer MAXPOOL, the feature map of the first stage is obtained; Step 4-5: Input the output feature map of step 4-4 into the second stage. The second stage consists of 3 bottlenecks. The first bottleneck increases the image channel by 4 times, and the second and third bottlenecks keep the channel and image size unchanged to obtain the feature map. Step 4-6: Input the output feature map of step 4-5 into the third stage. The third stage consists of 4 bottlenecks. The first bottleneck increases the image channel by 2 times and halves the image size. The second, third and fourth bottlenecks keep the channel and image size unchanged to obtain the feature map. Step 4-7: Input the output feature map of step 4-6 into the fourth stage. The fourth stage consists of 6 bottlenecks. The first bottleneck increases the image channel by 2 times and halves the image size. The second, third, fourth, fifth, and sixth bottlenecks keep the channel and image size unchanged to obtain the feature map. Step 4-8: Input the output feature map of step 4-7 into the fifth stage. The fifth stage consists of 3 bottlenecks. The first bottleneck increases the image channel by 2 times and halves the image size. The second and third bottlenecks keep the channel and image size unchanged to obtain the feature map. Step 4-9: The feature map obtained in step 4-8 is passed through an average pooling layer with an output size of 1×1 to reduce the number of calculated parameters and prevent overfitting, and then flattened to obtain an image feature vector with a dimension of D; Step 4-10: The Resnet network uses the NLLLoss loss function. In the NLLLoss loss function, the input is x, the target is y, the weight is w, the batch size is N, and the default is restored to "mean". Its calculation formula is shown in formula (6): Step 4-11: Update the network weights, calculate the loss of the predicted results and the actual results to update the network nodes; Step 4-12: Obtain the feature set of the training set, and pass the dataset through the feature extraction network updated by the above steps 4-1 to 4-11 to obtain an image feature set of size N and length D.
2. The method according to claim 1, characterized in that In step 1, collect and take pictures of marine fish, and organize the corresponding labels and quantities of fish species.
3. The method according to claim 2, characterized in that In step 2, the following steps are included: Step 2-1: Use the pre-trained multi-category target detection network model and the publicly available fish dataset to perform fish target detection training; Step 2-2: Locate the target of the fish image, retain a small amount of background information, and crop out a rectangular area containing only the fish target.
4. The method according to claim 1, characterized in that: In step 5, the following steps are specifically included: Step 5-1: Select the feature sample set. Use the image feature set obtained in step 4 to extract N samples from the image feature set by Bootstraping in each round to obtain a sample subset of size N, which is used to train a decision tree. Perform X rounds of extraction and generate X decision trees in total. Step 5-2: Generate a decision tree. The feature vector of each image in the image feature set is D-dimensional. In each round of decision tree generation, several features are randomly selected from the D features to form a new feature subset. The decision tree is generated by using the new feature subset. In the process of generating the decision tree, since the X decision trees are random in the selection of training sets and features, the decision trees are independent of each other. The feature selection of the decision tree nodes uses the "gini" quantitative evaluation standard, which indicates the probability of a randomly selected sample in the sample set being misclassified. The smaller the "gini" index is, the smaller the probability of the set and the selected sample being misclassified, that is, the higher the purity of the set is. The number of classification categories is K. Starting from k=1, the probability that the sample belongs to the kth category is p k , the probability distribution is calculated according to formula (7): Step 5-3: Combine decision trees. Since the generated X decision trees are independent of each other and the importance of each decision tree is equal, there is no need to consider their weights when combining them, or they can be considered to have the same weights. For classification problems, the final classification result is determined by voting of all decision trees.
5. The method according to claim 1, characterized in that In step 6, the Resnet feature extractor constructed in step 4 and the random forest classifier constructed in step 5 are integrated to construct a Resnet+random forest fish classification model.
Citation Information
Patent Citations
Retina OCT image classification method based on deep learning subnetwork feature extraction
CN110188820A
Face recognition method and face recognition system based on deep neural network and random forest
CN110414483A