Oyster meat yield prediction method based on multi-modal fusion
By employing a multimodal fusion method, this study extracts the appearance and shape features of oysters using a self-attention mechanism and a variational autoencoder, and constructs a multilayer perceptron regression model. This overcomes the limitations of existing technologies in estimating oyster meat yield and achieves higher-precision prediction and classification.
Patent Information
- Application Number
- CN202510987266.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies for estimating oyster meat yield using computer vision mainly rely on single-factor features, which has limitations, resulting in low sorting efficiency and large measurement errors, making it difficult to accurately classify oyster quality.
A multimodal fusion method is adopted, which extracts appearance image features through a self-attention mechanism encoder and obtains shape features through a variational autoencoder. The Concat method is combined to perform feature fusion, and a multilayer perceptron regression prediction model is constructed to predict the meat yield of oysters.
It significantly improves the accuracy of oyster meat yield prediction, avoids the problems of low efficiency and large measurement errors in manual sorting, and achieves more accurate oyster quality classification.
Smart Images

Figure CN120877012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oyster meat yield prediction technology, and specifically to a method for predicting oyster meat yield based on multimodal fusion. Background Technology
[0002] Oysters are an important source of nutrition in the human diet, and the daily demand is rapidly increasing. Aquaculture is one of the main ways to obtain oysters. From harvesting to sale, oysters need to be sorted. Currently, the market mainly uses manual weighing to separate oysters into different qualities, and then packages them for sale as equal quality. The oyster shell not only protects its soft tissues from external environmental pollution but also effectively slows down moisture loss, maintaining the oyster's freshness. However, the shielding effect of the oyster shell presents challenges for manual sorting.
[0003] Oyster meat yield refers to the ratio of the edible portion weight of an oyster to its total weight, and is an important indicator of oyster quality. A high meat yield indicates a good growing environment, tender meat, and superior taste. Conversely, if the oysters are contaminated or malnourished during growth, the meat yield is often lower. Therefore, accurate classification of oyster quality based on meat yield is crucial for maximizing the benefits of aquaculture. Computer vision technology, due to its non-invasive and non-radioactive characteristics, has become an important tool in calculating oyster meat yield. However, current computer vision technology mainly uses single-factor features to estimate weight, which has certain limitations in quality classification. Summary of the Invention
[0004] To address the above problems, this invention provides a method for predicting oyster meat yield based on multimodal fusion, comprising:
[0005] Raw images of oysters were collected, and appearance image data and shape feature data were obtained through a segmentation network to construct a multimodal dataset, which includes appearance image data and shape feature data.
[0006] For appearance image data, a self-attention mechanism encoder is used to extract global image features to obtain appearance feature vectors; for shape feature data, a variational autoencoder is used to map key features to obtain shape feature vectors.
[0007] Based on the Concat method, the appearance feature vector and the shape feature vector are fused to construct a fused feature vector;
[0008] A multilayer perceptron regression prediction model is constructed, and the multilayer perceptron regression prediction model is trained with the fused feature vector as input and the meat yield as the target output.
[0009] The effectiveness of the trained multilayer perceptron regression prediction model is validated using a test set. If the validation results meet the target requirements, the multilayer perceptron regression prediction model is saved.
[0010] Furthermore, the process of acquiring the appearance image data and the shape feature data includes:
[0011] Collect raw oyster image data and preprocess the image data;
[0012] The target region is segmented using a segmentation network model;
[0013] The segmentation network model is trained to generate segmentation masks and extract target contours to obtain shape feature data;
[0014] By using a mask to set the background area of the original image to black while preserving the target area, the apparent image data is obtained.
[0015] Furthermore, the self-attention mechanism captures different relationships and features of the input data through different self-attention heads. The attention weight formula is as follows:
[0016]
[0017] Where Attention(·) represents the vector multiplication operation, Q, K, and V represent the linear mappings of query, key, and value, respectively, T represents the matrix transpose, and d k It represents the mapping dimension, and softmax represents the activation function;
[0018] To enhance the model's expressive power, a multi-head self-attention mechanism is used, which involves calculating multiple sets of Q, K, and V, and then concatenating the results. The formula is as follows:
[0019] MultiHeadSelf Attention(Z)=Concat(head1, head2,..., head h W O ;
[0020] head i =Attention(ZW Qi ZW Ki ZW Vi );
[0021] Among them, W O Z represents the final output mapping weights, and Z represents the input sequence.
[0022] The output of the self-attention mechanism is further processed by a feedforward network, and the formula is as follows:
[0023] Y′=FeedForward(Y)=ReLU(TW F1 +b F1 W F2 +b F2 ;
[0024] Where Y′ represents the output of the feedforward network, Y represents the output of the self-attention mechanism, and ReLU represents the activation function;
[0025] By introducing residual connectivity and layer normalization, the flow of information between layers and the stability of the model are increased, as expressed in the following formula:
[0026] Y′=LayerNorm(Y+Z);
[0027] LayerNorm represents the layer normalization operation.
[0028] Furthermore, for the self-attention mechanism encoder extracting global image features, when the feature dimension is much larger than the number of samples, principal component analysis is selected to reduce the dimensionality of the extracted global features. By retaining the principal components with the largest variance, the feature directions that contribute the most to the data distribution are selected. First, the original data is centered, i.e., using the formula:
[0029] X centered =X-μ;
[0030] Among them, X centered Let μ represent the centered data matrix, and μ represent the mean vector of each feature dimension.
[0031] Through singular value pairs X centered The decomposition is performed, and the variance contribution rate is calculated based on the singular values. The expression is as follows:
[0032] X centered =U∑V T ;
[0033]
[0034] Where U represents the left singular vector matrix, ∑ represents the diagonal extension matrix containing singular values, V represents the right singular vector matrix, V represents the transpose, and σ i Let be the i-th singular value, representing the variance intensity of the corresponding principal component direction. To complete the data dimensionality reduction, the principal components corresponding to the first k largest singular values are selected to form the projection matrix w.
[0035] Finally, by multiplying the centered data matrix by the projection matrix, data dimensionality reduction is achieved, expressed as X. pca =X centered ·W.
[0036] Furthermore, the variational autoencoder obtains the shape feature vector by: inputting data into the encoder, which outputs two codes: one is the original code μ(μ1, μ2, μ3), and the other is the control noise code σ(σ1, σ2, σ3). The control noise code σ(σ1, σ2, σ3) is weighted with random noise e(e1, e2, e3). Finally, the original code and the weighted noise code are added together to obtain the latent vector z(z1, z2, z3) of the variational autoencoder model.
[0037] Furthermore, a loss function is used in the construction of the variational autoencoder model. The loss function includes the reconstruction error and the KL divergence, and its formula is:
[0038] L=D KL (q(z|x)||p(z))-E q(z|x) [log p(x|z)];
[0039] Where q(z|x) represents the latent distribution generated by the encoder, z represents the latent variable, x represents the input data, p(z) represents the prior distribution of z, and p(x|z) represents the output of the latent variable z transformed back into the original data space.
[0040] Furthermore, the multilayer perceptron regression prediction model is constructed as follows: an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer; the number of neurons in the input layer is the same as the dimension of the concatenated feature vector; the number of neurons in the first, second, and third hidden layers is halved layer by layer to gradually compress redundant information and retain key features; the output layer has a dimension of 1 and directly predicts the meat yield; batch normalization and ReLU activation functions are used between each layer of the multilayer perceptron regression prediction model; Dropout is set to 0.2 in the first and second hidden layers to suppress noise, and Dropout is set to 0.1 in the third hidden layer to retain effective features.
[0041] The beneficial effects of this invention are as follows:
[0042] 1. This invention constructs a multilayer perceptron regression prediction model. Specifically, by extracting features from the appearance image and shape parameters of oysters, and using the Concat method to fuse feature vectors, a prediction model suitable for oyster meat yield is constructed. Based on the prediction model with multi-factor input, it can better capture the interaction patterns between different dimensions of features, significantly improve the model prediction accuracy, and avoid the problems of low efficiency and large measurement error in manual sorting in the prior art.
[0043] 2. The present invention designs a dual-branch-segmentation architecture feature extraction network to achieve specialized division of labor. The self-attention branch strengthens the expression of key features by calculating pixel-level association weights, while the autoencoder branch constructs a probability distribution representation of shape parameters through latent space modeling. The synergistic effect of the two improves the prediction accuracy of the model. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the oyster meat yield prediction method based on multimodal fusion provided in an embodiment of the present invention;
[0045] Figure 2 The experimental comparison diagram is provided in the embodiments of the present invention to verify the effectiveness of the segmentation network model in extracting data;
[0046] Figure 3 A schematic diagram of the overall structure of the multimodal feature fusion learning model provided in an embodiment of the present invention;
[0047] Figure 4 This is a comparison chart of the model prediction results and the actual measurement results of the meat yield provided in the embodiments of the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0049] Figure 1 This is a schematic diagram of the oyster meat yield prediction method based on multimodal fusion provided in an embodiment of the present invention, as shown below. Figure 1 As shown, this invention provides a method for predicting oyster meat yield based on multimodal fusion, comprising:
[0050] Step S01: Acquire raw images of oysters and obtain appearance image data and shape feature data through a segmentation network to construct a multimodal dataset, which includes appearance image data and shape feature data;
[0051] Step S02: For appearance image data, extract global image features through a self-attention mechanism encoder to obtain appearance feature vectors; for shape feature data, map key features through a variational autoencoder to obtain shape feature vectors.
[0052] Step S03: Based on the Concat method, the appearance feature vector and the shape feature vector are fused to construct a fused feature vector;
[0053] Step S04: Construct a multilayer perceptron regression prediction model, using the fused feature vector as input and the meat yield as the target output to train the multilayer perceptron regression prediction model;
[0054] Step S05: Use the test set to verify the effectiveness of the trained multilayer perceptron regression prediction model. If the verification results meet the target requirements, save the multilayer perceptron regression prediction model.
[0055] Specifically, this embodiment selected 184 oyster samples from the same aquaculture farm in Shandong Province. A strict random sampling method was used to ensure the randomness and representativeness of the samples. Images were acquired using a Canon EOS 5D MarkIV camera (Canon China), and weighed using a JA3003 electronic analytical balance (sensitivity 1mg, Shanghai Licheng Bangxi Instrument Technology Co., Ltd., China). The wet weight of the oysters and the total wet weight were obtained separately. The formula for calculating the meat yield in this invention is: Meat yield = Wet weight of oysters / Total wet weight. Simultaneously, the length, width, and other shape features were manually measured using an S102-107-101 vernier caliper (Shanghai Yonghui Industrial Development Co., Ltd., China) to verify the effectiveness of the shape features measured by the machine learning method. The process of acquiring the original image data in this embodiment is as follows: oysters are placed on a black background, and the camera is fixed above the black background to take pictures from above. During the acquisition process, the distance between the camera and the table remains unchanged. The front and back of each oyster are photographed and acquired. When acquiring oyster images, the camera's exposure time is set to 1 / 80s, ISO 12800, focal length 100mm, and the image pixel size is fixed at 4480*4480.
[0056] Specifically, the process of acquiring appearance image data and shape feature data includes:
[0057] (1) Collect raw oyster image data and preprocess the image data, such as adjusting the image size and normalizing the image;
[0058] (2) Use a segmentation network model to segment the target region;
[0059] (3) A segmentation mask is generated by training a segmentation network model, the target contour is extracted, and then the shape feature numerical data is extracted using a variety of digital image processing methods.
[0060] (4) Use a mask to set the background area of the original image to black, and only retain the target area to obtain the apparent image data.
[0061] In this embodiment, the extraction of shape feature numerical data using a segmentation network model further includes: extracting the target contour using a segmentation mask, measuring multiple shape features, such as length, width, area, perimeter, convex shell length, and convex shell area, etc., and further standardizing the shape data so that the values of the shape feature attributes are independent of the size of the oyster, resulting in new shape feature attributes such as length eccentricity, width eccentricity, roughness, compactness, elongation, and fullness, as detailed in Table 1 below:
[0062] Table 1. Definitions and descriptions of oyster morphological characteristics.
[0063]
[0064] To verify the effectiveness of the segmentation network model in extracting shape feature numerical data, the experimentally measured major and minor axes were compared with those measured manually. The experimental results are as follows: Figure 2 As shown in the figure, due to space limitations, only the experimental results of a portion of the samples are presented. It can be seen from the figure that the data measured in the experiment are similar to the data measured manually. The data obtained by the two measurement methods have a high degree of consistency, which verifies the effectiveness of the model extraction.
[0065] This invention constructs an oyster meat yield prediction model based on multimodal fusion learning. This model deeply fuses the shape and appearance features of oysters, building a complete multimodal fusion learning framework that includes a shape feature variational autoencoder extraction network, an appearance feature self-attention mechanism extraction network, a feature fusion module, and regression prediction. Figure 3 As shown;
[0066] Specifically, a self-attention mechanism encoder is first used to extract global features of the image, and then principal component analysis is used to reduce the dimensionality of the high-dimensional features to address the overfitting problem that may occur when the feature dimension is much larger than the number of samples. The self-attention mechanism overcomes the limitations of convolutional neural networks in handling long-distance dependencies by adding class labels at the beginning of the sequence and performing classification in a linear layer. By capturing long-distance dependencies using the self-attention mechanism, the complexity of manually designing convolutional kernels is avoided, and interactive features can be learned adaptively.
[0067] In this embodiment, the image is divided into 16×16 blocks. A 224×224 image yields 196 blocks. Each block is flattened and embedded into a 768-dimensional vector through a linear layer, resulting in a two-dimensional matrix of size [196, 768] for each image. A special category label, clsToken, is then added to the embedded matrix as a trainable parameter. This label has the same data format as the other block vectors, being a single vector of size [1, 768]. The category label is concatenated before the block vectors to obtain a two-dimensional matrix of size [197, 768]. Position encoding uses trainable parameters, directly superimposed on the vectors element-wise. Therefore, the shape of the two-dimensional matrix remains unchanged before and after position encoding. This two-dimensional matrix is input into the encoder for feature extraction, and the output dimension remains unchanged after feature extraction. The first vector (the output vector of clsToken) is extracted from the output as the global feature representation of the entire image, and the apparent features extracted from the feature two-dimensional matrix [1, 768] are saved.
[0068] The self-attention mechanism captures different relationships and features of the input data through different self-attention heads. The attention weight formula is as follows:
[0069]
[0070] Where Attention(·) represents the vector multiplication operation, Q, K, and V represent the linear mappings of query, key, and value, respectively, T represents the matrix transpose, and d k It represents the mapping dimension, and softmax represents the activation function;
[0071] To enhance the model's expressive power, a multi-head self-attention mechanism is used, which involves calculating multiple sets of Q, K, and V, and then concatenating the results. The formula is as follows:
[0072] MultiHeadSelf Attention(Z)=Concat(head1,head2,…,head h W O ;
[0073] head i =Attention(ZW Qi ZW Ki ZW Vi );
[0074] Among them, W O Z represents the final output mapping weights, and Z represents the input sequence.
[0075] The output of the self-attention mechanism is further processed by a feedforward network, and the formula is as follows:
[0076] Y′=FeedForward(Y)=ReLU(TW F1 +b F1 W F2 +b F2 ;
[0077] Where Y′ represents the output of the feedforward network, Y represents the output of the self-attention mechanism, and ReLU represents the activation function;
[0078] By introducing residual connectivity and layer normalization, the flow of information between layers and the stability of the model are increased, as expressed in the following formula:
[0079] Y″ = LayerNorm(Y+Z);
[0080] LayerNorm represents layer normalization, which helps alleviate the vanishing gradient problem during training and improves the training stability of the model. This structure enables the model to better capture information in the input sequence while retaining the important features of the original input.
[0081] When the feature dimension (768) is much larger than the number of samples (368), the data is extremely sparse in the high-dimensional space, which can easily cause the model to remember the noise in the training data and fail to learn the generalization rules better. At this time, it is necessary to use principal component analysis to reduce the dimensionality of the extracted global features. By retaining the principal components with the largest variance, the feature directions that contribute the most to the data distribution are selected. First, the original data is centered, that is, by using the formula:
[0082] X centered =X-μ;
[0083] Among them, X centered Let X represent the centered data matrix, where μ represents the mean vector of each feature dimension. The purpose of this step is to eliminate data bias and ensure the stability of subsequent analysis. The centered data matrix X is... centered ∈R, where R contains 768-dimensional features from 368 samples;
[0084] Next, through singular value pairs X centered The decomposition is performed, and the variance contribution rate is calculated based on the singular values. The expression is as follows:
[0085] X centered =U∑V T ;
[0086]
[0087] Where U represents the left singular vector matrix, ∑ represents the diagonal extension matrix containing singular values, V represents the right singular vector matrix, V represents the transpose, and σ iLet be the i-th singular value, representing the variance intensity along the corresponding principal component direction. To achieve dimensionality reduction, we need to select the principal components corresponding to the k largest singular values. To prevent overfitting, we choose a data matrix reduced to 64 dimensions with a cumulative variance greater than 95%, i.e., extracting 64 columns from the right singular vector matrix V to form the projection matrix W∈R. 768×64 ;
[0088] Finally, by multiplying the centered data matrix by the projection matrix, its expression is X. pca =X centered ·W can map the original 768-dimensional features to a 64-dimensional low-dimensional space, thus completing the data dimensionality reduction. This process effectively compresses the data dimensionality by retaining the principal components with the largest variance, while preserving the original information to the greatest extent.
[0089] Specifically, the variational autoencoder obtains the shape feature vector by: inputting data into the encoder, which outputs two codes: one is the original code μ(μ1, μ2, μ3), and the other is the control noise code σ(σ1, σ2, σ3). The control noise code σ(σ1, σ2, σ3) is weighted with random noise e(e1, e2, e3). Finally, the original code and the weighted noise code are added together to obtain the latent vector z(z1, z2, z3) of the variational autoencoder model.
[0090] In this embodiment, by introducing weighted noise, the model can be more resistant to changes in input data during training, exhibiting robustness. The extracted features are more generalizable in the latent space, better capturing data variation factors, thereby improving the performance of downstream classification tasks. This invention innovatively proposes using the encoded latent vector z as the shape feature extraction result, which has significant representativeness and discriminativeness, providing a good data foundation for subsequent regression prediction.
[0091] Specifically, a loss function is used in constructing the variational autoencoder model. This loss function includes reconstruction error and KL divergence. Reconstruction error measures the difference between the original and reconstructed data, evaluating the model's ability to reconstruct the input data x given a latent variable z. KL divergence measures the difference between the latent distribution q(z|x) of the encoder output and the standard normal distribution p(z). The general formula is:
[0092] L=D KL (q(z|x)||p(z))-E q(z|x) [log p(x|z)];
[0093] Where q(z|x) represents the latent distribution generated by the encoder, z represents the latent variable, x represents the input data, p(z) represents the prior distribution of z, and p(x|z) represents the output of the latent variable z transformed back into the original data space;
[0094] More specifically, the original data input to this invention has a shape of [368, 6], and a latent dimension d is selected. During training, a complete model needs to be constructed to evaluate the effectiveness of feature extraction, including an encoder where fully connected layers are used for the hidden layers; the decoder and encoder have symmetrical structures, ReLU activation is used to enhance nonlinearity, and the Sigmoid output is adapted to the normalized data. To further determine the dimension of the latent space in the model for experimental verification, since the sample size is small and an excessively high latent dimension may lead to redundancy in some dimensions or failure to effectively encode useful information, it is necessary to control the latent dimension within a reasonable range. This invention ultimately determines the latent dimension d∈{4, 5, 6, 7}. The model performance is confirmed by comparing the reconstruction error and directly predicting the meat yield using shape features, and the most suitable dimension is selected. The specific results are shown in Table 2. The experimental results show that when d=6, the model achieves the best balance between reconstruction ability and generalization. Therefore, a 6-dimensional vector is finally selected as the result of shape feature extraction.
[0095] Table 2. Model performance under different latent dimensions
[0096] Potential Dimensions Reconstruction error <![CDATA[R 2 ]]> 4 0.12 0.85 5 0.09 0.88 6 0.08 0.92 7 0.07 087
[0097] This invention employs the Concat method to concatenate and fuse feature vectors, and then inputs the concatenated feature vector into a multilayer perceptron model for regression prediction to construct an oyster meat yield prediction model. In this embodiment, the feature fusion process specifically includes: first, reducing the 768-dimensional appearance feature vector extracted by the multi-head self-attention mechanism to 64 dimensions using principal component analysis, while retaining the 6-dimensional shape feature vector extracted by the variational autoencoder; then, concatenating the two using the Concat method to generate a 70-dimensional feature vector. This feature vector serves as the input to the regression predictor of the multilayer perceptron model for predicting oyster meat yield.
[0098] Specifically, to construct a multilayer perceptron regression prediction model, and considering an input feature dimension of 70, the neural network structure designed in this invention includes: an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. The number of neurons in the input layer is set to 70, the same as the dimension of the concatenated feature vector. The number of neurons in the first hidden layer is set to 128, approximately twice the input dimension, which expands the feature representation capability while avoiding parameter explosion. The number of neurons in the second hidden layer is set to 64, halved layer by layer, gradually compressing redundant information and retaining key features. The number of neurons in the third hidden layer is further compressed to 32 to reduce model complexity and prevent overfitting. The output layer has a dimension of 1, directly predicting the meat yield. Batch normalization and ReLU activation functions are used between each layer of the multilayer perceptron regression prediction model to stabilize the training process and improve generalization ability. Dropout is set to 0.2 in the first and second hidden layers to suppress noise, and Dropout is set to 0.1 in the third hidden layer to retain effective features. After multiple experiments, it has been verified that the 3-hidden-layer design can effectively capture the nonlinear relationship of input features, while avoiding the high training difficulty caused by the complexity of the network structure.
[0099] To verify the effectiveness of the constructed multilayer perceptron regression prediction model, the model prediction results were compared and analyzed with the actual measurement results. The experimental results are as follows: Figure 4 As shown in the comparison results, the prediction results are basically consistent with the actual data, indicating that the multilayer perceptron regression prediction model constructed in this invention has high feasibility and accuracy in predicting oyster meat yield.
[0100] More specifically, to verify the rationality of extracting certain parts in the predictive model, this invention conducts ablation experiments. First, the appearance feature extraction part is replaced with a ResNet residual network. As a classic convolutional network, ResNet can effectively capture local detail features through its residual structure. Second, the shape feature extraction part is replaced with an autoencoder (AE). Each part of the model is replaced individually, or both parts are replaced simultaneously, to comprehensively evaluate the impact of the feature extraction module. To evaluate the rationality of the regression model, the Multilayer Perceptron (MLP) regression prediction model and multinomial regression are compared. In this ablation experiment, all multimodal models use the Concat concatenation and fusion method.
[0101] Among them, the mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination (R²) are used. 2 The prediction results of each multimodal model are quantitatively evaluated, and the calculation formula is as follows:
[0102]
[0103] Among them, y i Represents the actual value. Indicates the predicted value. This represents the average of the true values.
[0104] represents the average of the predicted values, and N represents the number of predicted values;
[0105] The experimental results of each multimodal model are shown in Table 3 below:
[0106] Table 3. Experimental Comparison Results of Different Multimodal Models
[0107] RMSE MAE <![CDATA[R 2 ]]> ResNet+AE+MLP 0.0918 0.0553 0.4844 ResNet+VAE+MLP 0.0403 0.0318 0.9010 ViT-PCA+AE+Polynomial 0.1581 0.1237 -1.2272 ViT-PCA+AE+MLP 0.0283 0.0216 0.9512 This invention 0.0073 0.0050 0.9967
[0108] Based on the experimental results in Table 3, replacing the appearance feature and shape feature extraction networks in the model of this invention with ResNet and AE, respectively, significantly degrades the performance of the multimodal model, i.e., R... 2 The R-value is only 0.4844; analyzing each feature extraction module individually, the choice of apparent feature extraction method has a significant impact on model performance. The ViT model using the self-attention mechanism has the highest R-value. 2 The metrics reached 0.9512 and 0.9967 respectively, while the R-value using the ResNet model was... 2 The differences are 0.4844 and 0.9010, respectively. This difference indicates that feature extraction via ViT can more effectively capture global appearance features, while the local convolutional properties of ResNet may limit its ability to model the complex texture of oyster appearances. Among shape feature extraction methods, VAE significantly outperforms AE. The R of this invention... 2 =0.9967, which is 4.5% higher than ViT-PCA+AE+MLP's 0.9512, indicating that VAE can better extract discriminative shape features through the regularization constraint of the latent space. The experimental results of MLP and multinomial regression show that MLP performs well in nonlinear relationship modeling. In summary, the present invention shows obvious advantages in various multimodal models and also verifies the effectiveness of the feature extraction network in the prediction model constructed in the present invention.
[0109] In this embodiment, the computer used in the experiment was configured as follows: CPU frequency of 4.89GHz, memory of 16GB, operating system of Windows 11 (64-bit); programming language of Python 3.8, using the Adam optimizer, integrated development environment of Anaconda3, and the experiment was completed based on the PyTorch framework.
[0110] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting oyster meat yield based on multimodal fusion, characterized in that, include: Raw images of oysters are acquired, and appearance image data and shape feature data are obtained through a segmentation network to construct a multimodal dataset, which includes the appearance image data and the shape feature data. For the appearance image data, a self-attention mechanism encoder is used to extract global image features to obtain appearance feature vectors; for the shape feature data, a variational autoencoder is used to map key features to obtain shape feature vectors. The appearance feature vector and the shape feature vector are fused using the Concat method to construct a fused feature vector. A multilayer perceptron regression prediction model is constructed, and the multilayer perceptron regression prediction model is trained with the fused feature vector as input and the meat yield as the target output. The effectiveness of the trained multilayer perceptron regression prediction model is validated using a test set. If the validation results meet the target requirements, the multilayer perceptron regression prediction model is saved.
2. The oyster meat yield prediction method based on multimodal fusion according to claim 1, characterized in that, The process of acquiring the appearance image data and the shape feature data includes: Collect raw oyster image data and preprocess the image data; The target region is segmented using a segmentation network model; The segmentation network model is trained to generate segmentation masks and extract target contours to obtain shape feature data; By using a mask to set the background area of the original image to black while preserving the target area, the apparent image data is obtained.
3. The oyster meat yield prediction method based on multimodal fusion according to claim 1, characterized in that, The self-attention mechanism captures different relationships and features of the input data through different self-attention heads. The attention weight formula is as follows: Where Attention(·) represents the vector multiplication operation, Q, K, and V represent the linear mappings of query, key, and value, respectively, T represents the matrix transpose, and d k It represents the mapping dimension, and softmax represents the activation function; To enhance the model's expressive power, a multi-head self-attention mechanism is used, which involves calculating multiple sets of Q, K, and V, and then concatenating the results. The formula is as follows: MultiHeadSelfattention(Z)=Concat(head1,head2,…,head h )W0; headi=Attention(ZW Qi ,ZW Ki ,ZW Vi ); Where W0 represents the final output mapping weights, and Z represents the input sequence; The output of the self-attention mechanism is further processed by a feedforward network, and the formula is as follows: Y′=Feed Forward(Y)=ReLU(TW F1 +b F1 )W F2 +b F2 ; Where Y represents the output of the feedforward network, Y represents the output of the self-attention mechanism, and ReLU represents the activation function; By introducing residual connectivity and layer normalization, the flow of information between layers and the stability of the model are increased, as expressed in the following formula: Y″ = LayerNorm(Y+Z); Here, LaterNorm represents the layer normalization operation.
4. The oyster meat yield prediction method based on multimodal fusion according to claims 1 and 3, characterized in that, For the self-attention mechanism encoder extracting global image features, when the feature dimension is much larger than the number of samples, principal component analysis is selected to reduce the dimensionality of the extracted global features. By retaining the principal components with the largest variance, the feature directions that contribute the most to the data distribution are selected. First, the original data is centered, i.e., using the formula: X centered =X-μ; Among them, X centered Let μ represent the centered data matrix, and μ represent the mean vector of each feature dimension. Through singular value pairs X centered The decomposition is performed, and the variance contribution rate is calculated based on the singular values. The expression is as follows: X centered =U∑V T ; Where U represents the left singular vector matrix, ∑ represents the diagonal extension matrix containing singular values, V represents the right singular vector matrix, V represents the transpose, and σ i Let be the i-th singular value, representing the variance intensity of the corresponding principal component direction. To complete the data dimensionality reduction, the principal components corresponding to the first k largest singular values are selected to form the projection matrix W. Finally, by multiplying the centered data matrix by the projection matrix, data dimensionality reduction is achieved, expressed as X. pca =X centered ·W.
5. The oyster meat yield prediction method based on multimodal fusion according to claim 1, characterized in that, The variational autoencoder obtains the shape feature vector by: inputting data into the encoder, which outputs two codes: one is the original code μ(μ1, μ2, μ3), and the other is the control noise code σ(σ1, σ2, σ3). The control noise code σ(σ1, σ2, σ3) is weighted with random noise e(e1, e2, e3). Finally, the original code and the weighted noise code are added together to obtain the latent vector z(z1, z2, z3) of the variational autoencoder model.
6. The oyster meat yield prediction method based on multimodal fusion according to claim 5, characterized in that, A loss function is used in constructing the variational autoencoder model. This loss function includes reconstruction error and KL divergence, and its formula is: L=D KL (q(z|x)||p(z))-E q(Z|X) [log p(x|z)]; Where q(z|x) represents the latent distribution generated by the encoder, z represents the latent variable, x represents the input data, p(z) represents the prior distribution of z, and p(x|z) represents the output of the latent variable z transformed back into the original data space.
7. The oyster meat yield prediction method based on multimodal fusion according to claim 1, characterized in that, The multilayer perceptron regression prediction model comprises an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. The number of neurons in the input layer is the same as the dimension of the concatenated feature vector. The number of neurons in the first, second, and third hidden layers is halved layer by layer to gradually compress redundant information and retain key features. The output layer has a dimension of 1 and directly predicts the meat yield. Batch normalization and ReLU activation functions are used between each layer of the multilayer perceptron regression prediction model. Dropout is set to 0.2 in the first and second hidden layers to suppress noise, and Dropout is set to 0.1 in the third hidden layer to retain effective features.
Citation Information
Cited By
Beef cattle meat percentage prediction method and system based on dynamic physiology and AI three-dimensional body size fusion
CN121580339A
Beef cattle meat yield prediction method and system based on dynamic physiological and ai three-dimensional body size fusion
CN121580339B