Intelligent Detection Method for Quality Grading of Tuna Based on Multivariate Information Feature Fusion

Through multimodal feature fusion technology, fish body images, meat quality images and related text information are deeply integrated, solving the problems of low efficiency and poor accuracy of existing tuna quality detection methods, and achieving more comprehensive quality judgment and higher detection accuracy.

CN119832543BActive Publication Date: 2025-06-10CHINESE ACAD OF FISHERY SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411892566.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-06-10
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

The existing tuna quality detection methods rely on manual inspection or local image analysis, are inefficient and subjective, difficult to ensure consistency and accuracy, and fail to fully capture multi-dimensional information of fish body and meat quality.

Method used

The tuna quality grading intelligent detection method based on the fusion of multivariate information features is adopted. Through multimodal feature fusion technology, fish body images, fleshy images and related text information are deeply integrated, and the multi-head attention mechanism is used to capture the relationship between the various modal data to generate a more comprehensive feature representation.

Benefits of technology

It significantly improves the robustness and accuracy of tuna quality inspection, making the test results more comprehensive and reliable, and can provide accurate quality judgments in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832543B_ABST
    Figure CN119832543B_ABST
Patent Text Reader

Abstract

An intelligent detection method for tuna quality grading based on multi-source information feature fusion relates to the technical field of fish quality grading detection. Fish body images and meat images of several different tunas are obtained, and relevant text information of the tunas, including origin, variety, weight and body length, is collected to construct a data set; the images and relevant text information are preprocessed; multi-modal fusion feature vectors are generated through feature extraction and fusion; the multi-modal fusion feature vectors are input into a tuna quality level detection module; the deep learning model is trained, and subsequently, the quality grading detection of the tuna to be detected is carried out. By fully considering the fish body images, meat images and relevant text information, multi-modal feature fusion technology is used to generate multi-modal fusion features containing fish body, meat and relevant text information, and the quality grading detector is trained to improve the robustness and accuracy of the detection, making the detection results more comprehensive and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fish quality grading detection, and specifically to an intelligent detection method for tuna quality grading based on multi-source information feature fusion. Background Art

[0002] Tuna are oceanic migratory fish, and common species include yellowfin tuna, bigeye tuna, bluefin tuna, longfin tuna, and masu tuna. As a high-value marine economic product, the quality of tuna directly determines the market price and consumer choice. In scenarios such as processing plants and restaurants, accurate quality judgment is crucial for classifying tuna according to quality grades, thereby determining its uses and formulating corresponding processing procedures.

[0003] However, current quality evaluations often rely on manual inspections, requiring inspectors to have rich experience to judge the color, texture, and fat content of the fish meat, etc. This is inefficient and highly subjective, making it difficult to ensure consistency and accuracy. In addition, existing automatic detection methods for tuna quality mainly use tail cross-section images and fishing vessel information, combined with machine learning for auxiliary judgment, but this method has obvious limitations. First, it only analyzes the tail cross-section, ignoring the meat texture, color, and fat distribution of other parts. In fact, the meat quality information of different parts is crucial for evaluating the overall quality of tuna, and the integrity characteristics of the fish body should also be included in the evaluation scope to ensure the accuracy and comprehensiveness of the detection results. Second, there are significant differences in the meat quality characteristics of different varieties of tuna. For example, longfin tuna has a lighter meat color and lower oil content, making it suitable for canning, while bluefin tuna has a dark red meat color and extremely high oil content, and is commonly used for making sashimi, sushi, and grilled fish. Therefore, these variety differences also have an important impact on quality evaluation and market pricing. However, existing detection methods mainly rely on fishing vessel information and fail to combine multi-dimensional data such as origin, variety, weight, and body length, making it difficult for the model to capture the comprehensive characteristics of tuna quality and unable to meet the precise detection requirements for different application scenarios.

[0004] Generally speaking, the deficiencies caused by single information sources and local image analysis have become the bottleneck of current tuna quality detection. Therefore, there is an urgent need for a multi-modal information fusion solution that can flexibly analyze the overall image of the fish body and the meat images of any part, and combine multi-dimensional structured text data to provide technical support for more accurate evaluation of tuna quality. Summary of the Invention

[0005] To address the deficiencies in the background technology, the present invention provides an intelligent detection method for tuna quality grading based on multi-source information feature fusion. This method fully considers fish body images, meat images, and relevant text information, uses multi-modal feature fusion technology to generate multi-modal fusion features containing fish body, meat, and relevant text information, and trains through a quality grading detector to improve the robustness and accuracy of detection, making the detection results more comprehensive and reliable.

[0006] To achieve the above objectives, the present invention adopts the following technical solutions: An intelligent detection method for tuna quality grading based on multi-source information feature fusion includes four parts: a multi-source information collection module, a data preprocessing module, a feature extraction and fusion module, and a tuna quality level detection module. Among them, the feature extraction and fusion module includes a feature extraction model and a feature fusion model, and together with the network structure of the tuna quality level detection module, they form a deep learning model. The method includes the following steps:

[0007] Step 1: In the multi-source information collection module, obtain complete fish body images of several different tunas and meat images of any part, and at the same time collect relevant text information of the tuna, including origin, variety, weight, and body length. Then construct a data set, divide the quality of the tuna into five grades from high to low according to user needs, and add real grade labels to each tuna in the data set in the form of manual annotation;

[0008] Step 2: In the data preprocessing module, perform standard preprocessing on the fish body images and meat images, uniformly adjust all original images to the target size, and normalize the pixel values;

[0009] Step 3: In the data preprocessing module, preprocess the relevant text information into two categories. Among them, the weight and body length are normalized to obtain a normalized numerical feature vector, and the origin and variety are converted into dense vector representations using the embedding layer in the deep learning framework PyTorch to obtain discrete feature vectors after embedding processing. The normalized numerical feature vector and the discrete feature vector after embedding processing are concatenated to obtain the relevant text information feature F text ;

[0010] Step 4: In the feature extraction and fusion module, use the feature extraction model to input the fish body images and meat images preprocessed in Step 2 into a parallel convolutional neural network, and use several convolutional layers to extract local features of the fish body and meat respectively. The pooling layer is used to reduce the feature dimension and retain the significance of local features. Global average pooling is performed on the multi-dimensional feature maps output by the convolutional layer to generate two fixed-length feature vectors, including the feature vector of the fish body and the feature vector of the meat respectively representing the overall features of the fish body and the meat;

[0011] Step 5: In the feature extraction and fusion module, use the feature extraction model to input the relevant text information feature F in Step 3 text into the fully connected network to learn its latent features and output the extracted relevant text information feature

[0012] Step 6: In the feature extraction and fusion module, use the feature fusion model to adopt a multi-modal feature fusion strategy for the feature vector of the fish body obtained in Step 4 and the feature vector of the fish meat and the relevant text information feature obtained in Step 5 to perform integration, and use the multi-head attention mechanism for fusion to obtain the multi-modal fusion feature vector F containing the fish body image, fish meat image and relevant text information global ;

[0013] Step 7: In the tuna quality level detection module, input the multi-modal fusion feature vector F obtained in Step 6 global . The network structure of the tuna quality level detection module consists of a multi-layer perceptron and a Softmax layer, and the output is a probability distribution representing the probability that the tuna belongs to each of the five levels;

[0014] Step 8: Train the deep learning model, calculate the loss between the probability distribution output in Step 7 and the true level label. The true level label is given in the form of one-hot encoding, where the position corresponding to the actual quality level of the tuna is 1 and the other positions are 0. Then use the weighted cross-entropy loss function to evaluate the accuracy of the deep learning model, and adjust the parameters of the deep learning model using backpropagation according to the calculated loss value until the loss value no longer decreases, and obtain the trained deep learning model;

[0015] Step 9: By inputting the sample information of the tuna to be detected into the trained deep learning model obtained in Step 8, the quality grading detection of the tuna to be detected can be realized.

[0016] Furthermore, Step 2 specifically includes:

[0017] S2.1. Set the target size. If the ratio of the original image to the target size is inconsistent, for the original image smaller than the target size, fill it with white around it. For the original image larger than the target size, first scale it proportionally so that its long side is the same as the target size, and then perform central cropping to extract the middle pixel area with the same size as the target size;

[0018] S2.2. Use the standardization formula to adjust the pixel values of all images adjusted to the target size to have a mean of 0 and a standard deviation of 1, and normalize all image pixels to the range [0,1]. The calculation formula is as follows:

[0019]

[0020] Wherein, normalized pixel represents the normalized pixel, and original pixel represents the original pixel, μ is the mean value, and σ is the standard deviation.

[0021] Furthermore, the specific steps of Step 4 include:

[0022] S4.1. For the preprocessed fish body image I body , extract the local features of the fish body through a number of convolutional layers, including surface texture and damage. Suppose L b layers of convolutional layers are used, and the operation of each layer is:

[0023]

[0024] Wherein, is the feature map generated after the l-th layer of convolution of the fish body image, is the l-th layer of convolution operation of the fish body image, including the convolutional kernel weights and biases. At the first layer, ReLU() represents the activation function;

[0025] After each convolution operation, use a pooling layer to reduce the spatial dimension of the feature map while retaining the significance of the local features, which is expressed as:

[0026]

[0027] Wherein, is the feature map of the fish body image after pooling, and Pooling() represents pooling;

[0028] S4.2. For the preprocessed meat image I meat , extract the local features of the meat through a number of convolutional layers, including color and fat distribution. Suppose L m layers of convolutional layers are used, and the operation of each layer is:

[0029]

[0030]

[0031] Wherein, is the feature map generated after the l-th layer of convolution of the meat image, is the l-th layer of convolution operation of the meat image, including the convolutional kernel weights and biases. At the first layer,

[0032] After each convolution operation, a pooling layer is used to reduce the spatial dimension of the feature map while preserving the significance of local features, expressed as:

[0033]

[0034] In the formula, is the feature map after pooling of the meat image;

[0035] S4.3. Generate two fixed-length feature vectors through global average pooling, where:

[0036] For the feature maps of all fish body images Perform global average pooling to generate the feature vector of the fish body The formula is as follows:

[0037]

[0038] In the formula, H' b and W' b are the height and width of the feature map after the last pooling layer of the fish body image respectively;

[0039] For the feature maps of all meat images Perform global average pooling to generate the feature vector of the meat The formula is as follows:

[0040]

[0041] In the formula, H' m and W' m are the height and width of the feature map after the last pooling layer of the meat image respectively.

[0042] Furthermore, the specific steps of step six include:

[0043] S6.1. Project the extracted relevant text information features onto the same dimension as the image feature vector through a linear layer to obtain the relevant text information feature vector to achieve feature dimension alignment;

[0044] S6.2. Concatenate the feature vector of the fish body the feature vector of the meat and the relevant text information feature vector These three feature vectors into an input matrix F concat , and each row represents the feature vector of one modality. The multi-head attention mechanism is used to fuse the three feature vectors into a multi-modal fusion feature vector. The formula is as follows:

[0045]

[0046] Q = F concat W Q

[0047] K = F concat W K

[0048] V = F concat W V

[0049]

[0050] Head i = Attention(Q i , K i , V i )

[0051] MultiHead(Q, K, V) = Concat(Head 1 , Head 2 , Head 3 )W O

[0052] F fused (i, :) = MultiHead(Q, K, V)

[0053]

[0054] Wherein, Q, K, are the query, key, and value matrices respectively, W Q , W K , are the corresponding learnable linear transformation matrices respectively, Attention() represents the attention mechanism, softmax() normalizes each attention weight to represent it as a probability, d is the dimension of each feature vector, Head i represents the output of the i-th attention head in the multi-head attention mechanism, MultiHead() represents the final output of the multi-head attention, Concat represents concatenating the outputs of multiple heads column-wise, is the output linear transformation matrix, h is the number of heads of the multi-head attention, F fused (i, :) represents the fused feature of the i-th modality, is the multi-modal fused feature vector.

[0055] Compared with the prior art, the beneficial effects of the present invention are:

[0056] 1. The method of the present invention is no longer limited to the image input of a specific part. Whether it is the complete fish body image of tuna or the meat image of any part, it can effectively extract deep image features. Through the convolutional neural network, local information such as surface texture and damage can be extracted from the overall fish body image, and at the same time, microscopic information such as color and fat distribution can be extracted from the meat image. Therefore, even if only a local meat image of tuna is taken, the model can still give a relatively accurate quality assessment;

[0057] 2. The method of the present invention adopts a multi-modal feature fusion technology, which deeply fuses the fish body image, the meat image with its origin, variety, weight, and body length. Through the multi-head attention mechanism, it can effectively capture the mutual relationship between each modal data, make full use of the complementarity of multi-modal information, generate a more comprehensive feature representation, significantly improve the robustness and accuracy, and make the detection results more comprehensive and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is the flow chart of the method of the present invention;

[0059] Figure 2 is the structural diagram of the deep learning model in the present invention;

[0060] Figure 3 is the example diagram of the quality grading of tuna in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0062] Traditional tuna quality detection methods mainly rely on manual identification, which has problems of low efficiency and easy errors. At present, automatic identification methods use the tail fin cross-section image and fishing vessel information, and combine traditional machine learning models to analyze features. Although they can provide certain quality judgments, there are significant limitations. Specifically, the tail fin cross-section image can only provide the meat quality information of a local part of the fish body and cannot reflect the states of other parts of the fish body, such as the fish belly and fish back, thus affecting the accuracy of the overall judgment. In addition, the machine learning methods used rely on manually designed features. These feature extraction methods are relatively single and cannot capture the deep information in the image, especially complex image patterns and semantic information. Moreover, the way of manually designing features highly depends on personnel experience and is difficult to adapt to the diversity of tuna quality under different varieties and different fishing environments. Especially when solely relying on the tail fin cross-section image without providing variety information, the problem is more significant. The meat of different types of tuna shows different color characteristics. Without variety information, it is difficult to give accurate and consistent judgment results. In addition, there is a strong dependence on image data, especially on images of specific parts. In actual operation, if the image of the specific part is missing due to shooting conditions or other reasons, the detection performance of this method will significantly decline. This dependence on specific data makes the generalization ability of the model weak and it is difficult to cope with the data missing problem in practical applications.

[0063] In view of the deficiencies existing in the above-mentioned existing methods, such as Figures 1 to 2 shown, an intelligent detection method for tuna quality grading based on multi-source information feature fusion is proposed. The specific process is combined with Figure 1As shown in the figure, it includes four parts: a multi-source information collection module, a data preprocessing module, a feature extraction and fusion module, and a tuna quality level detection module. Among them, the feature extraction and fusion module contains a feature extraction model and a feature fusion model, and together with the network structure of the tuna quality level detection module, they form a deep learning model. In the multi-source information collection module, the system receives the complete fish body image of the tuna and the meat image of any part, and at the same time collects the relevant text information of the tuna, including data such as origin, variety, weight, and body length. These image data combinations provide multi-modal data. The fish body image and the meat image can capture the image features of the tuna, while the relevant text information provides background information and quantitative features. In the data preprocessing module, the system performs data preprocessing on the collected multi-modal data. In the feature extraction and fusion module, the system uses a convolutional neural network to process the fish body image and the meat image, extracts the key image features of the fish body and the meat, and at the same time uses a fully connected network to process the relevant text information, extracts numerical features, and then uses a multi-modal feature fusion strategy for integration, and forms a multi-modal fusion feature vector containing multi-level information through a multi-head attention mechanism. In the tuna quality level detection module, the multi-modal fusion feature vector is input into the quality grading detector for training, and is subsequently used for the detection of tuna quality grading, which specifically includes the following steps:

[0064] Step 1: In the multi-source information collection module, use a high-resolution camera to obtain several complete fish body images of different tunas and meat images of any part, and at the same time collect the relevant text information of the tuna, including origin, variety, weight, and body length, to achieve multi-source information collection. Furthermore, construct a dataset for tuna quality grading detection. According to user requirements, divide the quality of the tuna into five grades from high to low: super excellent grade (S grade), excellent grade (A grade), good grade (B grade), qualified grade (C grade), and unqualified grade (D grade). Add real grade labels to each tuna in the dataset in the form of manual annotation;

[0065] Step 2: In the data preprocessing module, perform standard preprocessing on the fish body image and the meat image. Uniformly adjust all the original images to the target size and normalize the pixel values to ensure the consistency of the subsequent input data distribution. Specifically:

[0066] S2.1. Set the target size, for example, 224×224 pixels. If the ratio of the original image to the target size is inconsistent, for the original image smaller than the target size, fill it with white around it. For the original image larger than the target size, first scale it proportionally so that its long side is the same as the target size, and then perform central cropping to extract the middle pixel area that is the same as the target size, so as to ensure that the adjusted fish body image and meat image can both maintain their important features and unify the input size for subsequent processing;

[0067] S2.2. Adjust all image pixel values after being adjusted to the target size to have a mean of 0 and a standard deviation of 1, and normalize all images to the range [0, 1] using the following calculation formula:

[0068]

[0069] In the formula, normalized pixel represents the normalized pixel, and original pixel represents the original pixel, μ is the mean, and σ is the standard deviation;

[0070] Step 3: In the data preprocessing module, preprocess the relevant text information in two categories, where:

[0071] Body weight and body length belong to numerical information and are normalized to obtain a normalized numerical feature vector to ensure that different features are modeled at the same scale. The calculation formula is as follows:

[0072]

[0073] In the formula, X' is the normalized value, X is the original value, X min is the minimum value in the dataset, and X max is the maximum value in the dataset.

[0074] Place of origin and variety belong to discrete information. Using the embedding layer in the deep learning framework PyTorch, each feature is converted into a dense vector representation to obtain an embedded discrete feature vector.

[0075] Then, concatenate the normalized numerical feature vector and the embedded discrete feature vector to obtain the relevant text information feature F text ;

[0076] Step 4: In the feature extraction and fusion module, as shown in Figure 2 , use the feature extraction model to input the preprocessed fish body image and meat image in step 2 into a parallel convolutional neural network. Use several convolutional layers to extract local features such as the surface texture and damage of the fish body and local features such as the color and fat distribution of the meat. Reduce the feature dimension through the pooling layer and retain the significance of the local features. Perform global average pooling on the multi-dimensional feature maps output by the convolutional layer to generate two fixed-length feature vectors, including the feature vector of the fish body and the feature vector of the meat represent the overall features of the fish body and the meat respectively. Specifically:

[0077] S4.1. For the preprocessed fish body image Ibody Extract the local features of the fish body, including surface texture and damage, etc., through several convolutional layers. Suppose L b convolutional layers are used, and the operation of each layer is as follows:

[0078]

[0079]

[0080] In the formula, is the feature map generated after the l-th layer of convolution of the fish body image, is the l-th layer of convolution operation of the fish body image, including the convolutional kernel weights and biases. At the first layer, i.e., the fish body image itself, and ReLU() represents the activation function.

[0081] After each convolution operation, use the pooling layer to reduce the spatial dimension of the feature map while retaining the significance of the local features, which is expressed as:

[0082]

[0083] In the formula, is the feature map of the fish body image after pooling, and Pooling() represents pooling;

[0084] S4.2. For the preprocessed meat image I meat , extract the local features of the meat, including color and fat distribution, etc., through several convolutional layers. Suppose L m convolutional layers are used, and the operation of each layer is as follows:

[0085]

[0086]

[0087] In the formula, is the feature map generated after the l-th layer of convolution of the meat image, is the l-th layer of convolution operation of the meat image, including the convolutional kernel weights and biases. At the first layer, i.e., the meat image itself.

[0088] Similarly, after each convolution operation, use the pooling layer to reduce the spatial dimension of the feature map while retaining the significance of the local features, which is expressed as:

[0089]

[0090] In the formula, is the feature map of the meat image after pooling;

[0091] S4.3. Generate two fixed-length feature vectors through global average pooling, where:

[0092] For the feature maps of all fish body images Perform global average pooling to generate the feature vector of the fish body The formula is as follows:

[0093]

[0094] In the formula, H' b and W' b are the height and width of the feature map after the last layer of pooling of the fish body image respectively. For the feature maps of all meat images Perform global average pooling to generate the feature vector of the meat The formula is as follows:

[0095]

[0096] In the formula, H' m and W' m are the height and width of the feature map after the last layer of pooling of the meat image respectively;

[0097] Step Five: In the feature extraction and fusion module, as shown in Figure 2 , use the feature extraction model to input the relevant text information feature F text in Step Three into the fully connected network to learn its latent features and output the extracted relevant text information feature

[0098] Step Six: In the feature extraction and fusion module, as shown in Figure 2 , use the feature fusion model to adopt a multi-modal feature fusion strategy for the feature vector of the fish body and the feature vector of the meat obtained in Step Four and the relevant text information feature obtained in Step Five for integration, and use the multi-head attention mechanism for fusion to obtain the multi-modal fusion feature vector F global containing the fish body image, meat image and relevant text information. Specifically:

[0099] S6.1. Project the extracted relevant text information feature onto the same dimension as the image feature vector through a linear layer to obtain the relevant text information feature vector to achieve feature dimension alignment;

[0100] S6.2. Concatenate the feature vector of the fish body , the feature vector of the meat and the relevant text information feature vector into an input matrix F concat, each row represents the feature vector of a modality. To capture the deep interaction relationships between different modalities, a multi-head attention mechanism is adopted to fuse the three feature vectors into a multi-modal fusion feature vector. The multi-head attention mechanism can adaptively allocate attention weights to each modality, thereby capturing the mutual influence between modalities. The formula is as follows:

[0101]

[0102] Q = F concat W Q

[0103] K = F concat W K

[0104] V = F concat W V

[0105]

[0106] Head i = Attention(Q i , K i , V i )

[0107] MultiHead(Q, K, V) = Concat(Head 1 , Head 2 , Head 3 )W O

[0108] F fused (i, :) = MultiHead(Q, K, V)

[0109]

[0110] In the formula, Q, K, are the query, key, and value matrices respectively, W Q , W K , are the corresponding learnable linear transformation matrices respectively, Attention() represents the attention mechanism, softmax() normalizes each attention weight to represent it as a probability, d is the dimension of each feature vector, Head i represents the output of the i-th attention head in the multi-head attention mechanism, MultiHead() represents the final output of the multi-head attention, Concat represents concatenating the outputs of multiple heads column-wise, is the output linear transformation matrix, h is the number of heads of the multi-head attention, F fused(i, :) represents the fusion features of the i-th modality (fish body, meat quality, or related text information), which is a multi-modal fusion feature vector;

[0111] Step 7: In the tuna quality level detection module, input the multi-modal fusion feature vector F obtained in Step 6 global . The network structure of the tuna quality level detection module consists of a multi-layer perceptron and a Softmax layer. First, the multi-layer perceptron captures the complex relationships between the inputs, and the Softmax layer is used to convert the output into a probability distribution, representing the probability that the tuna belongs to each of the five grades (S, A, B, C, D);

[0112] Step 8: Train the deep learning model. To evaluate the performance of the deep learning model, it is necessary to calculate the loss between the probability distribution output in Step 7 and the true grade label. The true grade label is given in the form of one-hot encoding, where the position corresponding to the actual quality grade of the tuna is 1, and the other positions are 0. Then, a weighted cross-entropy loss function is used to evaluate the accuracy of the deep learning model. According to

[0113] the calculated loss value is used to adjust the parameters of the deep learning model by backpropagation until the loss value no longer decreases, and the trained deep learning model is obtained;

[0114] Step 9: By inputting the sample information of the tuna to be detected into the trained deep learning model obtained in Step 8, the quality grading detection of the tuna to be detected can be achieved.

[0115] The present invention learns the quality grading features of tuna based on a deep learning model of multi-modal feature fusion. Among them: the image feature extraction network adopts different convolutional layer structures to extract their respective local features in parallel for the different features of the fish body image and meat quality image of the tuna. The fish body image focuses on appearance features such as surface texture and damage, while the meat quality image extracts internal features such as color and fat distribution; the relevant text information feature extraction network is used to learn the latent features of the relevant text information and generate a vector representation that matches the image features; the feature fusion network uses the multi-head attention mechanism to capture the complex interaction relationships between the fish body image, meat quality image, and relevant text information, so as to be able to adaptively focus on the features that are most important for quality judgment and fully fuse the information of each modality; the detector network is used to output the quality grading result of the tuna.

[0116] Different from the existing methods, the method of the present invention comprehensively utilizes fish body images, meat images, and relevant text information to achieve a more accurate assessment of the quality of tuna. Compared with the existing methods that solely rely on the tail cross-section images and fishing vessel information, the method of the present invention can capture more dimensional information of the fish body and meat, thus providing a more comprehensive basis for quality judgment. In terms of feature extraction, the present invention adopts different convolutional network structures for fish body images and meat images to adapt to the uniqueness of their respective features. The extraction of fish body images focuses on appearance features such as surface texture and damage, while meat images focus on internal features such as color and fat distribution. This parallel processing not only improves the accuracy of feature extraction but also enhances the robustness. In addition, a multi-head attention mechanism is introduced to capture the complex relationships between fish body, meat, and relevant text information, identify the importance of different features for tuna quality judgment, and further improve the accuracy of the assessment. This innovative method effectively overcomes the limitations of the existing methods based on tail cross-section images, makes the quality grading more comprehensive and accurate, improves the intelligent level of the system, and can better meet the market's demand for tuna quality detection.

[0117] Embodiment

[0118] Combine Figure 3 As shown, in this embodiment, a high-resolution camera is used to obtain the fish body image and meat image of tuna, and relevant text information is collected. The obtained images are adjusted to a unified size of 224×224 pixels, and the pixels are normalized. The collected data information is normalized, and the origin and variety information are converted into vector representations through an embedding layer. The origin, variety, weight, and body length data are concatenated to obtain the relevant text information features. A convolutional neural network is used to extract the local features of the fish body image and meat image, and a multi-head attention mechanism is adopted to capture the deep interaction information between the images and relevant text information. The fused multi-modal fusion feature vector is sent into a quality grading detector for detection, and the quality grading result of the tuna sample is obtained as the extra-superior grade (S grade).

[0119] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent conditions of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

[0120] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An intelligent detection method for tuna quality grading based on multivariate information feature fusion, characterized by: The method comprises four parts: a multivariate information collection module, a data preprocessing module, a feature extraction and fusion module, and a tuna quality grade detection module. The feature extraction and fusion module comprises a feature extraction model and a feature fusion model, and together with the network structure of the tuna quality grade detection module, constitutes a deep learning model. The method comprises the following steps: Step 1: In the multivariate information collection module, obtain the complete body images of several different tunas and the meat images of any parts, and collect the relevant text information of the tuna, including the origin, species, weight and length, and then build a data set. According to user needs, the quality of tuna is divided into five levels from high to low, and the real grade label is added to each tuna in the data set through manual annotation; Step 2: In the data preprocessing module, the fish body images and meat quality images are standardized and preprocessed, all the original images are uniformly adjusted to the target size, and the pixel values ​​are normalized; Step 3: In the data preprocessing module, the relevant text information is divided into two categories for preprocessing. Among them, weight and body length are normalized to obtain normalized numerical feature vectors. Origin and variety are converted into dense vector representation using the embedding layer in the deep learning framework PyTorch to obtain embedded discrete feature vectors. The normalized numerical feature vectors and embedded discrete feature vectors are concatenated to obtain the relevant text information feature F. text ; Step 4: In the feature extraction and fusion module, the fish image and meat image preprocessed in step 2 are input into the parallel convolutional neural network using the feature extraction model. Several convolutional layers are used to extract the local features of the fish body and meat quality, respectively. The feature dimension is reduced through the pooling layer while retaining the significance of the local features. The multi-dimensional feature map output by the convolutional layer is globally averaged and pooled to generate two fixed-length feature vectors, including the feature vector of the fish body. and the feature vector of meat quality They represent the overall characteristics of the fish body and the overall characteristics of the meat respectively; Step 5: In the feature extraction and fusion module, the feature extraction model is used to extract the relevant text information features F in step 3. text Input into the fully connected network to learn its potential features, and output the extracted relevant text information features Step 6: In the feature extraction and fusion module, the feature fusion model is used to adopt a multimodal feature fusion strategy to extract the feature vector of the fish body obtained in step 4. and the feature vector of meat quality Related text information features obtained in step 5 The multi-head attention mechanism is used to integrate and fuse to obtain a multimodal fusion feature vector F containing fish body images, meat quality images and related text information. global ; Step 7: In the tuna quality level detection module, input the multimodal fusion feature vector F obtained in step 6 global ,The network structure of the tuna quality grade detection module is composed of a ,multilayer perceptron and a Softmax layer, and the output is a probability ,distribution representing the probability of the tuna belonging to each of the five grades; Step 8: Train the deep learning model and calculate the loss between the probability distribution output in step 7 and the true grade label. The true grade label is given in the form of one-hot encoding, where the position corresponding to the grade of the actual quality of tuna is 1 and the other positions are 0. Then, the weighted cross entropy loss function is used to evaluate the accuracy of the deep learning model. According to the calculated loss value, the parameters of the deep learning model are adjusted using back propagation until the loss value no longer decreases, thereby obtaining the trained deep learning model. Step nine: By inputting the sample information of the tuna to be tested into the trained deep learning model obtained in step eight, the quality grading test of the tuna to be tested can be achieved.

2. The intelligent detection method for tuna quality grading based on multivariate information feature fusion according to claim 1 is characterized in that: The step 2 specifically includes: S2.

1. Set the target size. If the original image is not proportional to the target size, fill the original image with white if it is smaller than the target size. If the original image is larger than the target size, scale it proportionally to make its long side consistent with the target size, and then perform center cropping to extract the pixel area in the middle that is consistent with the target size. S2.

2. Use the normalization formula to adjust the mean of all image pixels after they are resized to the target size to 0 and the standard deviation to 1, and normalize all image pixels to the range of [0,1]. The calculation formula is as follows: In the formula, normalized pixel Represents normalized pixels, original pixel represents the original pixel, μ is the mean, and σ is the standard deviation.

3. The intelligent detection method for tuna quality grading based on multivariate information feature fusion according to claim 1, characterized in that: The step 4 specifically includes: S4.

1. For the preprocessed fish image I body , extract the local features of the fish body, including surface texture and damage, through several convolutional layers. Suppose L b The operation of each convolutional layer is: In the formula, is the feature map generated after the l-th layer convolution of the fish image, is the lth layer convolution operation of the fish image, including the convolution kernel weight and bias. In the first layer, ReLU() represents the activation function; After each convolution operation, a pooling layer is used to reduce the spatial dimension of the feature map while retaining the saliency of local features, expressed as: In the formula, It is the feature map after the fish image is pooled, and Pooling() represents pooling; S4.

2. Preprocessed meat image I meat , extract the local features of meat quality, including color and fat distribution, through several convolutional layers. Suppose L m The operation of each convolutional layer is: In the formula, is the feature map generated after the l-th layer convolution of the meat image, This is the convolution operation of the meat image layer l, including the convolution kernel weight and bias. In the first layer, After each convolution operation, a pooling layer is used to reduce the spatial dimension of the feature map while retaining the saliency of local features, expressed as: In the formula, It is the feature map after meat image pooling; S4.

3. Generate two fixed-length feature vectors through global average pooling, where: Feature map for all fish images Perform global average pooling to generate the feature vector of the fish body The formula is as follows: Where H' b and W' b They are the height and width of the feature map after the last layer of pooling of the fish image; Feature map for all meat images Perform global average pooling to generate the feature vector of meat quality The formula is as follows: Where H' m and W' m They are respectively the height and width of the feature map after the last layer of pooling of the meat image.

4. The intelligent detection method for tuna quality grading based on multivariate information feature fusion according to claim 1, characterized in that: The step six specifically includes: S6.

1. The extracted relevant text information features are transformed into Project to the same dimension as the image feature vector to obtain the relevant text information feature vector Achieve feature dimension alignment; S6.

2. The feature vector of the fish body Characteristic vector of meat quality and related text information feature vector These three eigenvectors are concatenated into an input matrix F concat , each row represents a feature vector of a modality, and a multi-head attention mechanism is used to fuse the three feature vectors into a multi-modal fusion feature vector. The formula is as follows: Q=F concat W Q K=F concat W K V=F concat W V Head i =Attention(Q i ,K i ,V i ) MultiHead(Q,K,V)=Concat(Head1,Head2,Head3)W O F fused (i,:)=MultiHead(Q,K,V) In the formula, are query, key, and value matrices respectively, are the corresponding learnable linear transformation matrices, Attention() represents the attention mechanism, softmax() normalizes each attention weight to represent its probability, d is the dimension of each feature vector, Head i Represents the output of the i-th attention head in the multi-head attention mechanism, MultiHead() represents the final output of the multi-head attention, and Concat represents concatenating the outputs of multiple heads by column. is the linear transformation matrix of the output, h is the number of heads of multi-head attention, F fused (i,:) represents the fusion feature of the i-th modality, is the multimodal fusion feature vector.

Citation Information

Patent Citations

  • Artwork classification method and system based on multi-modal fusion

    CN114170460A

  • Joint attention-based cross-modal deep hash retrieval method and system and medium

    CN115203442A