Pig intramuscular fat content in-vivo prediction method, system and equipment based on multi-modal deep neural network
By constructing a multimodal deep neural network and combining ultrasound images with auxiliary physiological information, the accuracy and stability problems of predicting intramuscular fat content in pigs in existing technologies have been solved. High-precision prediction under high IMF content and low-quality images has been achieved, which is applicable to live breeding of various pig breeds.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUN YAT SEN UNIV
- Filing Date
- 2026-03-12
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies lack high-precision and robust models for predicting intramuscular fat content (IMF) in pigs. In particular, their prediction performance is poor under conditions of large samples and low-quality images, making it difficult to meet the breeding requirements for pig breeds with high IMF content. Furthermore, the lack of a multimodal information fusion mechanism affects the accuracy and stability of predictions.
A multimodal deep neural network was constructed, combining ultrasound images and auxiliary physiological information. Through grayscale texture feature extraction and wavelet decomposition, EfficientNet-B0 was used as the backbone network for image feature extraction. Multi-scale features and structured data were integrated to achieve prediction of high IMF content and low-quality images.
It achieves high-precision and robust in vivo prediction of intramuscular fat content in pigs, improves the model's predictive ability under different pig breeds and image quality conditions, and has cross-breed applicability and stability.
Smart Images

Figure CN121837271A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of animal live detection and computer image data processing technology, specifically relating to a method, system and device for predicting intramuscular fat content in pigs based on a multimodal deep neural network. Background Technology
[0002] Intramuscular fat (IMF) refers to adipose tissue located between skeletal muscle fibers, including fat stored between muscle bundles and fibers within the muscle tissue structure. It is an important component of muscle tissue. IMF content, as a key characteristic of meat quality, determines the meat's antioxidant properties, tenderness, flavor, and juiciness. Generally, meat with higher IMF content is more tender and juicy, with a richer flavor, and IMF content is closely related to consumer acceptance. Traditional IMF determination requires post-slaughter sensory evaluation or chemical methods (Soxhlet extraction). The former is highly subjective, lacks standardized criteria, and has a large margin of error, while the latter is costly and inefficient. Neither method can be used for live animal selection, limiting the genetic selection of meat quality traits. The emergence and application of ultrasound technology offers a solution. Ultrasound technology was first used in human medicine, and later, due to its non-invasive, live animal testing advantages, it has been gradually extended to the assessment of live animal traits in livestock production and breeding. Compared to the main economic traits of livestock carcasses such as backfat thickness and eye muscle area, the research and application of ultrasound detection technology in the determination of IMF content started relatively late. Early research mainly focused on cattle, and ultrasound prediction of IMF content in pigs has gradually developed since then.
[0003] The existing methods for predicting IMF content in pigs using ultrasound mainly include the following: (1) Linear regression model: For example, Newcom et al. (2002) used 207 purebred Duroc pigs as the research object, collected their live ultrasound images, backfat thickness and other data and IMF data after slaughter (mean IMF=3.76%), screened variables through linear regression analysis, and constructed a prediction model containing ultrasound backfat thickness and 5 image parameters; and then verified it with data from 619 Duroc pigs and Yorkshire pigs, and concluded that real-time ultrasound image analysis can be used to predict the percentage of IMF in live pigs. Among them, the correlation coefficient between the predicted value and the measured value in the Duroc pig population reached 0.60, and the prediction effect was the best (Newcom DW, Baas TJ, Lampe JF. Prediction of intramuscularfat percentage in live swine using real-time ultrasound. J Anim Sci 2002;80:3046-52.). Peppmeier et al. (2023) further expanded the sample size and population representativeness, selecting 574 experimental pigs including Duroc pigs, purebred Large White pigs and commercial pigs, and integrated their B-ultrasound images and IMF measurements (mean IMF=2.21%) to construct an equation-based IMF prediction model with a model determination coefficient of 0.56 (Peppmeier ZC, Howard JT, Knauer MT, Leonard SM. Estimating backfat depth, loin depth, and intramuscular fat percentage from ultrasound images in swine. Animal 2023;17.). (2) Traditional machine learning model: Wu et al. (2025) collected real-time ultrasound images and IMF data (mean IMF=2.43%) of 336 pigs (Duroc × (Landrace × Large White) crossbred pigs and purebred Large White pigs), and extracted image feature parameters through computer image processing technology. A prediction model was built using three methods: multiple linear regression (MLR), support vector machine (SVM), and backpropagation artificial neural network (BPANN).The correlation coefficients in the validation and test sets ranged from 0.72 to 0.82 (Wu J, Yang YY, Yang W, et al. Non-destructive and efficient prediction of intramuscular fat in livepigs based on ultrasound images and machine learning. Comput Electron Agric2025;234.). (3) Deep learning model: Kvam and Kongsro (2017) first applied deep learning to predict IMF in pigs. Their study used 303 Duroc, Landrace, and Large White pigs as research subjects, collecting and cropping ultrasound images of specific body parts. A deep convolutional neural network (CNN) was used to construct the model. After training and validation, it was found that the model performed best on samples with low to medium IMF (<6%), with a correlation coefficient of 0.82 (Kvam J, Kongsro J. & IT In vivo & IT prediction of intramuscular fat using ultrasound and deep learning. Comput Electron Agric 2017;142:521-3.). Liu et al. (2024) collected ultrasound images of 945 commercial pigs (including Landrace × Large White crossbred pigs and Duroc × (Landrace × Large White) crossbred pigs) and their corresponding IMF data (mean). The correlation coefficient of the ResNet50 model trained on this dataset reached 0.68 on the test set, and an online tool, PIMFP, was developed (Liu Z, Du H, LaoFD, et al. PIMFP: an accurate tool for the prediction of intramuscular fat percentage in live pigs using ultrasound images based on deep learning. Comput Electron Agric 2024;217.).
[0004] However, existing technologies still have many shortcomings: (1) There is a lack of high-precision prediction models that have been validated by large-scale datasets: Most current studies are based on small samples (less than 1,000 heads) to build prediction models, resulting in insufficient model training and limited generalization ability. Although some methods have reported high correlation coefficients, their performance has not been fully validated in real-world scenarios with large samples, making it difficult to support actual breeding applications. (2) Existing technologies are mainly applicable to lean-type pigs with low IMF content, which is difficult to meet the breeding needs of high IMF: Existing models are mostly developed in lean-type pig herds with IMF content concentrated in 1%–4%, and their predictive ability for local pig breeds with high IMF is significantly reduced. They cannot accurately identify individuals with high IMF, which limits their application value in the genetic improvement of high-quality pork. (3) Existing technologies are highly sensitive to the quality of ultrasound images, which limits their practicality: Existing methods generally rely on high-quality ultrasound images acquired in a standard posture. In actual farm environments, due to differences in the technical level of operators, pig agitation, probe angle deviation, and other factors, blurry, low-contrast, or partially occluded images are often obtained. Existing models exhibit poor prediction stability and significantly increased errors when faced with such low-quality images, making it difficult to achieve large-scale, robust field deployment and promotion. (4) Existing technologies lack effective multimodal information fusion mechanisms: Existing technologies typically rely solely on a single ultrasound image modality for modeling, failing to effectively integrate multi-scale texture and structural features in the image, and also failing to deeply fuse key auxiliary physiological parameters (such as ultrasound backfat thickness) with image features. This limitation in information utilization restricts the model's ability to represent the complex biological mechanisms of IMF, becoming a bottleneck for further improving prediction accuracy. Summary of the Invention
[0005] This invention aims to solve the above-mentioned technical problems by constructing a large-sample ultrasound-IMF paired dataset of multiple pig breeds and designing a novel multimodal neural network architecture that effectively integrates multi-scale features of ultrasound images with auxiliary physiological information. This improves the robustness and prediction accuracy of the model under low-quality images and high IMF content samples, enabling high-precision, non-destructive, and generalizable prediction of IMF content in live animals, thus serving precision breeding and meat quality assessment.
[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for in vivo prediction of intramuscular fat content in pigs based on a multimodal deep neural network, comprising the following steps: S1. Acquire ultrasound images of pigs; S2, Image preprocessing; S3. Image ROI grayscale texture feature extraction, including: S3-1. Convert the image to a single-channel grayscale image; S3-2, Analyze the pixel proportion of key grayscale ranges; S3-3, Global Gray-Level Dispersion Evaluation: Calculate the standard deviation of pixel values in the entire gray-level image to measure the drastic degree of overall gray-level variation in the image; S3-4, Wavelet Vertical Detail Energy Analysis: Perform a single-layer two-dimensional discrete wavelet transform on the grayscale image, and use the Daubechies 1 wavelet basis to decompose the detail information in the horizontal, vertical and diagonal directions. Focus on extracting the detail sub-bands in the vertical direction and calculate the variance of their pixel values. S3-5, Multi-scale gray-level co-occurrence texture modeling: Based on the gray-level co-occurrence matrix theory, a texture co-occurrence relationship model of the image is constructed at different spatial distances; S3-6, Feature integration and structured output; S4. Input the data into the trained multimodal model for processing. The multimodal model uses a convolutional neural network as the backbone network for extracting image features, and splices and fuses the image features with structured numerical auxiliary information (auxiliary physiological parameters and texture features) in the feature dimension, and outputs the predicted value of IMF content through a fully connected layer.
[0007] Preferably, in step S1, the location for acquiring the ultrasound image is 5 centimeters away from the midline of the back on the left side of the pig's body, from the last four ribs.
[0008] Preferably, in step S2, the ROI region is selected by cropping to remove irrelevant interference information at the image edges.
[0009] Preferably, in step S3-2, the image grayscale value range of 0–255 is divided into multiple consecutive 10-level intervals, and the percentage of pixels in each interval is counted to the total number of pixels.
[0010] More preferably, in step S3-2, the proportions of the two adjacent intervals 20–30 and 30–40 are added together to form the "low-medium gray concentration" feature; the proportions of the intervals 100–110 and 110–120 are added together to form the "medium-high gray concentration" feature.
[0011] Preferably, in steps S3-5, homogeneity is calculated under adjacent pixels and horizontal direction conditions to characterize the consistency and uniformity of gray levels in local areas; contrast, dissimilarity, and entropy are calculated under the same horizontal direction conditions with a one-pixel interval to reflect the intensity of local gray level differences, degree of dissimilarity, and texture complexity, respectively.
[0012] Preferably, in steps S3-6, the various texture features are organized according to image samples, with each image corresponding to a record containing the sample name and multiple quantitative features. After summarizing the feature data of all samples, the data is saved as a CSV file in the form of a structured table.
[0013] Preferably, in step S4, the structured numerical auxiliary information includes: backfat thickness, eye muscle depth, IMF content measured by ultrasound, and texture statistics calculated from the original B-ultrasound image.
[0014] Preferably, in step S4, the training method of the multimodal model is as follows: the mean squared error is used as the loss function, the optimizer is Adam, the initial learning rate is set to 0.001, the batch size is 32, and the total training cycle is 200 rounds; in each training round, the model performs forward propagation, loss calculation, backpropagation, and parameter update on the training set; then, gradient-free inference is performed on the validation set, the prediction results and the true labels are collected, and the validation loss, root mean square error, and mean absolute error are calculated as performance evaluation indicators.
[0015] Preferably, in step S4, the CustomMultiModalDataset class is used to synchronously integrate image data and structured numerical auxiliary information, and align it with the target label. This class can automatically read image files and their associated metadata in a specified directory, convert the image to RGB format, and convert the structured numerical auxiliary information and IMF content labels into floating-point tensor forms suitable for use by deep learning frameworks, thereby providing standardized input for multimodal neural networks.
[0016] Preferably, in step S4, the construction of the multimodal model includes the following steps: S4-1. Select a lightweight and efficient convolutional neural network as the backbone network for image feature extraction, which is used to automatically extract high-level semantic features of images. S4-2. Embed the above feature extractor into a custom multimodal fusion model, which receives two types of inputs: image tensors and structured additional feature vectors; S4-3. Inside the model, the image embedding representation is first obtained through the feature extractor. Then, the embedding is concatenated with the additional feature vector along the feature dimension, and a single continuous value prediction result is output through the fully connected regression head. S4-4. Deploy the completed model to the GPU.
[0017] Preferably, in step S4, the iterative training of the multimodal model includes: for each training epoch, traversing all batches in the training data loader; for each batch of data, performing the following sub-steps: (a) Migrate images, additional features, and labels to a specified computing device; (b) Zero out the optimizer gradient; (c) Forward propagation: Input the image and additional features into the multimodal model to obtain the prediction output; (d) Calculate the MSE loss between the predicted value and the true label; (e) Backpropagation is used to calculate the gradient and the model parameters are updated through the optimizer; (f) Accumulate the total training loss for the current epoch; After each epoch, the average training loss is calculated and recorded.
[0018] Preferably, in step S4, the verification of the multimodal model includes: entering verification mode after each training epoch. (a) Disable gradient calculation; (b) Traverse the validation data loader, perform forward inference for each batch, obtain the predicted value and accumulate the validation loss; (c) Collect the prediction results and true labels for all validation samples; (d) Calculate regression performance metrics based on the collected results, including root mean square error and mean absolute error; (e) Record the average validation loss, RMSE, and MAE.
[0019] In step S4, the training of the multimodal model further includes: After each epoch is completed, the state dictionary of the current model is saved to the log directory in the format of "epoch_{n}.pth" for easy selection of subsequent models or resumption of training from breakpoints; All training / validation losses and evaluation metrics are organized by epoch and written to a CSV file; Plot the training loss and validation loss as a function of epochs and save the plot as a PNG image for intuitive analysis of model convergence and overfitting.
[0020] Secondly, the present invention provides a live prediction system for intramuscular fat content in pigs based on a multimodal deep neural network, comprising: A B-mode ultrasound machine is used to acquire B-mode images; The image preprocessing module is used to preprocess ultrasound images; The image ROI grayscale texture feature extraction module is used to systematically extract texture features from grayscale image ROIs. The multimodal neural network training module is used to process data input into the trained multimodal model. The results output module is used to output the predicted IMF content.
[0021] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the in vivo prediction method for intramuscular fat content in pigs as described in the present invention.
[0022] The core technology of this invention lies in constructing a high-precision, highly generalizable, and robust porcine intramuscular fat (IMF) in vivo prediction system, specifically including: Construction of a large-scale, multi-breed pig dataset: Establish an ultrasound-IMF paired database covering multiple pig breeds (such as lean commercial pigs, Duroc pigs, and rural black pigs) with a large sample size of 1,657 pigs.
[0023] MultiModal Model: A novel deep neural network architecture is proposed, employing EfficientNet-B0 as the backbone network for image feature extraction. This model achieves direct concatenation of ultrasound image features (from CNN) with structured supplementary data (ultrasound backfat thickness, ultrasound oculomotor depth, and various texture statistical features extracted from the image) along the feature dimension, forming a joint feature representation, which is then used for regression prediction through fully connected layers.
[0024] A systematic image texture feature extraction method: For B-ultrasound images, a multi-dimensional quantization method is provided that integrates gray-level distribution statistics (such as the pixel ratio of a specific gray-level range), frequency domain wavelet decomposition (such as wavelet vertical detail energy variance), and spatial gray-level co-occurrence relationship modeling (homogeneity, contrast, entropy, etc. of GLCM).
[0025] High accuracy and robustness: It achieves excellent prediction capabilities for samples with high IMF content (>4%) and low-quality ultrasound images, improving the practicality of the model. Attached Figure Description
[0026] Figure 1 The diagram illustrates ultrasound image acquisition and ROI selection. (A) Schematic diagram of ultrasound image detection; (B) Ultrasound image; (C) ROI region.
[0027] Figure 2 The process of extracting grayscale texture information is shown.
[0028] Figure 3 The multimodal model processing flow of the present invention is shown.
[0029] Figure 4 The IMF liveness prediction process is shown.
[0030] Figure 5 Partial prediction results are shown.
[0031] Figure 6 The linear fitting analysis of the prediction results is shown. Detailed Implementation
[0032] To facilitate understanding of the present invention, a more complete description will be given below with reference to specific embodiments. Preferred embodiments of the invention are shown in the accompanying drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0034] Unless otherwise specified, the experimental methods used in the following examples and comparative examples are conventional methods, and the materials and reagents used are commercially available unless otherwise specified.
[0035] This invention provides a method for predicting the in vivo IMF content of pigs based on a multimodal deep neural network, specifically including the following steps: 1. Data Collection and Preprocessing (1) Ultrasound image acquisition like Figure 1 As shown, the specific operation process for ultrasound data acquisition is as follows: First, the pig to be tested needs to be herded into a restraining cage, which should restrict the pig's movement as much as possible. The operator determines the measurement location by touching the ribs. The measurement location is located on the left side of the pig's body, 5 cm away from the midline of the back, from the last four ribs. Before the measurement, the measurement area needs to be cleaned; for pigs with thick hair, only shaving is necessary. After cleaning, apply sufficient coupling gel or vegetable oil to the measurement area to eliminate air interference between the probe and the skin; then, the operator holds the ultrasound probe vertically against the skin, and by slightly adjusting the angle and position of the probe, allows the ultrasound beam to penetrate the subcutaneous tissue, acquiring a clear longitudinal tomographic image in real time on the ultrasound screen. Figure 1 (A and B). At this point, the outline of the skin layer, backfat layer (including fat and connective tissue), and deep oculomotor muscles (longissimus dorsi) can be observed from top to bottom. The image is saved when the boundary between the oculomotor muscles and the ribs is visible. Multiple clear ultrasound images of each pig are acquired through multi-angle, multi-frame acquisition.
[0036] (2) Image preprocessing The ultrasound examination used was a Unicorn Vet portable veterinary ultrasound scanner (AdveEMA BV), with a frequency set to 3.5 MHz. The exported images were three-channel images of 940 pixels * 720 pixels, saved in PNG format. Before image analysis, Region of Interest (ROI) selection was performed by cropping to remove irrelevant interference information such as individual and parameter information from image edges. Figure 1 (C).
[0037] (3) Image ROI grayscale texture analysis This invention provides a systematic texture feature extraction method for grayscale images. By fusing grayscale distribution statistics, frequency domain wavelet decomposition, and spatial co-occurrence relationship modeling, it achieves multi-dimensional quantification of local structural characteristics of images. This method is applicable to IMF content liveness prediction, and the operation flowchart is as follows. Figure 2 As shown, the specific steps include: Step 1: Image Reading and Grayscale Conversion All image files are read sequentially from the specified folder and uniformly converted into single-channel grayscale images. These grayscale images serve as the basis for all subsequent texture analysis, ensuring that different feature calculations are based on a consistent data representation.
[0038] Step 2: Pixel Proportion Analysis in Key Gray-Scale Ranges The image grayscale value range (0–255) is divided into multiple consecutive 10-level intervals (e.g., 0–10, 10–20…250–256), and the percentage of pixels in each interval is calculated. Specifically, the percentages of the two adjacent intervals 20–30 and 30–40 are added together to form the "low-to-medium grayscale concentration" feature; similarly, the percentages of the intervals 100–110 and 110–120 are added together to form the "medium-to-high grayscale concentration" feature. These two features are used to reflect the energy distribution characteristics of the image within a specific grayscale range.
[0039] Step 3: Global Gray-Scale Discrepancy Assessment The standard deviation of pixel values in the entire grayscale image is calculated to measure the drasticness of overall grayscale changes. This metric reflects the image's contrast level and noise intensity, and is a fundamental statistical measure for describing the macroscopic visual characteristics of an image.
[0040] Step 4: Wavelet Vertical Detail Energy Analysis A single-layer two-dimensional discrete wavelet transform is performed on the grayscale image, and the Daubechies 1 wavelet basis is used to decompose the image into detail information in the horizontal, vertical, and diagonal directions. The vertical detail sub-bands are extracted with a focus on extraction, and the variance of their pixel values is calculated. This variance reflects the activity level of horizontal edges or stripe structures in the image and is sensitive to directional textures (such as fibers or cracks).
[0041] Step 5: Multi-scale grayscale co-occurrence texture modeling Based on the Gray-Level Co-occurrence Matrix (GLCM) theory, a texture co-occurrence relationship model of an image is constructed at different spatial distances: Homogeneity is calculated under the conditions of adjacent pixels (distance of 1) and horizontal direction (angle of 0 degrees) to characterize the consistency and uniformity of gray levels in local areas; At a distance of 2 pixels, and in the same horizontal direction, the three indices of contrast, dissimilarity, and entropy are calculated respectively, reflecting the intensity of local grayscale differences, the degree of dissimilarity, and the texture complexity.
[0042] By setting different distance parameters, this method can simultaneously capture texture structure features in both the microscopic neighborhood and a slightly larger area.
[0043] Step 6: Feature Integration and Structured Output The texture features described above are organized according to image samples, with each image corresponding to a record containing the sample name and multiple quantitative features. After summarizing the feature data of all samples, the data is saved as a CSV file in structured table format, which is convenient for subsequent use in machine learning model training, multimodal fusion analysis, or quality assessment systems.
[0044] (4) Meat sample collection and chemical determination of IMF content Each pig was slaughtered within 24 hours after the ultrasound examination. After slaughter, 100g of the longissimus dorsi muscle sample was collected from the third and fourth ribs from the bottom left side of each pig.
[0045] The Soxhlet extraction method was used to determine the IMF (intra-fat content) in pork: After removing visible connective tissue and fat from the pork sample, the sample was wrapped in a filter paper tube and placed in the extractor. Anhydrous diethyl ether was used as the solvent, and the extraction was carried out by heating and circulating for several hours. After recovering the diethyl ether, the receiving flask was dried, cooled, and weighed. The percentage of fat content was calculated by the difference in mass before and after extraction, and the IMF content was obtained. This measured value was used as the true value for model learning.
[0046] (5) Dataset construction and partitioning This invention acquired ultrasound images of 1657 pigs, including multi-dimensional information such as backfat thickness, eye muscle depth, and texture features. Each image was tagged with IMF content, which was precisely measured using chemical methods. All relevant data—such as ultrasound image file names, backfat thickness, eye muscle depth, image texture analysis results, and IMF content—were recorded in a single table for easy data management and access.
[0047] This technology uses the CustomMultiModalDataset class to synchronously integrate image data with structured numerical auxiliary information (including backfat thickness and oculomotor depth measured by ultrasound, as well as texture statistics calculated from the original B-ultrasound image) and align them with target labels. This class can automatically read image files and their associated metadata from a specified directory, convert the images to RGB format, and convert the structured numerical auxiliary information and IMF content labels into floating-point tensors suitable for deep learning frameworks, thus providing standardized input for multimodal neural networks.
[0048] To evaluate model performance and prevent overfitting, the entire dataset was divided into three independent subsets: a test set, a validation set, and a training set.
[0049] The specific division strategy is as follows: Test Set: 330 samples are randomly selected from all 1657 samples as the test set. The main purpose of the test set is to provide an unbiased estimate of the model's performance on unseen data after training, in order to evaluate the model's true predictive ability and generalization level.
[0050] Training and Validation Sets: The remaining 1327 samples were further divided into training and validation sets in an 8:2 ratio. 1062 samples were allocated to the training set for learning model parameters; the 265 samples formed the validation set, used during training to adjust hyperparameters, monitor model performance, and prevent overfitting.
[0051] 2. Construction and training of multimodal neural networks (1) Multimodal model construction ( Figure 3 ) The multimodal neural network architecture of this technology includes two core modules: an image feature extraction backbone network and a multi-source feature fusion adapter.
[0052] The image feature extraction backbone network employs a replaceable pre-trained convolutional neural network (EfficientNet-B0) to automatically learn deep semantic features from the input ultrasound images. The ultrasound images are uniformly resized to 224×224 pixels before input and converted into three-channel tensors to be compatible with standard visual models. The high-dimensional feature map output by the backbone network is globally flattened to form image feature vectors.
[0053] At the same time, the additional non-image data is organized into a fixed-length numerical vector.
[0054] In this embodiment, the additional data includes: backfat thickness and extraocular muscle depth measured by ultrasound, as well as texture statistics calculated from the original B-ultrasound image. This additional data is reshaped into a one-dimensional tensor before being fed into the model, and then concatenated with the image feature vector in the feature dimension to form a joint multimodal feature representation.
[0055] Subsequently, the concatenated joint features are input into a fully connected adaptation layer (i.e., a linear regression head), which maps the high-dimensional fused features into a single continuous output value for predicting IMF content.
[0056] (2) Model training During the model training phase, the model was deployed on a GPU, using mean squared error (MSELoss) as the loss function, Adam as the optimizer, with an initial learning rate of 0.001, a batch size of 32, and a total training duration of 200 epochs. In each training epoch, the model performed forward propagation, loss calculation, backpropagation, and parameter updates on the training set; subsequently, gradient-free inference was performed on the validation set, collecting prediction results and ground truth labels, and calculating validation loss, root mean squared error (RMSE), and mean absolute error (MAE) as performance evaluation metrics.
[0057] Step 1: Data Preprocessing and Loading The raw image data is standardized and preprocessed, including scaling the input image to a uniform size of 224×224 pixels and converting it to tensor format; Construct a custom multimodal dataset class (CustomMultiModalDataset), in which each sample contains a quadruple: (1) RGB image data; (2) structured additional feature vectors associated with the image (additional_data); (3) continuous target labels (labels); The training and validation sets are loaded from the specified paths (such as data / train and data / val) respectively, and are loaded in batches using PyTorch's DataLoader with a set batch size (batch_size=32). The training set is randomly shuffled (shuffle=True), while the validation set is loaded sequentially (shuffle=False).
[0058] Step 2: Multimodal Model Construction Lightweight and efficient convolutional neural networks (such as EfficientNet-B0) are selected as the backbone network for image feature extraction to automatically extract high-level semantic features of images. The aforementioned feature extractor is embedded into a custom multimodal fusion model (MultiModalModel), which receives two types of inputs: image tensors and structured additional feature vectors; Inside the model, the image embedding representation is first obtained through a feature extractor. Then, the embedding is concatenated with the additional feature vector along the feature dimension, and a single continuous value prediction result is output through a fully connected regression head. Deploy the completed model to the GPU.
[0059] Step 3: Initialize training configuration The loss function is defined as the mean squared error (MSELoss), which is suitable for regression tasks; The Adam optimizer was used to perform end-to-end optimization of the model parameters, with the initial learning rate set to 0.001. Set the total number of training epochs to 200, and store the model weights, evaluation metrics, and visualization results during the training process.
[0060] Step 4: Iterative Training and Validation (A) Training phase: For each training epoch, iterate through all batches in the training data loader; for each batch of data, perform the following sub-steps: (a) Migrate images, additional features, and labels to a specified computing device; (b) Zero out the optimizer gradient; (c) Forward propagation: Input the image and additional features into the multimodal model to obtain the prediction output; (d) Calculate the MSE loss between the predicted value and the true label (after unsqueeze(1) to expand the dimensions); (e) Backpropagation is used to calculate the gradient and the model parameters are updated through the optimizer; (f) Accumulate the total training loss for the current epoch.
[0061] After each epoch, the average training loss is calculated and recorded.
[0062] (B) Validation Phase: After each training epoch, the system enters validation mode. (a) Disable gradient computation (torch.no_grad()); (b) Traverse the validation data loader, perform forward inference for each batch, obtain the predicted value and accumulate the validation loss; (c) Collect the prediction results and true labels for all validation samples; (d) Calculate regression performance metrics based on the collected results, including root mean square error (RMSE) and mean absolute error (MAE). (e) Record the average validation loss, RMSE, and MAE.
[0063] Step 5: Model saving and result recording After each epoch, the current model's state dictionary (state_dict) is saved to the log directory in the format "epoch_{n}.pth" for easy model selection or resuming training from breakpoints. All training / validation losses and evaluation metrics are organized by epoch and written to a CSV file; Plot the training loss and validation loss as a function of epochs and save the plot as a PNG image for intuitive analysis of model convergence and overfitting.
[0064] Step 6: Training terminated After completing the preset 200 training epochs, the training process will automatically terminate, and the final output will include a complete training log, model weight sequence, and performance evaluation report.
[0065] 3. In vivo prediction of IMF content in pigs This invention provides a method for automatically predicting new samples based on a trained multimodal neural network model, which is suitable for intramuscular fat (IMF) content regression tasks that integrate ultrasound images and structured auxiliary features.
[0066] This prediction method ensures standardized input data, efficient model inference, and traceable results, and specifically includes the following steps: Step 1: Test Data Preparation The ultrasound image to be predicted, along with its corresponding texture information and auxiliary physiological information (back fat thickness, eye muscle depth), is organized according to a preset format. Step 2: Load the trained multimodal model The initialization process uses the same multimodal model architecture as the training phase, with the image feature extraction backbone network being EfficientNet-B0, followed by a multimodal fusion regression head. Load the pre-trained and saved model weight file from the specified path and map the weights to the current computing device (such as a GPU). Set the model to evaluation mode and disable training-specific mechanisms such as batch normalization and random deactivation to ensure a stable and consistent inference process.
[0067] Step 3: Batch Reasoning and Result Collection Use the data loader to iterate through the test set sample by sample with a batch size of 1 (supports any number of samples). For each sample, its image and auxiliary features are simultaneously fed into the model, forward propagation is performed, and a single continuous value of IMF content prediction result is obtained; Compared with the prior art, the technical solution of the present invention has the following advantages: (1) Significantly improved prediction accuracy: On the independent test set, the Pearson correlation coefficient of this model is R = 0.88 and MAPE = 15.76%, which is better than the existing technology; (2) Strong generalization ability: This technology has performed well in multiple pig breeds, including lean-type commercial pigs (Landrace × Landrace), Guangdong Small Spotted Pig, rural black pigs and Duroc × Guangdong Small Spotted Pig, proving that the model has cross-breed applicability; (3) Excellent robustness: The predictive ability for low-quality ultrasound images (such as blurry, low signal-to-noise ratio) and high IMF samples (>4%) is significantly better than existing models; (4) Innovation of multimodal fusion mechanism: The multimodal early fusion strategy model directly concatenates image features (from CNN) with additional data (ultrasound back fat thickness, texture features, etc.) at the feature level to achieve early fusion.
[0068] Application Example: Cross-breed IMF live prediction Subjects: 210 pigs, including 70 commercial white pigs (average IMF=2.43%), 70 Duroc pigs (average IMF=4.17%), and 70 rural black pigs and Laiwu black pigs (average IMF=7.01%).
[0069] The operation process is as follows Figure 4 As shown.
[0070] Data preparation: Collect and preprocess ultrasound images according to the above process, and extract back fat thickness, eye muscle depth and texture features.
[0071] Model prediction: Input the prepared data into a multimodal model loaded with pre-trained weights and set the model to evaluation mode.
[0072] Output: After performing forward propagation, continuous IMF content predictions are automatically output.
[0073] The experimental results are shown in Table 1. Figure 5 , Figure 6 As shown.
[0074] Table 1. Prediction Results of Pig Breeds
[0075] Depend on Figure 5 As can be seen, the IMF content predicted by the model of this invention is close to the measured results and can accurately reflect the fat content in meat samples.
[0076] Correlation analysis As shown in Table 1, in all test groups, the Pearson correlation coefficient between predicted and actual values was above 0.895, demonstrating a very strong linear correlation. Figure 6 The correlation coefficient was highest in the Duroc pig population, reaching 0.917, indicating that the model's predictive performance for the medium IMF content range is extremely accurate. In the Xiangxia Hei and Laiwu pig populations, which represent high IMF content, the correlation coefficient also reached 0.895, far exceeding the predictive ability of existing technologies for high IMF samples.
[0077] Error Analysis (MAPE) The MAPE value represents the average relative error between the predicted and actual values; a lower value indicates higher model accuracy. As shown in Table 1, the MAPE for the Duroc pig is the lowest, at only 6.74%, indicating that the model's prediction bias for this population is extremely small. The MAPE values for lean white pigs and rural black / Laiwu pigs are 11.83% and 11.06%, respectively, both at low levels, demonstrating the high reliability of the prediction results of this invention across different IMF content gradients.
[0078] On the other hand, the present invention also provides a live prediction system for intramuscular fat content in pigs based on a multimodal deep neural network, comprising: A B-mode ultrasound machine is used to acquire B-mode images; The image preprocessing module is used to preprocess ultrasound images; The image ROI grayscale texture feature extraction module is used to systematically extract texture features from grayscale image ROIs. The multimodal neural network training module is used to process data input into the trained multimodal model. The results output module is used to output the predicted IMF content.
[0079] On the other hand, the present invention also provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the in vivo prediction method for intramuscular fat content in pigs as described in the present invention.
[0080] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for in vivo prediction of intramuscular fat content in pigs based on multimodal deep neural networks, characterized in that, Includes the following steps: S1. Acquire ultrasound images of pigs; S2, Image preprocessing; S3. Image ROI grayscale texture feature extraction, including: S3-1. Convert the image to a single-channel grayscale image; S3-2, Analyze the pixel proportion of key grayscale ranges; S3-3, Global Gray-Level Dispersion Evaluation: Calculate the standard deviation of pixel values in the entire gray-level image to measure the drastic degree of overall gray-level variation in the image; S3-4, Wavelet Vertical Detail Energy Analysis: Perform a single-layer two-dimensional discrete wavelet transform on the grayscale image, and use the Daubechies 1 wavelet basis to decompose the detail information in the horizontal, vertical and diagonal directions. Focus on extracting the detail sub-bands in the vertical direction and calculate the variance of their pixel values. S3-5, Multi-scale gray-level co-occurrence texture modeling: Based on the gray-level co-occurrence matrix theory, a texture co-occurrence relationship model of the image is constructed at different spatial distances; S3-6, Feature integration and structured output; S4. Input the data into the trained multimodal model for processing. The multimodal model uses a convolutional neural network as the backbone network for extracting image features, and splices and fuses the image features with structured numerical auxiliary information in the feature dimension. The predicted value of IMF content is output through a fully connected layer.
2. The method according to claim 1, characterized in that, In step S2, the ROI region is selected by cropping to remove irrelevant interference information from the image edges.
3. The method according to claim 1, characterized in that, In step S3-2, the image grayscale value range of 0–255 is divided into multiple consecutive 10-level intervals, and the percentage of pixels in each interval is counted to the total number of pixels.
4. The method according to claim 1, characterized in that, In steps S3-5, homogeneity is calculated under adjacent pixels and horizontal direction conditions to characterize the consistency and uniformity of gray levels in local areas; contrast, dissimilarity, and entropy are calculated under the same horizontal direction conditions with a one-pixel interval to reflect the intensity of local gray level differences, degree of dissimilarity, and texture complexity, respectively.
5. The method according to claim 1, characterized in that, The structured numerical auxiliary information includes: backfat thickness and eye muscle depth measured by ultrasound, as well as texture statistics calculated from the original B-ultrasound image.
6. The method according to claim 1, characterized in that, The training method of the multimodal model is as follows: the mean squared error is used as the loss function, the optimizer is Adam, the initial learning rate is set to 0.001, the batch size is 8, and the total training cycle is 200 rounds; in each training round, the model performs forward propagation, loss calculation, backpropagation and parameter update on the training set; then gradient-free inference is performed on the validation set, the prediction results and the true labels are collected, and the validation loss, root mean square error and mean absolute error are calculated as performance evaluation indicators.
7. The method according to claim 1, characterized in that, In step S4, the CustomMultiModalDataset class is used to synchronously integrate image data and structured numerical auxiliary information and align it with the target label. This class automatically reads the image files and their associated metadata in the specified directory, converts the image into RGB format, and converts the structured numerical auxiliary information and IMF content label into floating-point tensor form suitable for use by deep learning frameworks, thereby providing standardized input for multimodal neural networks.
8. The method according to claim 1, characterized in that, In step S4, the construction of the multimodal model includes the following steps: S4-1. Select a lightweight and efficient convolutional neural network as the backbone network for image feature extraction, which is used to automatically extract high-level semantic features of images. S4-2. Embed the above feature extractor into a custom multimodal fusion model, which receives two types of inputs: image tensors and structured additional feature vectors; S4-3. Inside the model, the image embedding representation is first obtained through the feature extractor. Then, the embedding is concatenated with the additional feature vector along the feature dimension, and a single continuous value prediction result is output through the fully connected regression head. S4-4. Deploy the completed model to the GPU.
9. A live in vivo prediction system for intramuscular fat content in pigs based on a multimodal deep neural network, characterized in that... include: A B-mode ultrasound machine is used to acquire B-mode images; The image preprocessing module is used to preprocess ultrasound images; The image ROI grayscale texture feature extraction module is used to systematically extract texture features from grayscale image ROIs. The multimodal neural network training module is used to process data input into the trained multimodal model. The results output module is used to output the predicted IMF content.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for predicting intramuscular fat content in pigs as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Pig intramuscular fat content in-vivo prediction method, device and equipment and storage medium
CN116898482A
Sheep intramuscular fat content estimation method and system based on ultrasonic image
CN119205708A
Live pig intramuscular fat detection method and device based on multiple modes
CN119279632A
Multi-mode-based sheep non-injury intramuscular fat content detection method, system, equipment and medium
CN121622123A