Method for predicting intramuscular fat in living body and related product
The problem of measuring intramuscular fat content in vivo was solved by collecting live images through ultrasound equipment and using improved YOLOv5 and Transformer models to identify and predict intramuscular fat content in vivo, and support for pork quality improvement and breeding work.
Patent Information
- Application Number
- CN202510415983.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art cannot accurately measure the intramuscular fat content in the live state of pigs, resulting in hindering breeding work.
The ultrasonic device was used to collect live images, and the image quality was identified through the trained YOLOv5 model, and the training Transformer model was input to predict the intramuscular fat content after passing the qualification, combining image coding and population coding to improve prediction accuracy.
It has achieved accurate prediction of the fat content in living muscles, supported the accurate screening of breeding work, and improved pork quality and industrial development.
Smart Images

Figure CN120241129A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of in-vivo detection technology. More specifically, this disclosure relates to a method, a processing device, and a computer-readable storage medium for predicting intramuscular fat in vivo. Background Art
[0002] Pork occupies a significant share in China's meat consumption, accounting for more than 50% of the global total output and consumption. Moreover, its industrial chain is extremely extensive, covering links such as live pig breeding, slaughtering and processing, cold chain transportation, and terminal sales, containing a huge market scale. The quality of pork largely determines its market price, and the intramuscular fat content (IMF) is a key indicator for measuring the quality of pork. Phospholipids in fat will degrade into volatile substances such as aldehydes and ketones during the heating process, and these substances endow pork with a unique flavor. Generally, the higher the fat content, the stronger the flavor of pork. For example, pork with a high content of "marbling" can reach more than twice the price of ordinary pork. Therefore, increasing the intramuscular fat content of pigs to optimize the quality of pork and meet consumers' demand for high-quality pork is of great significance for promoting the high-quality development of the pork industry.
[0003] Currently, the main means to improve the quality of pork and increase the intramuscular fat content is genetic breeding optimization, that is, selecting pigs with a high intramuscular fat content as breeding pigs to reproduce offspring. However, at present, the determination of pork fat content must be carried out after the pigs are slaughtered, which brings great obstacles to the breeding work. Therefore, there is an urgent need in the industry for a new technology that can accurately determine the fat content in the live state of pigs.
[0004] In view of this, there is an urgent need to provide a method, a processing device, and a computer-readable storage medium for predicting intramuscular fat in vivo, so as to accurately predict the intramuscular fat content of live bodies such as pigs. Summary of the Invention
[0005] In order to solve at least one or more of the above-mentioned technical problems, this disclosure proposes a method, a processing device, and a computer-readable storage medium for predicting intramuscular fat in vivo in multiple aspects.
[0006] In a first aspect, an embodiment of this disclosure proposes a method for predicting intramuscular fat in vivo, including: collecting a target image of a live body using an ultrasonic device; inputting the target image into a first model to determine the quality of the target image; if the quality of the target image is qualified, inputting the target image into a second model to predict the intramuscular fat content of the live body.
[0007] In some embodiments, before inputting the target image into the second model to predict the intramuscular fat content of the living body, it further includes: globally denoising the target image; performing histogram equalization on the globally denoised target image.
[0008] In some embodiments, collecting the target image of the living body by using an ultrasonic device includes: controlling the probe of the ultrasonic device to move between the 3rd and 4th ribs from the bottom of the lower back of the living body, 5 to 7 cm away from the dorsal midline, so as to collect the target image.
[0009] In some embodiments, the first model is an improved Yolov5 model, and the activation function in the improved Yolov5 model is the GELU activation function.
[0010] In some embodiments, the second model is an improved Transformer model, and the improved Transformer model includes an image encoder and a decoder; inputting the target image into the second model to predict the intramuscular fat content of the living body includes: the image encoder performing image encoding on the target image to obtain an image encoding vector; the decoder predicting the intramuscular fat content of the living body according to the image encoding vector.
[0011] In some embodiments, the improved Transformer model further includes a population encoder;
[0012] In some embodiments, before inputting the target image into the second model to predict the intramuscular fat content of the living body, it includes: obtaining the population information of the living body; inputting the target image into the second model to predict the intramuscular fat content of the living body includes: the image encoder performing image encoding on the target image to obtain an image encoding vector; the population encoder performing population encoding on the population information to obtain a population encoding vector; the decoder predicting the intramuscular fat content of the living body according to the image encoding vector and the population encoding vector.
[0013] In some embodiments, before inputting the target image into the first model to determine the quality of the target image, it includes: collecting an image data set collected by using an ultrasonic device; annotating the image data set according to a preset standard to divide each image in the image data set into two categories: qualified or unqualified, where the preset standard includes clear backfat and eye muscle, clear ribs in the field of view and requiring more than 4 ribs; inputting the image data set into the first model for training to obtain the trained first model.
[0014] In some embodiments, before inputting the target image into the second model to predict the intramuscular fat content of the living body, the following steps are included: collecting an image dataset acquired by an ultrasound device; screening out a target image dataset with qualified quality from the image dataset; annotating the intramuscular fat content of each image in the target image dataset; and inputting the target image dataset annotated with the intramuscular fat content into the second model for training to obtain the trained second model.
[0015] In a second aspect, an embodiment of the present disclosure provides a processing device, including: a processor configured to execute program instructions; and a memory configured to store program instructions, which, when loaded and executed by the processor, cause the processor to execute the method described in the first aspect and any of its embodiments above.
[0016] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium storing program instructions, which, when loaded and executed by a processor, cause the processor to execute the method described in the first aspect and any of its embodiments above.
[0017] Through the method for predicting intramuscular fat of a living body provided as above, it successfully acquires a target image containing intramuscular fat feature information of the living body by using an ultrasound device, and determines the quality of the acquired target image through a first model. Only when the quality of the target image is qualified, it is input into the second model to predict the intramuscular fat content of the living body. Through this solution, accurate prediction of the intramuscular fat content of the living body can be achieved, which helps to accurately screen breeding pigs and support the breeding work. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0019] Figure 1 An exemplary flowchart of a method for predicting intramuscular fat of a living body showing some embodiments of the present disclosure;
[0020] Figure 2 An exemplary flowchart of a method for predicting intramuscular fat of a living body showing some embodiments of the present disclosure;
[0021] Figure 3 An exemplary network structure diagram of a first model showing some embodiments of the present disclosure;
[0022] Figure 4 An exemplary network framework diagram of a second model showing some embodiments of the present disclosure;
[0023] Figure 5 Schematic diagram of an image acquired by an ultrasound device showing some embodiments of the present disclosure;
[0024] Figure 6 Exemplary structural block diagram of a processing device showing some embodiments of the present disclosure. Detailed implementation manners
[0025] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0026] It should be understood that the terms "comprising" and "including" used in the specification and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0027] It should also be understood that the terms used in the specification of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. As used in the specification and claims of the present disclosure, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the specification and claims of the present disclosure refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0028] As used in this specification and the claims, the term "if" may be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" may be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.
[0029] The following will describe in detail the specific implementation manners of the present disclosure with reference to the accompanying drawings.
[0030] Exemplary application scenarios
[0031] Pork occupies an important position in China's meat consumption, and its production and consumption account for more than 50% of the world's total. From the perspective of the industrial chain, the pork industry runs through multiple links such as pig breeding, slaughtering and processing, cold chain transportation and terminal sales, and contains huge market potential. In recent years, relying on technological upgrades and variety improvement, the pork industry is steadily moving towards intensification and standardization.
[0032] The quality of pork plays a decisive role in its market price, and the intramuscular fat content is one of the key indicators for measuring pork quality. During the heating process, the phospholipids in pork fat will degrade into volatile substances such as aldehydes and ketones, which give pork a unique aroma. Generally speaking, the higher the fat content, the stronger the aroma of pork. For example, pork with a high content of "marbling" usually costs more than twice that of ordinary pork. Therefore, increasing the intramuscular fat content of pigs to optimize pork quality and meet consumer demand for high-quality pork is of great significance to promoting the high-quality development of the pork industry.
[0033] At present, improving pork quality and increasing intramuscular fat content mainly rely on genetic breeding optimization, that is, selecting pigs with high intramuscular fat content as breeding pigs. However, the fat content of pork can only be measured after the pigs are slaughtered, which greatly hinders the advancement of breeding work. Therefore, a new technology is urgently needed to accurately measure the fat content of pigs in the live state.
[0034] Exemplary Application Scenarios
[0035] In view of this, the disclosed embodiment provides a method for predicting intramuscular fat in living bodies, which successfully collects a target image containing intramuscular fat feature information of a living body by using an ultrasonic device, and determines the quality of the collected target image by a first model. Only when the target image quality is qualified will it be input into a second model to predict the intramuscular fat content of the living body. This solution can achieve accurate prediction of the intramuscular fat content of a living body, thereby facilitating accurate screening of breeding pigs and supporting breeding work.
[0036] Figure 1 An exemplary flow chart of a method 100 for predicting intramuscular fat in vivo according to some embodiments of the present disclosure is shown. Figure 1 As shown, the method 100 includes steps S101 to S103.
[0037] First, at step S101, an ultrasonic device is used to collect a target image of a living body. It can be understood that to predict the intramuscular fat content of a living body, it is necessary to use an ultrasonic device to collect a target image in a detection area of the living body, and the detection area is a common or concentrated distribution area of intramuscular fat in the living body, so that the collected target image contains as much characteristic information of intramuscular fat as possible.
[0038] In some embodiments, the ultrasound device includes a probe. During image acquisition, a coupling agent is first applied to the surface of the detection area of the living body. The function of the coupling agent is to exclude the air between the probe and the skin, enabling ultrasonic waves to better penetrate into the body, reducing the reflection and attenuation of ultrasonic energy, and thus improving the image quality. Then, by controlling the probe of the ultrasound device to slowly move on the surface of the detection area of the living body in a vertically light-pressing manner until a target image with a clear boundary between the backfat and the muscle is acquired. It should be noted that due to the different reflection and absorption characteristics of the backfat and the muscle to ultrasonic waves, they are presented as a bright layer and a dark layer respectively in the target image. Therefore, during image acquisition, the operator can judge the acquisition effect by closely observing the imaging conditions of the bright layer and the dark layer, as well as whether the boundary between the two is clear, until a target image with a clear boundary between the backfat and the muscle is acquired.
[0039] It can be understood that during the image acquisition process, parameters such as the gain and gray scale of the ultrasound device can be continuously adjusted to acquire a target image with a clear boundary between the backfat and the muscle.
[0040] In some embodiments, there may be body hair on the surface of the common or concentrated distribution areas of intramuscular fat in some living bodies. Therefore, to avoid affecting the acquisition of the target image, the body hair in the detection area of the living body can be removed first, and then the above-mentioned step S101 can be executed.
[0041] For example, taking pigs as an example, the detection area is between the 3rd and 4th ribs from the bottom of the lower back, 5 to 7 cm away from the dorsal midline. Therefore, the body hair between the 3rd and 4th ribs from the bottom of the lower back of the pig, 5 to 7 cm away from the dorsal midline, can be removed, and after removal, a coupling agent is applied to ensure that the probe can be in close contact with the pig's skin. Then, control the probe of the ultrasound device to slowly move between the 3rd and 4th ribs from the bottom of the lower back of the pig, 5 to 7 cm away from the dorsal midline, in a vertically light-pressing manner to acquire a target image with a clear boundary between the backfat and the muscle.
[0042] In some embodiments, the ultrasound device can also be equipped with relevant measurement tools and software so that after the target image is acquired, the vertical distance from the outer layer of the backfat to the muscle layer can be further marked using the built-in measurement tools or software of the ultrasound device, thereby assisting in the evaluation of the intramuscular fat content.
[0043] Although the target image is collected using ultrasound equipment, its collection quality is controlled by the operator, which leads to large individual differences in image quality. Different operators have different proficiency levels, preferences for setting ultrasound equipment parameters, and accuracy in locating the collection site. These factors will significantly affect the clarity of the final collected image, the recognition of the boundary between back fat and muscle, and the integrity of the intramuscular fat feature information in the image, thereby affecting the subsequent determination of the intramuscular fat content of the living body. To solve this defect, the following step S102 is executed next.
[0044] In step S102, the target image is input into the first model to determine the quality of the target image. In the disclosed embodiment, the first model is a trained model used to identify the quality of the input image. After the target image is acquired, it is input into the first model, and the first model identifies the quality of the target image and outputs the quality result of the target image. The quality result indicates whether the target image is qualified. The quality of the target image can be accurately identified by the first model, thereby avoiding the individual differences caused by the operator's technique, experience and other factors when the target image is manually acquired, and ensuring that the acquired target image is stable and reliable.
[0045] Next, a specific embodiment of training the first model is given to improve the quality detection capability of the first model. Training the first model should be performed before the above step S102.
[0046] Specifically, first, it is necessary to collect an image data set collected by ultrasonic equipment. The image data set covers all species of living things and is collected by different operators to ensure the richness of the data. Then the image data set is annotated according to the preset standard to divide each image in the image data set into two categories: qualified or unqualified. The preset standard can be that the back fat and eye muscles are clear, the ribs in the field of view are clear, and the number is required to be more than 4. Requiring more than 4 ribs is used to ensure that the area of all images remains consistent, so as to avoid the image reflecting the situation of different parts due to the deviation of the collection position. It can be understood that images that meet the preset standard are marked as qualified, and images that do not meet the preset standard are marked as unqualified. Figure 5 The schematic diagram of images collected by the ultrasound device of some embodiments of the present disclosure is shown, and both images a and b are target images collected from the third to fourth ribs from the lower back of the pig, 5 to 7 cm from the dorsal midline, where image a is an unqualified image and image b is a qualified image. Finally, the image data set is input into the first model for training, so that the trained first model can accurately identify whether the quality of the image is qualified.
[0047] In some embodiments, the aforementioned labeled image dataset is also manually labeled. Therefore, to avoid errors caused by manual labeling, experienced personnel can be selected to label the image dataset, thereby improving the accuracy of image labeling.
[0048] Finally, in step S103, if the quality of the target image is qualified, the target image is input into the second model to predict the intramuscular fat content of the living body. In the embodiments of the present disclosure, the second model is also a trained model. By analyzing the target image with qualified quality through the second model, the specific value of the intramuscular fat content of the living body can be accurately predicted. It is not difficult to understand that the intramuscular fat content refers to the amount of fat contained in muscle tissue.
[0049] It can be understood that if the quality of the target image is unqualified, step S103 is not executed, but instead, the above-mentioned step S101 is jumped to to re-collect the image until a target image with qualified quality is collected.
[0050] Next, a specific implementation method for training the second model is given to improve the prediction ability of the second model. It can be understood that the training of the second model should be executed before the above-mentioned step S103.
[0051] Specifically, first collect an image dataset collected by an ultrasonic device. This image dataset covers all breeds of living bodies and is collected by different operators, thus ensuring the richness of the data. Further, select a target image dataset with qualified quality from this image dataset. As an example, the image dataset can be input into the trained first model to determine the quality of each image in the image dataset. Then, a target image dataset with qualified quality is selected. After that, further label the intramuscular fat content of each image in the target image dataset. Finally, input the target image dataset labeled with the intramuscular fat content into the second model for training, so that the trained second model can accurately predict the intramuscular fat content of the input image.
[0052] In some embodiments, a target image with qualified quality of the living body can be obtained first. After slaughtering the living body, measure the intramuscular fat content in the detection area and label the target image with this intramuscular fat content.
[0053] The above combination Figure 1 describes method 100 for predicting intramuscular fat of living bodies in some embodiments of the present disclosure. It successfully collects a target image containing intramuscular fat feature information of a living body by using an ultrasonic device, and determines the quality of the collected target image through the first model. Only when the quality of the target image is qualified, it is input into the second model to predict the intramuscular fat content of the living body. Through this solution, accurate prediction of the intramuscular fat content of living bodies can be achieved, which helps to accurately screen breeding pigs and support breeding work.
[0054] However, it can be understood that Figure 1 the method shown is exemplary rather than restrictive, and those skilled in the art can make flexible adjustments as needed. For example, the first model can be embedded in an ultrasonic device so that the ultrasonic device can immediately identify the image quality using the first model during the image acquisition process until a target image with qualified quality is acquired. Based on such a setting, the acquisition rate of the target image with qualified quality can be accelerated.
[0055] Figure 3 An exemplary network structure diagram of the first model showing some embodiments of the present disclosure. As Figure 3 shown, the first model is a Yolov5 (You Only Look Once version 5) model, which consists of three parts: Backbone (backbone network), Neck (neck network), and Head (detection head). For the specific details of each component, reference can be made to Figure 3 the network structure shown, which will not be elaborated in detail here.
[0056] It can be understood that the YOLOv5 model is an end-to-end single-stage object detection model. It innovatively transforms the object detection task into a regression problem and, through a simple and efficient single-stage network, achieves accurate object localization and classification in one step. Specifically, in terms of training optimization, the YOLOv5 model adopts the Mosaic data augmentation technique. This technique randomly stitches and combines four images, greatly enriching the diversity of training data. It not only improves the model's adaptability to different scenarios and objects but also significantly enhances the training efficiency and accelerates the model's convergence speed. At the same time, combined with the strategies of adaptive anchor box calculation and adaptive image scaling, intelligent preprocessing of the input image is carried out. Adaptive anchor box calculation can automatically adjust the size and ratio of the anchor box according to the characteristics of the dataset, enabling the model to better adapt to objects of different sizes and shapes. Adaptive image scaling can flexibly adjust the size of the input image while ensuring detection accuracy, reducing the computational load and improving the model's running speed. From the perspective of the network structure, the Backbone part of YOLOv5 is based on the powerful CSPDarknet53 (Cross Stage Partial Darknet 53) backbone network, and a unique Focus module is introduced. The Focus module enhances the feature extraction ability while reducing the computational load through slicing operations on the input image, enabling it to more effectively capture key information in the image. The Neck part adopts a structure that combines FPN (Feature Pyramid Network) and PAN (Path Aggregation Network). By fusing shallow and deep layer features, the model's detection accuracy for small objects is significantly improved, effectively making up for the deficiencies of single-stage detection models in small object detection. The Head part is responsible for outputting multi-scale prediction results, and then the non-maximum suppression algorithm is used to screen the prediction boxes, finally obtaining accurate detection boxes to achieve precise object detection.
[0057] With the above advanced design concepts and technical means, the YOLOv5 model demonstrates excellent high real-time performance and outstanding accuracy.
[0058] Furthermore, as Figure 3 shown, the activation function of Yolov5 is the SiLU (Sigmoid Linear Unit) activation function. The Sigmoid operation of this activation function SiLU may introduce numerical instability problems at low precision. Therefore, in some embodiments, Yolov5 is improved by specifically changing the original activation function SiLU to the GELU (Gaussian Error Linear Unit) activation function.
[0059] The mathematical form of the GELU activation function is GELU(x) = xΦ(x), where Φ(x) is the cumulative distribution function of the standard normal distribution. This form is relatively smooth and can transform inputs in different value ranges more naturally when processing data, avoiding problems such as discontinuity or gradient explosion / vanishing that some activation functions may have. Therefore, the GELU activation function has more advantages when combined with mixed-precision training and can bring better stability and faster optimization speed to the training of the first model.
[0060] The above combination Figure 3 has comprehensively elaborated on the network structure of the first model in some embodiments of this disclosure. The first model is an improved version based on the Yolov5 model. The key improvement lies in replacing the SiLU activation function used in the original Yolov5 model with the GELU activation function. This change is significant. Due to its unique mathematical properties, the GELU activation function shows good adaptability when combined with mixed-precision training, making the improved Yolov5 model more stable during training and significantly improving the optimization speed. This optimization not only helps the first model learn image features more efficiently but also reduces training fluctuations to a certain extent, accelerates model convergence, and thus improves the performance of the first model in quality identification of target images.
[0061] However, it can be understood that the first model described in the embodiments of this disclosure is exemplary rather than restrictive, and those skilled in the art can make flexible adjustments according to needs.
[0062] In some embodiments, the second model is a Transformer model, which includes an image encoder and a decoder. The image encoder consists of multiple stacked self-attention layers and a feed-forward neural network, and is used to extract the global features of the input target image and generate an image encoding vector. The decoder then predicts the intramuscular fat content based on the image encoding vector output by the image encoder through a masked self-attention mechanism, which can prevent information leakage and avoid the model cheating by "peeking" at future information, thereby improving the generalization ability of the model.
[0063] It can be understood that the image encoder of the Transformer model is a self-attention mechanism, which is stacked by multiple identical image encoder layers, usually 6-12 layers. Each layer contains two core sub-layers: (1) The multi-head self-attention layer captures the global dependencies of elements within the sequence. (2) The feed-forward neural network enhances the feature expression ability through non-linear transformation. Before entering the image encoder layer, positional encoding is also required to capture the positional relationships between inputs, that is, by using sine / cosine functions or learnable parameters to add positional information to sequence elements to make up for the defect that the self-attention mechanism is insensitive to order. After being processed by multiple layers of image encoders, a high-dimensional context vector is output, which contains the global semantics and structural information of the sequence and provides input for the decoder. The decoder is the core component of the model to generate the target sequence. It outputs the results step by step in an autoregressive manner. Similar to the image encoder, it is also stacked by multiple identical encoder-decoder layers. The difference is that its components are: (1) The masked multi-head self-attention layer: only allows the current position to attend to previous positions, aiming to prevent leakage of future information. (2) The encoder-decoder attention layer: uses the Key-Value pairs output by the encoder to establish the association between the input and the target sequence.
[0064] In some embodiments, referring to Figure 4 , Figure 4 FIG. shows an exemplary network framework diagram of a second model according to some embodiments of the present disclosure. As Figure 4 shown, to improve the prediction performance, the Transformer model is improved. The specific improvement point is that the image encoder is changed to a VIT (Vision Transformer Encoder) encoder. The ViT encoder divides the input image into fixed-size image patches (e.g., a grid of 16×16 pixels). After each patch is flattened, it is converted into an embedding vector (i.e., the image encoding vector) through linear projection, simulating the input form of word sequences in natural language processing, directly establishing the global dependencies between image patches, which is superior to the local convolution operation of convolutional neural networks. After pre-training on large datasets, ViT achieves or exceeds the performance of traditional convolutional neural networks in tasks such as ImageNet (Image Network) and has multi-modal adaptability.
[0065] Based on this, when performing the previous step S103, the specific execution process of the second model is as follows: The VIT encoder performs image encoding on the target image to obtain the image encoding vector. Then the decoder predicts the intramuscular fat content of the living body according to the image encoding vector.
[0066] In some embodiments, such as Figure 4As shown, to further improve the prediction performance, a population encoder is additionally added to perform population encoding on the population information of the live body to obtain a population encoding vector. And the output of the population encoder is connected in series with the decoder, so that the decoder can predict the intramuscular fat content of the live body based on the population encoding vector output by the population encoder and the image encoding vector output by the image encoder. Furthermore, to enable population encoding, before performing the previous step S103, the population information of the live body needs to be obtained. Furthermore, when performing the previous step S103, the specific execution process of the second model is as follows: The VIT image encoder performs image encoding on the target image to obtain an image encoding vector, and the population encoding performs population encoding based on the input population information to obtain a population encoding vector. This population information can be the breed of the live body. Then the decoder predicts the intramuscular fat content of the live body based on the image encoding vector and the population encoding vector.
[0067] In summary, in the disclosed embodiment, specific adjustments are made to the Transformer model: only the image encoder is replaced with a VIT image encoder, and at the same time, a population encoder connected to the decoder is added. It should be emphasized that except for these two encoders, the other network structures of the Transformer model remain unchanged. That is to say, the overall network architecture of the second model in the disclosed embodiment is basically the same as that of the Transformer model, and the difference lies only in the specific image encoder (replaced by the VIT image encoder from the encoder of the original model) and the added population encoder.
[0068] In some embodiments, the population encoder encodes the input population information using a query mechanism.
[0069] As an example, the population encoder initializes a set of parameters with dimensions NxC. N represents all the types of pigs. For example, when there are only three types, namely Duroc, Landrace, and Yorkshire, N = 3. C represents the dimension, which is the same as the dimension of the vector after Vit encoding to ensure more convenient decoding later. NxC is denoted as the population encoding vector V. That is, when a pig corresponds to a certain breed, the corresponding population encoding vector is retrieved through a look-up table. For example, its dimension is: 1xC, and then 1xC is output to the decoder.
[0070] Furthermore, in some embodiments, since the VIT image encoder is used, the model has multi-modal adaptability. Therefore, the decoder can use the cross-attention mechanism to fuse the image encoding vector and the population encoding vector. It can be understood that cross-attention is a variant of the attention mechanism, which allows one sequence (query sequence, which is the population encoding vector in this case) to dynamically focus on another sequence (key-value sequence, which is the image encoding vector after Vit encoding), thereby establishing cross-modal information interaction and fusing the information of both.
[0071] It should be noted that the population encoder is not a mandatory option. If the second model does not receive the population information of the living body, then the image encoder encodes the target image to obtain an image encoding vector, while the population encoder does not perform encoding, and its decoder can directly predict the intramuscular fat content based on the image encoding vector. If the population information of the living body is received, the image encoder encodes the target image to obtain an image encoding vector, and the population encoder encodes the population information to obtain a population encoding vector. The decoder predicts the intramuscular fat content based on the image encoding vector and the population encoding vector.
[0072] The above combination Figure 4 Specifically describes the network structure of the second model of some embodiments of the present disclosure. By improving the Transformer model, the second model is obtained. The improvement lies in replacing the image encoder with a VIT image encoder to improve the performance of the second model and make it have multimodal adaptability, thereby improving the prediction accuracy of the second model. Based on the multimodal adaptability of the second model, in some embodiments, a population encoder connected to the decoder is additionally added to perform population encoding according to the input population information of the living body to obtain a population encoding vector. And the decoder can predict a more accurate intramuscular fat content by fusing the image encoding vector and the population encoding vector.
[0073] However, it can be understood that the second model described in the embodiments of the present disclosure is exemplary rather than restrictive, and those skilled in the art can make flexible adjustments according to needs.
[0074] Figure 2 An exemplary flowchart of a method 200 for predicting intramuscular fat of a living body according to some embodiments of the present disclosure is shown. As Figure 2 shown, the method 200 includes step S201 and step S201, and the method 200 can be executed before step S103 above.
[0075] In step S201, global noise reduction is performed on the target image. The target image refers to a target image with qualified quality. In some embodiments, global noise reduction of the target image can be achieved by performing a convolution operation on the target image.
[0076] For example, to achieve effective noise reduction, a 3x3 convolutional kernel can be used to perform a convolution operation on the target image to achieve global noise reduction. Specifically, for each pixel point in the target image, the 3x3 convolutional kernel will operate on all pixels within its neighborhood range, calculate the gray average value of these neighborhood pixels. Then, the obtained gray average value is used to replace the value of the current pixel point. Through this pixel-by-pixel and meticulous traversal method, global noise reduction of the target image can be achieved, thereby effectively improving the quality of the image and reducing noise interference.
[0077] In the embodiments of the present disclosure, no specific limitation is imposed on the size of the convolutional kernel. This is because the choice of the convolutional kernel size will significantly affect the effect of image denoising, and there are different trade - off relationships. A larger - sized convolutional kernel may be smoother but more blurred, while a smaller convolutional kernel retains more details. Those skilled in the art can flexibly select different - sized convolutional kernels to perform global denoising on the target image according to actual needs. For example, in other embodiments, a 5x5 convolutional kernel is selected to perform a convolutional operation on the target image to achieve global denoising of the target image.
[0078] In step S202, histogram equalization is performed on the globally denoised target image. Specifically, the execution process of histogram equalization is as follows: First, count the number of pixels of each gray level in the target image to obtain a gray - level histogram. Then, calculate the cumulative distribution function (CDF) of each gray level according to the gray - level histogram. The cumulative distribution function represents the proportion of pixels less than or equal to a certain gray level in the image. Next, normalize the cumulative distribution function so that its value range is between 0 and 255. Finally, map the gray - level value of each pixel in the target image according to the normalized cumulative distribution function to obtain the target image after histogram equalization. By adjusting the distribution of the image gray levels, the gray levels originally concentrated in a narrow area are extended to the entire dynamic range, thereby achieving the purpose of enhancing the image contrast.
[0079] In some embodiments, this operation can be completed by using the histogram function in Opencv. OpenCV (Open Source Computer Vision Library), namely the open - source computer vision library, is a powerful open - source library widely used in the field of computer vision. The histogram equalization processing of the target image can be automatically realized by calling the histogram function provided by this library.
[0080] The above combination Figure 2 has further described the method for predicting intramuscular fat in some embodiments of the present disclosure. After obtaining a target image with qualified quality, it further performs global denoising on the target image and histogram equalization on the globally denoised target image, thereby further improving the quality of the target image so that the second model can accurately measure the intramuscular fat content of the living body.
[0081] However, it can be understood that Figure 2 the method shown is exemplary rather than restrictive, and those skilled in the art can make flexible adjustments according to needs.
[0082] To implement the method steps described in the foregoing of the present disclosure in terms of software and hardware, embodiments of the present disclosure also provide a processing device, and this processing device can be the processing device as Figure 6 shown.Figure 6 shows an exemplary structural block diagram of the processing device 60 according to an embodiment of the present disclosure. As Figure 6 shown, the processing device 60 of the present disclosure may include a processor 610 and a memory 620. Among them, an executable program is stored on the memory 620, and the processor 610 can load and execute the executable program, so that the processing device 60 implements any of the method steps described above.
[0083] In an example scenario, the processor 610 can be used to control the memory 620. Further, the processor 610 may be a central processing unit (CPU), an application processor (AP), etc. integrated in the processing device 60; and the memory 620, as the hardware for implementing the storage function, may be a read-only memory (ROM), a dynamic RAM (DRAM), etc.
[0084] The embodiment of the present disclosure also provides a computer-readable storage medium, in which program instructions are stored. When the program instructions are executed by the processor of the processing device, the processor is enabled to execute the method steps described in any embodiment of the present disclosure.
[0085] In the embodiment of the present disclosure, a computer program product is also provided, including a computer program or instruction. When the computer program or instruction is executed by the processor, the method described in any embodiment of the present disclosure is implemented.
[0086] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art can think of many changes, alterations, and alternative ways without departing from the spirit and scope of the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed in practicing the present disclosure. The appended claims are intended to define the scope of protection of the present disclosure and thus cover equivalents or alternatives within the scope of these claims.
Claims
1. A method for predicting intramuscular fat in vivo, characterized in that, Comprising: Collecting a target image of a living body using an ultrasonic device; Inputting the target image into a first model to determine the quality of the target image; If the quality of the target image is qualified, inputting the target image into a second model to predict the intramuscular fat content of the living body.
2. The method according to claim 1, characterized in that, Before the step of inputting the target image into the second model to predict the intramuscular fat content of the living body, it further includes: Performing global noise reduction on the target image; Performing histogram equalization on the target image after global noise reduction.
3. The method according to claim 1, wherein The step of collecting a target image of a living body using an ultrasonic device includes: Controlling the probe of the ultrasonic device to move between the 3rd and 4th ribs from the bottom of the back of the living body, 5 to 7 centimeters away from the mid-dorsal line, to collect the target image.
4. The method according to claim 1, wherein The first model is an improved Yolov5 model, and the activation function in the improved Yolov5 model is the GELU activation function.
5. The method according to claim 1, wherein The second model is an improved Transformer model, and the improved Transformer model includes an image encoder and a decoder; The step of inputting the target image into the second model to predict the intramuscular fat content of the living body includes: The image encoder performs image encoding on the target image to obtain an image encoding vector; The decoder predicts the intramuscular fat content of the living body according to the image encoding vector.
6. The method according to claim 5, characterized in that, The improved Transformer model further includes a population encoder; Before the step of inputting the target image into the second model to predict the intramuscular fat content of the living body, it includes: Obtaining the population information of the living body; The step of inputting the target image into the second model to predict the intramuscular fat content of the living body includes: The image encoder performs image encoding on the target image to obtain an image encoding vector; The population encoder performs population encoding on the population information to obtain a population encoding vector; The decoder predicts the intramuscular fat content of the living body according to the image encoding vector and the population encoding vector.
7. The method according to claim 1, characterized in that Before the step of inputting the target image into the first model to determine the quality of the target image, it includes: Collecting an image dataset collected using an ultrasonic device; Annotating the image dataset according to a preset standard to divide each image in the image dataset into two categories: qualified or unqualified, where the preset standard includes clear backfat and eye muscle, clear ribs in the field of view and requiring more than 4 ribs; Inputting the image dataset into the first model for training to obtain the trained first model.
8. The method according to claim 1, characterized in that Before the step of inputting the target image into the second model to predict the intramuscular fat content of the living body, it includes: Collecting an image dataset collected by an ultrasonic device; Screening out a target image dataset with qualified quality from the image dataset; Annotating the intramuscular fat content of each image in the target image dataset; Inputting the target image dataset annotated with intramuscular fat content into the second model for training to obtain the trained second model.
9. A processing device, characterized in that, Comprising: A processor configured to execute program instructions; And A memory configured to store program instructions that, when loaded and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing program instructions, characterized in that, When the program instructions are loaded and executed by a processor, they cause the processor to perform the method according to any one of claims 1 to 8.
Citation Information
Cited By
Deep learning-based methods for selecting breeding pigs, electronic devices, computer-readable storage media, and program products.
CN122575500A