Live pig weight measurement method based on depth estimation and multi-modal feature fusion

By combining a monocular camera with depth estimation and multimodal feature fusion, the automatic measurement of pig weight is achieved, which solves the problems of large errors, low efficiency and high cost in traditional methods, improves measurement accuracy and efficiency, and is suitable for smart farming environments.

CN120628254APending Publication Date: 2025-09-12GUANGDONG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510876426.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional pig weight measurement methods have large errors, low efficiency and high costs, and are difficult to meet the needs of precision agriculture, especially in large-scale farming.

Method used

A monocular camera combined with depth estimation technology is used to analyze a single RGB image for depth prediction. Combined with multimodal feature fusion, this system can automatically measure the weight of pigs, avoiding manual intervention and expensive hardware requirements.

Benefits of technology

The accuracy and efficiency of weight measurement are improved, labor and equipment costs are reduced, and the robustness and generalization ability in complex environments are enhanced. The error is controlled within ±4%, meeting the needs of smart farming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120628254A_ABST
    Figure CN120628254A_ABST
Patent Text Reader

Abstract

The invention provides a live pig weight measurement method based on depth estimation and multi-modal feature fusion, which comprises the following steps: firstly, initializing and reading camera setting, capturing an RGB image, then converting the RGB image into a depth image sequence through a depth estimation algorithm, carrying out feature extraction on the depth image by using a vit model, and carrying out feature extraction on the depth image; and carrying out feature fusion on the extracted features and rgb features extracted in depth estimation, and predicting the weight of the pig through a multi-layer perceptron regression model. According to the method, a monocular RGB image is utilized, three-dimensional space structure information of a target is reconstructed through a depth estimation neural network, an advanced monocular depth estimation network is adopted, correlation and influence between different images are captured through a Transform encoder, and the network gradient propagation efficiency is improved in combination with residual connection. Therefore, the robustness of the posture change of the pig, the shielding area and the background complexity is enhanced, the labor cost is greatly reduced through automatic measurement, and the weight measurement precision and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pig weight measurement, and in particular to a pig weight measurement method based on depth estimation and multimodal feature fusion. Background Art

[0002] Pig weight is an important indicator for assessing the growth and development of pigs, and is also a key basis for breeding, reproductive capacity assessment and precise feeding.

[0003] Traditional pig weight measurement methods have significant limitations and are unable to meet the growing demands of precision agriculture. Visual weight estimation relies on the owner's experience. While simple and easy, it suffers from significant errors and is subject to human error, making it inaccurate and particularly inadequate in large-scale farming. Body measurement involves using a tape measure or a speedometer to measure the pig's length, height, chest circumference, and other parameters, and then estimating weight based on these data. While relatively simple, manual measurement is inefficient and requires the pig to maintain a stable posture for accurate measurement, which places high demands on the pig's behavior and mobility in a farming environment and increases operational complexity. The weighbridge method, which uses a platform scale to measure pig weight, offers high accuracy, but the equipment is expensive and requires individual weighing for each measurement, which is time-consuming and labor-intensive. This makes it impractical in large-scale farming and fails to meet the needs of efficient management. Depth cameras are also currently being used for weight estimation, but the high cost of the equipment increases both procurement and operation costs for farms. Summary of the Invention

[0004] In response to the problems of the background technology, the present invention provides a pig weight measurement method based on depth estimation and multimodal feature fusion. By combining a monocular camera with depth estimation technology, it can obtain rich spatial information by analyzing a single RGB image and performing depth prediction without the need for special hardware support. This method not only avoids the stress response that may be caused to pigs by manual measurement, thus meeting animal welfare requirements, but also can greatly reduce labor costs through automated measurement, thereby improving the accuracy and efficiency of weight measurement.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method for measuring pig weight based on depth estimation and multimodal feature fusion, comprising the following steps:

[0007] (1) RGB image acquisition: Use a monocular camera to collect RGB images of pigs;

[0008] (2) Depth estimation: After feature extraction of the collected RGB image, a depth image is generated using a depth estimation algorithm;

[0009] (3) Depth feature extraction: perform depth feature extraction of depth images;

[0010] (4) Cross-modal fusion: Use cross-attention fusion network to fuse deep features and RGB features;

[0011] (5) Weight prediction: Use the MLP regression model to predict the weight of pigs.

[0012] Furthermore, the RGB image acquisition formula in step (1) is: ,in Represents a set of real numbers, H refers to the height of the image, W refers to the width of the image, and 3 corresponds to the three color channels of the RGB image.

[0013] Furthermore, step (2) adopts the DPT-Hybrid architecture in the MiDaS network to achieve depth estimation for a single RGB image. DPT (Dense Prediction Transformer) uses Vision Transformer (ViT) as its backbone and has strong global modeling capabilities. It can still output high-quality relative depth maps under conditions such as complex backgrounds, occlusions, and non-rigid bodies. The intermediate layer output of the Transformer encoder of the MiDaS network is extracted as RGB features, represented as F rgb .

[0014] Furthermore, the specific steps of step (3) are:

[0015] After obtaining the depth image, it is first divided into patches of fixed size and embedded into a vector sequence, which is then input into a lightweight Transformer encoder; this module extracts the three-dimensional geometric structure features contained in the depth image, which is recorded as

[0016] Furthermore, the specific steps of step (4) are:

[0017] Using RGB features as queries and depth features as keys and values, the corresponding relationship between the two modalities is established through the attention mechanism:

[0018]

[0019] in:

[0020] Q=F rgb

[0021] K,V=F depth

[0022] d k is the feature dimension normalization factor

[0023] Features output after fusion Characterizes the association between the pig's geometric structure and two-dimensional information

[0024] where N is the number of feature groups and d is the feature dimension of each feature group.

[0025] Furthermore, the specific steps of step (5) are:

[0026] After the fusion features are globally averaged and pooled, a global description vector is obtained. Then input a three-layer MLP network to complete the nonlinear mapping from fusion features to weight:

[0027]

[0028] in, is the final predicted pig weight, σ is the ReLU activation function, W i and b i are the weight and bias parameters of the MLP.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] (1) Reduce costs while improving estimation capabilities

[0031] This method relies solely on a monocular camera and uses a depth estimation algorithm to recover 3D structural information from a single RGB image frame, eliminating the need for expensive depth camera hardware and significantly reducing system deployment costs. Relatively accurate depth information can be obtained even under natural lighting and complex backgrounds, providing key spatial features for weight estimation. Compared to traditional solutions based solely on 2D images, weight prediction error is reduced by approximately 25%-35% in scenes with large pose variations or occlusion, achieving highly cost-effective estimation results.

[0032] (2) Lightweight and highly adaptable solution

[0033] Monocular depth estimation is integrated with an attention mechanism to form an end-to-end lightweight network framework suitable for edge computing environments (such as embedded devices in piggeries). A cross-modal attention module is introduced during the feature fusion stage to dynamically adjust the fusion ratio of RGB and estimated depth based on the scene, effectively mitigating the impact of depth estimation errors. Increased attention to key areas (such as the back contour and leg structure) ensures that feature representations are more accurately aligned with the actual body shape.

[0034] (3) Automated continuous monitoring supports smart farming

[0035] Compared with contact weighing methods, weight estimation based on monocular vision has the advantage of being non-invasive. The measurement process does not require human intervention, which reduces the stress level of pigs.

[0036] (4) Improve robustness in complex environments

[0037] Through the guided fusion of an attention mechanism and a multi-scale feature extraction strategy, the model significantly improves stability in complex scenarios such as overlapping pigs, occlusion, and changing lighting. The feature fusion module adaptively captures and distinguishes foreground pigs from background interference during training, resulting in a model with enhanced generalization across pig breeds and sizes. In experiments, the overall prediction error was kept within ±4%, and the missed detection rate dropped below 6%, meeting the needs of pig management throughout their lifecycle at different stages.

[0038] (5) Data-driven intelligent analysis capabilities

[0039] By generating synchronized RGB and pseudo-depth information from monocular images, the system achieves multimodal feature alignment and learning. Residual connections and a multi-head attention mechanism improve the model's responsiveness to key areas (such as backfat thickness and trunk length), enabling the network to more efficiently learn the mapping relationship between weight and visual features. Compared to a single-input model, the fusion model achieved an approximately 18% lower RMSE on the validation set, providing reliable data support for subsequent precision farming. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:

[0041] Figure 1 Flow chart of the method of the present invention.

[0042] Figure 2 This is the overall flow chart of the MiDaS network.

[0043] Figure 3 Flowchart of the cross-modal feature fusion method.

[0044] Figure 4 Flowchart of the MLP weight regression method. DETAILED DESCRIPTION

[0045] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0046] like Figure 1 As shown, a pig weight measurement method based on depth estimation and multimodal feature fusion of the present invention includes the following steps:

[0047] 1. Data collection and preprocessing

[0048] Use a monocular RGB camera installed above the pig house to capture images of a single pig and obtain the original color image This image serves as the input to both the depth estimation network and the subsequent RGB feature extraction module.

[0049] 2. Depth Estimation (MiDaS Network)

[0050] To obtain the 3D geometric structure of pigs, this paper uses the DPT-Hybrid architecture within the MiDaS network to achieve depth prediction from a single RGB image. The DPT (Dense Prediction Transformer), built on the Vision Transformer (ViT), possesses strong global modeling capabilities and can output high-quality relative depth maps despite complex backgrounds, occlusions, and non-rigid objects.

[0051] Figure 2 The overall process structure of the MiDaS network is shown.

[0052] 3. Feature extraction module

[0053] 3.1 Deep Feature Extraction

[0054] After obtaining the depth image, it is first divided into fixed-size patches (such as 16×16) and embedded as a vector sequence, which is then input into a lightweight Transformer encoder. This module extracts the three-dimensional geometric structure features contained in the depth image, which are recorded as Will serve as the Key and Value in cross attention.

[0055] 3.2 RGB feature extraction

[0056] The intermediate layer output of the Transformer encoder of the MiDaS network is extracted as RGB features (without going through an additional encoder), denoted as F rgb , which will serve as the Query of the cross attention module.

[0057] 4. Cross-modal feature fusion (cross attention)

[0058] In order to effectively fuse the two-dimensional information in the RGB image and the spatial structure information in the depth map, the present invention introduces a cross attention mechanism, such as Figure 3 The core idea is to use RGB features as query, deep structure features as key and value, and establish the correspondence between the two modalities through the attention mechanism:

[0059]

[0060] in:

[0061] Q=F rgb

[0062] K,V=F depth

[0063] d k is the feature dimension normalization factor

[0064] Features output after fusion The pig's geometric structure and two-dimensional information are jointly represented.

[0065] 5. MLP weight regression module

[0066] like Figure 4 As shown, after the fusion features are globally averaged and pooled, the global description vector is obtained. Then it is input into a three-layer multi-layer perceptron (MLP) network to complete the nonlinear mapping from fusion features to weight:

[0067]

[0068] in, is the final predicted pig weight, σ is the ReLU activation function, W i and b i are the weight and bias parameters of the MLP.

[0069] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A pig weight measurement method based on depth estimation and multimodal feature fusion, characterized by: The following steps are involved: (1) RGB image acquisition: Use a monocular camera to collect RGB images of pigs; (2) Depth estimation: After feature extraction of the collected RGB image, a depth image is generated using a depth estimation algorithm; (3) Depth feature extraction: perform depth feature extraction of depth images; (4) Cross-modal fusion: Use cross-attention fusion network to fuse deep features and RGB features; (5) Weight prediction: Use the MLP regression model to predict the weight of pigs.

2. The method for measuring pig weight based on depth estimation and multimodal feature fusion according to claim 1, characterized in that: The formula for obtaining the RGB image in step (1) is: ,in Represents a set of real numbers, H refers to the height of the image, W refers to the width of the image, and 3 corresponds to the three color channels of the RGB image.

3. The method for measuring pig weight based on depth estimation and multimodal feature fusion according to claim 2, characterized in that: The step (2) adopts the DPT-Hybrid architecture in the MiDaS network to achieve depth estimation of a single RGB image.

4. The method for measuring pig weight based on depth estimation and multimodal feature fusion according to claim 3, characterized in that: In step (2), the intermediate layer output of the MiDaS network Transformer encoder is extracted as RGB features, denoted as F rgb .

5. The method for measuring pig weight based on depth estimation and multimodal feature fusion according to claim 4, characterized in that: The specific steps of step (3) are: After obtaining the depth image, it is first divided into patches of fixed size and embedded into a vector sequence, which is input into a lightweight Transformer encoder; this module extracts the three-dimensional geometric structure features contained in the depth image, which is recorded as 6. The method for measuring pig weight based on depth estimation and multimodal feature fusion according to claim 5, characterized in that: The specific steps of step (4) are: Using RGB features as queries and depth features as keys and values, the corresponding relationship between the two modalities is established through the attention mechanism: in: Q=F rgb K,V=F depth d k is the feature dimension normalization factor Features output after fusion The joint representation of the geometric structure and two-dimensional information of the pig is characterized, where N refers to the number of feature groups and d represents the feature dimension of each feature group.

7. The method for measuring pig weight based on depth estimation and multimodal feature fusion according to claim 6, characterized in that: The specific steps of step (5) are: After the fusion features are globally averaged and pooled, a global description vector is obtained. Then input a three-layer MLP network to complete the nonlinear mapping from fusion features to weight: in, is the final predicted pig weight, σ is the ReLU activation function, W i and b i are the weight and bias parameters of the MLP.

Citation Information

Cited By

  • A visual measurement method for pig body size

    CN122492765A