Cow accurate feeding method and system based on vision and artificial intelligence
By applying visual and artificial intelligence technology in cattle feeding management, automated feeding trough scoring and feeding plan adjustments are achieved, the problem of inefficiency of traditional manual inspections is solved, and the scoring accuracy and breeding benefits are improved.
Patent Information
- Application Number
- CN202510268038.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional cattle feeding management relies on manual inspection, which is inefficient and prone to human errors, resulting in inaccurate trough scoring, which increases feeding costs and affects breeding benefits.
The cattle precise feeding method based on vision and artificial intelligence is adopted, and the automatic feeding trough scoring and feeding plan adjustment is achieved through image acquisition, histogram equalization, image defog removal, YOLOv5 object detection and ViT classification algorithm.
It improves the efficiency and accuracy of the trough score, reduces labor costs and human operation errors, achieves more accurate feeding management, and reduces feeding waste and breeding costs.
Smart Images

Figure CN120126178A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent feeding technology. Specifically, it relates to a precise feeding method and system for cattle based on vision and artificial intelligence. Background Art
[0002] In the beef cattle fattening and breeding system, the feeding cost accounts for an important part of the total cost, exceeding 70%. Therefore, optimizing the feeding technology and effectively controlling the cost are particularly crucial for improving the overall efficiency of the ranch. However, traditional small and medium-sized ranches often adopt the method of restricted feeding, such as feeding only twice a day. However, this mode has two main problems. The first is the waste problem. If the leftover feed in the beef cattle farm is cleaned up, it will cause a certain economic loss. But if it is not cleaned up, it will affect the quality of the daily ration, and then lead to a decrease in the feed intake of cattle. Secondly, if the proportion of high-concentration feed in the fattening ration is relatively high, it is easy to cause metabolic problems such as acidosis, which will then cause a large fluctuation in the feed intake.
[0003] In order to better evaluate and manage the feeding situation of beef cattle, a feed trough scoring system can be adopted. This system details different amounts of leftover feed and their characteristics, providing valuable reference for ranch managers.
[0004] The key to feed trough scoring lies in grasping the time and sequence. Usually, it is recommended to conduct scoring half an hour to one hour before the morning feeding, and evaluate all feed troughs one by one in the feeding order.
[0005] During the scoring process, several different situations of leftover feed need to be noted. For the situation of "a large amount of leftover feed", this usually indicates a sudden decrease in the feed intake of cattle, and the reason needs to be found out and the feeding amount adjusted in time; for the situation of "sporadic leftover feed", this may just be because the feeding amount the previous day was slightly more, and the feeding amount for the current day should be appropriately reduced at this time. In addition, the situation of "dry feed trough" indicates that the feed amount is insufficient and the cattle have finished eating a long time ago, so the feeding amount needs to be increased to ensure that the cattle can obtain enough feed. The ideal state of the feed trough is "wet feed trough", which means that the cattle just finish eating all the feed after the feeding, being in a suitable state.
[0006] The feed intake of cattle largely determines their growth status and health level. By observing the feeding behavior of cattle, analyzing the feeding pattern, accurately identifying the cattle with abnormal feeding in a timely manner, and adjusting the feeding amount in real time, feed waste can be effectively reduced and the breeding efficiency can be improved.
[0007] The management of feed trough scoring in the traditional cattle feeding process usually relies on manual inspection. These methods are inefficient, prone to human errors, which will affect the accuracy of inventory data, and are very labor-consuming.
[0008] Currently, most of the trough scoring management systems on the market are based on solutions such as drones and inspection robots, which are costly and cumbersome to operate. With the popularization of intelligentization and the continuous development of the field of deep learning, there is a need for low-cost and high-efficiency solutions to improve the existing situation. Summary of the Invention
[0009] In view of this, the present application provides a method and system for precise feeding of cattle based on vision and artificial intelligence to reduce the cost of trough scoring during the feeding process and improve the efficiency of providing trough scoring, so as to achieve the purpose of timely adjusting and modifying the trough plan.
[0010] To achieve the above object, the technical solution adopted by the present application is as follows: A method for precise feeding of cattle based on vision and artificial intelligence, comprising: S1: Collect relevant images of the trough in the cattle pen. S2: Perform histogram equalization and image dehazing operations on the collected images to improve the contrast and clarity of the images. S3: Based on the YOLOv5 object detection algorithm, detect and locate the trough in the image, obtain the trough area image and extract it. S4: Classify the extracted trough image through the ViT (Vision Transformer) classification algorithm to obtain the scoring situation of the trough. S5: Comprehensively analyze the trough scoring situation, and automatically adjust and modify the feeding plan for the next meal in combination with the growth situation of the cattle in the pen. S6: Store the historical data, feeding records and algorithm models of the cattle, and perform data analysis and algorithm optimization.
[0011] Further, the histogram equalization is specifically: S2.11: Calculate the histogram of the original image Count the number of occurrences of each pixel value in the image.
[0012] S2.12: Calculate the normalized histogram Divide the number of occurrences of each pixel value in the histogram by the total number of pixels in the image to obtain the probability distribution of each pixel value. The formula is:
[0013] Where: represents the kth pixel value, represents the pixel value The number of occurrences, and N represents the total number of pixels in the image.
[0014] S2.13: Calculate the cumulative distribution function (CDF) Accumulate the normalized histogram to obtain the cumulative distribution function. The formula is:
[0015] S2.14: Map pixel values Map the original pixel values to new pixel values according to the CDF. The mapping formula is:
[0016] Where: represents the pixel value after mapping, L represents the maximum value of the pixel value, and round represents rounding.
[0017] S2.15: Generate the equalized image Replace the pixel values of the original image with the mapped pixel values to generate the equalized image.
[0018] Furthermore, the image defogging operation is specifically: S2.21: Calculate the dark channel For each pixel of the image, take the minimum value of the RGB three color channels within its local window. The formula is:
[0019] Where: represents the local window centered on pixel x and represents the value of pixel y in color channel c.
[0020] S2.22: Estimate the atmospheric light value A Select the brightest 0.1% pixels from the dark channel, and use the highest brightness value of these pixels in the original image as the atmospheric light value.
[0021] S2.23: Estimate the transmittance Assume that the transmittance is constant within the local area, and use the dark channel prior to estimate the transmittance. The formula is:
[0022] Where ω is a tuning parameter, is the camera parameter, represents the value of pixel in color channel c.
[0023] S2.24: Restore the fog-free image Restore the fog-free image according to the degradation model. The formula is:
[0024] Where: is the lower limit of the transmittance, used to avoid too small a denominator, represents the pixel before defogging value, represents the pixel after defogging value.
[0025] Furthermore, the model structure of YOLOv5 includes a backbone network, a neck network, and a detection head. The backbone network uses CSPDarknet53 to reduce the computational load and improve the feature extraction ability; the neck network uses PANet as the neck network, and through top-down and bottom-up path aggregation, enhances the multi-scale feature fusion ability of the feature pyramid; and an SPP (Spatial Pyramid Pooling) module is introduced in PANet, and through pooling operations of different scales, the receptive field of the model is enhanced; the detection head is responsible for generating the final detection results. The detection head consists of multiple convolutional layers, outputs feature maps of three scales, and each feature map corresponds to targets of different sizes; YOLOv5 uses predefined Anchor Boxes to predict the bounding boxes of targets, and each scale of feature map corresponds to a group of Anchor Boxes.
[0026] Furthermore, the loss function of YOLOv5 consists of three parts, namely classification loss (ClassificationLoss), localization loss (Localization Loss), and confidence loss (Confidence Loss). The classification loss uses cross-entropy loss to measure the difference between the predicted class and the true class; the localization loss uses CIoU (Complete Intersection over Union) loss to evaluate the deviation between the bounding box predicted by the model and the true bounding box. CIoU not only considers the overlapping area between the bounding boxes, but also introduces the distance between the center points and the aspect ratio, so as to more accurately measure the similarity between the bounding boxes; the confidence loss uses binary cross-entropy loss to measure the confidence of whether the predicted box contains the target.
[0027] Furthermore, in the post-processing stage of YOLOv5, through the non-maximum suppression (NMS) algorithm, redundant boxes with a high degree of overlap are removed, and the most likely detection results are retained; and according to the confidence threshold and the class probability threshold, the detection results with low confidence are filtered out; the final output results include the class, confidence, and the coordinates (x, y, width, height) of the bounding box of each detected target, where x represents the X-axis coordinate of the center of the bounding box in the image, y represents the y-axis coordinate of the center of the bounding box in the image, width represents the width of the bounding box, and height represents the height of the bounding box.
[0028] Furthermore, the ViT classification algorithm includes an input preprocessing stage, specifically: S4.1: Divide the input image into N fixed-size patches, where the size of each patch is , H and W represent the height and width of the input image respectively, C represents the number of channels, and R represents the pixel space. Therefore, the number of patches is:
[0029] S4.2: Flatten the patches, where each patch is flattened into a vector with a dimension of C, and each flattened patch is transformed into a D-dimensional vector through a learnable linear mapping to obtain patch embeddings: where, is the linear mapping matrix, is the i th flattened vector of the patch; S4.3: Add positional encoding: Since the Transformer itself has no positional information, ViT introduces learnable positional encoding to retain the spatial position information of the patches: where, is the i th positional encoding of the patch; is the input vector after adding the positional encoding.
[0030] S4.4: Add a learnable class token at the beginning of the patch embeddings sequence for the final classification task: where, .
[0031] Furthermore, the core of the ViT classification algorithm is the Transformer encoder, which is stacked by multiple Transformer blocks. Each Transformer block includes the following components:
[0032] Multi-Head Self-Attention (MSA): It is used to complete self-attention calculation, that is, for the input sequence Z, calculate the relationship between each patch and other patches. The calculation formula is: Among them, Q, K, and V are the query, key, and value obtained through linear transformation respectively matrix.
[0033] Feed-Forward Network (FFN): It is used to complete non-linear transformation, that is, each attention output is non-linearly transformed through a two-layer feed-forward neural network: Among them, , , and are network parameters.
[0034] Residual Connection & Layer Normalization Residual connection: Add a residual connection after each sub-layer (MSA and FFN) to alleviate the problem of gradient disappearance: .
[0035] Furthermore, the loss function of the ViT classification algorithm is cross-entropy loss, and the cross-entropy loss (Cross-Entropy Loss) is used to calculate the difference between the predicted class and the true label. The specific formula is: Among them, is the true label, is the predicted probability.
[0036] A cattle precise feeding system based on vision and artificial intelligence is used to execute the foregoing cattle precise feeding method based on vision and artificial intelligence in this application, and specifically includes the following modules: A vision acquisition module, which is used to acquire relevant images of the feed trough in the cattle pen; A vision processing module, which is used to perform histogram equalization and image dehazing operations on the acquired images to improve the contrast and clarity of the images; The target detection module is used to detect and locate the feeding trough in the image based on the YOLOv5 target detection algorithm, obtain the feeding trough area image and extract it; The intelligent classification module is used to classify the extracted feeding trough image through the ViT (Vision Transformer) classification algorithm to obtain the scoring situation of the feeding trough; The feeding decision-making module is used to comprehensively analyze the scoring situation of the feeding trough, combine the growth situation of the cattle in the pen, and automatically adjust and modify the feeding plan for the next meal; The data management and optimization module is used to store the historical data, feeding records and algorithm models of the cattle, and conduct data analysis and algorithm optimization.
[0037] Compared with the prior art, the beneficial effects of this application are: 1. Automatic feeding trough scoring: Replacing the manual observation and scoring method, it realizes the automatic scoring of the feeding trough in the cattle pen; 2. High-precision detection: The YOLOV5 target detection algorithm trained on a large scale improves the accuracy of target detection and performs excellently in complex environments; 3. High-precision classification: Through the ViT algorithm for image classification based on the Transformer architecture, it realizes the fast and efficient classification and scoring of the feeding trough image; 4. Cost reduction and efficiency improvement: While greatly reducing the manual time cost, it improves the efficiency of scoring the feeding trough in the cattle pen and ensures the accuracy of the feeding trough scoring. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of this application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0039] Figure 1 It is a flowchart of a precise cattle feeding method based on vision and artificial intelligence of this application; Figure 2 It is a structural block diagram of a precise cattle feeding system based on vision and artificial intelligence of this application. Detailed Embodiments
[0040] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them.
[0041] Such asFigure 1 As shown in the figure, a precise feeding method for cattle based on vision and artificial intelligence includes: S1: Collect relevant images of the feed trough in the cattle pen. S2: Perform histogram equalization and image dehazing on the collected images to improve the contrast and clarity of the images. Preprocess the collected images, specifically including histogram equalization processing and image dehazing processing, to improve the contrast and clarity of the images, reduce noise and light interference, thereby improving the accuracy of subsequent artificial intelligence analysis. In practice, opencv can be used to perform histogram equalization and image dehazing operations on the images.
[0042] S3: Based on the YOLOv5 object detection algorithm, detect and locate the feed trough in the image, obtain the feed trough area image and extract it. S4: Classify the extracted feed trough image through the ViT (Vision Transformer) classification algorithm to obtain the scoring situation of the feed trough. S5: Comprehensively analyze the scoring situation of the feed trough, and combine it with the growth situation of the cattle in the pen to automatically adjust and modify the feeding plan for the next meal. According to the scoring situation of the feed trough, timely adjust the feeding plan and ratio for the next meal based on the historical data and growth data of cattle feeding. S6: Store the historical data, feeding records and algorithm models of cattle, and conduct data analysis and algorithm optimization.
[0043] In addition, this step can also display the comprehensive analysis results, and can manually adjust the final feeding plan according to special situations (such as special situations where cattle are sick in the pen).
[0044] As a further implementation method, the histogram equalization is specifically as follows: S2.11: Calculate the histogram of the original image Count the number of occurrences of each pixel value in the image.
[0045] For example, for an 8-bit grayscale image (pixel value range 0-255), count the number of occurrences of each pixel value (0, 1, 2,..., 255).
[0046] S2.12: Calculate the normalized histogram Divide the number of occurrences of each pixel value in the histogram by the total number of pixels in the image to obtain the probability distribution of each pixel value. The formula is:
[0047] Where: represents the k-th pixel value, Indicates the pixel value The number of occurrences, and N represents the total number of pixels in the image.
[0048] S2.13: Calculate the cumulative distribution function (CDF) Accumulate the normalized histogram to obtain the cumulative distribution function. The formula is:[[]]
[0049] S2.14: Map the pixel values Map the original pixel values to new pixel values according to the CDF. The mapping formula is:[[]]
[0050] Where:[[]] Indicates the pixel value after mapping; L represents the maximum value of the pixel value. For an 8-bit image, L = 256; round represents rounding.
[0051] S2.15: Generate the equalized image Replace the pixel values of the original image with the mapped pixel values to generate the equalized image.
[0052] Furthermore, the image defogging operation is specifically as follows: S2.21: Calculate the dark channel For each pixel of the image, take the minimum value of the RGB three color channels within its local window (such as 15*15). The formula is:[[]]
[0053] Where:[[]] Indicates the local window centered on the pixel x Indicates the pixel y The value in the color channel c.
[0054] S2.22: Estimate the atmospheric light value A Select the brightest 0.1% pixels from the dark channel, and use the highest brightness value of these pixels in the original image as the large Atmospheric light value.
[0055] S2.23: Estimate the transmittance Assume that the transmittance is constant within the local area, and use the dark channel prior to estimate the transmittance. The formula is:[[]]
[0056] Where ω is the adjustment parameter, usually taking the value 0.95; Is the camera parameter; Indicates the pixel The value in the color channel c.
[0057] S2.24: Restore the haze-free image Restore the haze-free image according to the degradation model. The formula is:
[0058] Where: is the lower limit of the transmittance (usually taken as 0.1) to avoid too small a denominator, represents the pixel value before haze removal value, represents the pixel value after haze removal value.
[0059] The YOLOV5 object detection algorithm described above specifically includes an input preprocessing stage, which specifically includes: ① Image scaling: The input image is scaled to a fixed size (such as 1280x1280) to meet the input requirements of the model.
[0060] ② Normalization: The image pixel values are normalized to the range [0, 1], usually achieved by dividing by 255.
[0061] ③ Data augmentation: An optional step, including random cropping, rotation, flipping, etc., to increase the diversity of data.
[0062] As a further implementation method, the model structure of YOLOv5 includes a backbone network, a neck network, and a detection head. The backbone network uses CSPDarknet53, which is based on Darknet53 and introduces a Cross Stage Partial (CSP) structure to reduce the computational amount and improve the feature extraction ability; the neck network uses PANet as the neck network, and through top-down and bottom-up path aggregation, enhances the multi-scale feature fusion ability of the feature pyramid; and an SPP (Spatial Pyramid Pooling) module is introduced in PANet, and through pooling operations at different scales, the receptive field of the model is enhanced; the detection head is responsible for generating the final detection results. The detection head consists of multiple convolutional layers, and feature maps of three scales are output, and each feature map corresponds to targets of different sizes; YOLOv5 uses predefined Anchor Boxes to predict the bounding boxes of targets, and each scale of feature map corresponds to a group of Anchor Boxes.
[0063] As a further implementation, the loss function of YOLOv5 consists of three parts, namely classification loss, localization loss, and confidence loss. The classification loss uses cross-entropy loss to measure the difference between the predicted class and the true class. The localization loss uses CIoU (Complete Intersection over Union) loss to evaluate the deviation between the bounding box predicted by the model and the true bounding box. CIoU not only considers the overlapping area between the bounding boxes but also introduces the distance between the center points and the aspect ratio, thus being able to more accurately measure the similarity between the bounding boxes. The confidence loss uses binary cross-entropy loss to measure the confidence of whether the predicted bounding box contains the target.
[0064] As a further implementation, during the prediction stage, YOLOv5 generates a large number of candidate bounding boxes. In the post-processing stage, YOLOv5 uses the non-maximum suppression (NMS) algorithm to remove redundant bounding boxes with high overlap and retains the most likely detection results. And according to the confidence threshold and class probability threshold, it filters out the detection results with low confidence. The final output results include the class, confidence, and the coordinates (x, y, width, height) of the bounding box for each detected object, where x represents the X-axis coordinate of the center of the bounding box in the image, y represents the y-axis coordinate of the center of the bounding box in the image, width represents the width of the bounding box, and height represents the height of the bounding box.
[0065] As a further implementation, the ViT classification algorithm includes an input preprocessing stage, specifically: S4.1: Patch Embedding: The input image is divided into N fixed -sized patches, and the size of each patch is , where H and W represent the height and width of the input image respectively, C represents the number of channels, and R represents the pixel space. Therefore, the number of patches is:
[0066] S4.2: Linear Projection: Flatten the patches, and each patch is flattened into a vector with a dimension of C. Each flattened patch is converted into a D-dimensional vector through a learnable linear mapping to obtain patch embeddings: Among them, is a linear mapping matrix, is the i flattened vector of the S4.3: Add Positional Encoding: Since the Transformer itself has no positional information, ViT introduces learnable Positional Encoding to preserve the spatial positional information of patches: is the i positional encoding of the patch;
[0067] is the input vector after adding the positional encoding. S4.4: Add Class Token: Add a learnable Class Token at the beginning of the patch embeddings sequence .
[0068] As a further implementation, the core of the ViT classification algorithm is the Transformer encoder, which is stacked by multiple
[0069] Transformer Blocks. Each Transformer block includes the following components: Multi-Head Self-Attention (MSA): Used to complete self-attention calculation, that is, for the input sequence Z, calculate the relationship between each patch and
[0070] other patches, and the calculation formula is: where Q, K, and V are the query, key, and value Among them, 、 、 and are network parameters.
[0071] Residual Connection & Layer Normalization Residual Connection: Add a residual connection after each sub-layer (MSA and FFN) to alleviate the vanishing gradient problem: .
[0072] Layer Normalization: Apply Layer Normalization after each sub-layer.
[0073] After passing through all Transformer blocks, map the classification token to a vector of the number of classes K through a fully connected layer (MLP). quantity K.
[0074] As a further implementation, the loss function of the ViT classification algorithm is cross-entropy loss, and the cross-entropy loss (Cross-Entropy Loss) is used to calculate the difference between the predicted class and the true label. The specific formula is: Among them, is the true label, is the predicted probability.
[0075] As Figure 2 shown, a precise cattle feeding system based on vision and artificial intelligence is used to execute the aforementioned precise cattle feeding method based on vision and artificial intelligence in this application, and specifically includes the following modules: Visual acquisition module 210, which is used to acquire relevant images of the feed trough in the cattle pen; Visual processing module 220, which is used to perform histogram equalization and image dehazing operations on the acquired images to improve the contrast and clarity of the images; Target detection module 230, which is used to detect and locate the feed trough in the image based on the YOLOv5 target detection algorithm, obtain the feed trough area image and extract it; Intelligent classification module 240, which is used to classify the extracted feed trough images through the ViT (Vision Transformer) classification algorithm to obtain the scoring situation of the feed trough; The feeding decision-making module 250 is used to comprehensively analyze the trough scoring situation, combine the growth situation of the cattle in the pen, and automatically adjust and modify the feeding plan for the next meal; The data management and optimization module 260 is used to store the historical data, feeding records and algorithm models of the cattle, and perform data analysis and algorithm optimization.
[0076] This application can achieve high-precision trough positioning and classification scoring under different lighting conditions and complex environments, effectively improving the efficiency and accuracy of trough scoring management in the cattle feeding process, being applicable to a variety of complex trough scenarios, and significantly reducing labor costs and human operation errors.
[0077] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. A method for precise feeding of cattle based on vision and artificial intelligence, characterized in that: include: S1: Collect relevant images of the feed trough in the cattle pen; S2: Perform histogram equalization and image defogging operations on the collected image to improve the contrast and clarity of the image; S3: Based on the YOLOv5 target detection algorithm, the trough in the image is detected and located, and the trough area image is obtained and extracted; S4: classify the extracted trough image through the ViT classification algorithm to obtain the trough score; S5: Comprehensively analyze the trough score and automatically adjust the feeding plan for the next meal based on the growth of the cattle in the pen; S6: Stores cattle historical data, feeding records and algorithm models for data analysis and algorithm optimization.
2. A method for precise feeding of cattle based on vision and artificial intelligence as claimed in claim 1, characterized in that: The histogram equalization is specifically as follows: S2.11: Calculate the histogram of the original image Count the number of occurrences of each pixel value in the image; S2.12: Calculate the normalized histogram Divide the number of occurrences of each pixel value in the histogram by the total number of pixels in the image to obtain the probability distribution of each pixel value; the formula is: in: represents the kth pixel value, Represents pixel value The number of occurrences, N represents the total number of pixels in the image; S2.13: Calculate the cumulative distribution function The normalized histograms are accumulated to obtain the cumulative distribution function; the formula is: S2.14: Mapping pixel values The original pixel value is mapped to the new pixel value according to the CDF; the mapping formula is: in: Represents the pixel value after mapping, L represents the maximum value of the pixel value, and round represents rounding; S2.15: Generate equalized image The mapped pixel values are used to replace the pixel values of the original image to generate an equalized image.
3. A method for accurate feeding of cattle based on vision and artificial intelligence as claimed in claim 2, characterized in that: The image defogging operation is specifically as follows: S2.21: Calculate dark channel For each pixel of the image, take the minimum value of the three RGB color channels in its local window; the formula is: in: In pixels x is the local window centered on Represents pixels y The value in color channel c; S2.22: Estimation of atmospheric light value A Select the brightest 0.1% pixels from the dark channel and use the highest brightness value of these pixels in the original image as the large Air light value; S2.23: Estimation of Transmittance Assuming that the transmittance is constant in the local area, the transmittance is estimated using the dark channel prior; the formula is: Among them, ω is the adjustment parameter, are the camera parameters, Represents pixels The value in color channel c; S2.24: Restoring a haze-free image According to the degradation model, the haze-free image is restored; the formula is: in: is the lower limit of transmittance, used to avoid the denominator being too small, Indicates the pixel before defogging The value of Represents the pixel after defogging The value of .
4. A method for precise feeding of cattle based on vision and artificial intelligence as claimed in claim 3, characterized in that: The YOLOv5 model structure includes a backbone network, a neck network and a detection head. The backbone network adopts CSPDarknet53 is used to reduce the amount of calculation and improve the feature extraction capability; PANet is used as the neck network to enhance the multi-scale feature fusion capability of the feature pyramid through top-down and bottom-up path aggregation; the SPP module is introduced in PANet to enhance the receptive field of the model through pooling operations at different scales; the detection head is responsible for generating the final detection result. The detection head consists of multiple convolutional layers and outputs feature maps of three scales, each of which corresponds to targets of different sizes; YOLOv5 uses predefined Anchor Boxes to predict the bounding box of the target, and each scale of the feature map corresponds to a set of Anchor Boxes.
5. A method for accurate feeding of cattle based on vision and artificial intelligence as claimed in claim 4, characterized in that: The loss function of YOLOv5 consists of three parts: classification loss, positioning loss and confidence loss. The classification loss uses cross entropy loss to measure the difference between the predicted category and the true category; the positioning loss uses CIoU loss to evaluate the deviation between the bounding box predicted by the model and the true bounding box. CIoU not only considers the overlapping area between the bounding boxes, but also introduces the center point distance and aspect ratio, so that the similarity between the bounding boxes can be measured more accurately; the confidence loss uses binary cross entropy loss to measure the confidence of whether the predicted box contains the target.
6. A method for accurate feeding of cattle based on vision and artificial intelligence as claimed in claim 5, characterized in that: In the post-processing stage, YOLOv5 uses the non-maximum suppression algorithm to remove redundant boxes with high overlap and retain the most likely detection results; and filters out low-confidence detection results based on the confidence threshold and category probability threshold; the final output results include the category, confidence, and coordinates of the bounding box (x, y, width, height) of each detected target, where x represents the x-axis coordinate of the center of the bounding box in the image, y represents the y-axis coordinate of the center of the bounding box in the image, width represents the width of the bounding box, and height represents the height of the bounding box.
7. A method for accurate feeding of cattle based on vision and artificial intelligence as claimed in claim 6, characterized in that: The ViT classification algorithm includes an input preprocessing stage, specifically: S4.1: Input image It is divided into N fixed-size patches, each with a size of Xiaowei , H and W represent the height and width of the input image, C represents the number of channels, and R represents the pixel space; therefore, the number of patches is: S4.2: Flatten patches. Each patch is flattened into a vector with dimension C, flatten each The patch is converted into a D-dimensional vector through a learnable linear mapping to obtain patch embeddings: in, is the linear mapping matrix, It is i The flattened vector of patches; S4.3: Adding position encoding: Since the Transformer itself has no position information, ViT introduces a learnable position Set the encoding to preserve the spatial location information of patches: in, It is i The position encoding of each patch; is the input vector after adding the position encoding; S4.4: Add a learnable classification token at the beginning of the sequence of patch embeddings , for the final Classification task: in, .
8. A method for accurate feeding of cattle based on vision and artificial intelligence as claimed in claim 7, characterized in that: The core of the ViT classification algorithm is the Transformer encoder, which consists of multiple Transformer blocks; Stacked; each Transformer block includes the following components: Multi-head self-attention mechanism: used to complete self-attention calculation, that is, for the input sequence Z, calculate the relationship between each patch and other patches. The calculation formula is: Among them, Q, K, and V are the query, key, and value matrices obtained through linear transformation, respectively; Feedforward neural network: used to complete nonlinear transformation, that is, each attention output is transformed nonlinearly through a two-layer feedforward neural network: in, , , and are network parameters; Residual connection and layer normalization Residual connection: Add residual connection after each sub-layer to alleviate the gradient disappearance problem: 。 9. A method for accurate feeding of cattle based on vision and artificial intelligence as claimed in claim 8, characterized in that: The loss function of the ViT classification algorithm is the cross entropy loss, which is used to calculate the difference between the predicted category and the true label. The specific formula is: in, is the true label, is the predicted probability.
10. A precision cattle feeding system based on vision and artificial intelligence, characterized in that: A method for accurately feeding cattle based on vision and artificial intelligence for executing any one of claims 1 to 9, specifically comprising the following modules: A visual acquisition module is used to collect images of feed troughs in cattle pens; A visual processing module is used to perform histogram equalization and image defogging operations on the collected images to improve the contrast and clarity of the images; The target detection module is used to detect and locate the trough in the image based on the YOLOv5 target detection algorithm, obtain the trough area image and extract it; Intelligent classification module, used to classify the extracted trough images through ViT classification algorithm to obtain the trough score; The feeding decision module is used to comprehensively analyze the trough score and automatically adjust the feeding plan for the next meal based on the growth of the cattle in the pen. The data management and optimization module is used to store cattle historical data, feeding records and algorithm models, and perform data analysis and algorithm optimization.
Citation Information
Patent Citations
Manger hay temperature image processing method based on artificial intelligence and active ball machine
CN111985472A
Papillary thyroid carcinoma lymph node metastasis prediction method based on Transform-MIL
CN114188020A
Electronic skin spatial resolution detection method
CN118411331A
Method and system for predicting feed intake of perinatal cows in pasture breeding scene
CN119168425A
Method for processing a substrate
KR102817738B1
Cited By
Video stream-based beef cattle feed intake intelligent monitoring and analysis method
CN121095200A