Broiler instance segmentation method based on thermal imaging and RGB fusion
By combining thermal imaging and RGB images deep learning methods, cross-modal attention mechanism and multi-source position coding are used to solve the problem of insufficient segmentation accuracy of broiler instances in high-density flocks and complex backgrounds, and efficient and real-time broiler individual identification and health monitoring are achieved.
Patent Information
- Application Number
- CN202510527451.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The prior art is difficult to effectively distinguish individual chickens in high-density flocks, complex backgrounds and dynamic environments, and traditional RGB images provide limited information in low-light and occlusion scenarios, resulting in insufficient segmentation accuracy.
The RGB visible light camera and thermal imager are used to collect data, and the thermal imaging and RGB images are deeply fused through deep learning models. The cross-modal attention mechanism and multi-source position coding are used to optimize the cross-modal information interaction to generate the final broiler instance segmentation mask.
It improves the system's adaptability in different environments, improves the overall performance of instance segmentation in complex scenarios, and realizes real-time individual segmentation and health monitoring of broiler with low power consumption and low latency.
Smart Images

Figure CN120070897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to a method for segmenting broiler instances based on the fusion of thermal imaging and RGB. Background Art
[0002] With the continuous expansion of the scale of broiler farming, traditional manual monitoring methods can no longer meet the requirements of modern production for accuracy and efficiency. Individual broiler identification, behavior analysis, and health monitoring have become important components of intelligent management. However, in high-density, complex background, and dynamic environments, existing deep learning technologies still face many challenges.
[0003] 1. Challenges in broiler instance segmentation: In high-density chicken flocks, occlusion and overlap between chickens, as well as complex backgrounds (such as feed troughs, straw, etc.), make it difficult for existing object detection and instance segmentation algorithms to effectively distinguish individual chickens. Although new network architectures (such as Mask2Former) have advantages in processing global information, they still face difficulties in capturing detailed features. In addition, traditional RGB images provide limited information in low-light and occlusion scenarios, resulting in insufficient segmentation accuracy. 2. Potential of multimodal data fusion: Thermal imaging technology can provide additional data that cannot be obtained from RGB images by capturing temperature information, especially in low-light and occlusion environments, which helps to improve the accuracy of target recognition. However, in the current broiler farming scenario, the technical research on combining thermal imaging and RGB images is still scarce, and their complementary advantages have not been fully exploited. 3. Edge computing and real-time processing requirements: Intelligent monitoring systems in agricultural fields usually need to run on edge devices, which requires the segmentation model to not only have high accuracy but also low latency and low power consumption. Existing deep learning models have high computational resource requirements and are difficult to meet the requirements of real-time inference and efficient processing.
[0004] Therefore, how to combine thermal imaging and RGB images to improve the accuracy and robustness of broiler instance segmentation through deep learning, especially in high-density, occluded, and complex backgrounds, and solve the deficiencies of existing methods in environmental adaptability, real-time performance, and segmentation accuracy is an issue that needs to be considered currently. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for segmenting broiler instances based on the fusion of thermal imaging and RGB, which solves the deficiencies existing in the prior art.
[0006] The purpose of the present invention is achieved through the following technical solutions: A method for segmenting broiler instances based on the fusion of thermal imaging and RGB, the segmentation method comprising: Step 1: Use an RGB visible light camera and a thermal imager as acquisition devices to obtain image data in two modalities, RGB and thermal imaging. According to the annotation strategy, annotate the acquired image data to obtain bimodal annotations of the same instance, and perform data preprocessing to generate annotated training data for the network model to learn. Step 2: Use the annotated training data after data preprocessing as the training set and input it into the constructed network model for training to obtain an optimized network model. Step 3: Re-input the acquired image data into the optimized network model, extract the feature information of the RGB and thermal imaging modalities in the image, fuse the feature information of the two modalities through a cross-modal attention mechanism, optimize the cross-modal information interaction by combining a deformable attention mechanism and multi-source position encoding, and finally generate the final broiler instance segmentation mask through a decoder.
[0007] The construction of the network model in Step 2 specifically includes the following: A1. Use a dual-branch MobileViT as the visual backbone structure to receive RGB images and thermal images respectively. The RGB branch receives the input of a three-channel RGB image, and the thermal branch receives the input of a single-channel thermal image, and extracts multi-scale feature maps through hierarchical convolution and Transformer. A2. After extracting the feature maps from each MobileViT branch, connect a CBAM module for feature recalibration, purify and strengthen the RGB and thermal feature maps at each scale to highlight the information related to the chicken body. A3. Construct a feature pyramid structure, upsample the high-level semantic features and fuse them with the low-level detail features to obtain rich and high-resolution feature maps, and introduce an ASPP module on the highest-level semantic features to obtain a feature representation with both local details and global context, better distinguishing adjacent broiler individuals and handling complex backgrounds. A4. Achieve deep fusion at the feature level through a cross-modal attention fusion module.
[0008] In the step of A2, for the RGB branch, the CBAM module adaptively assigns weights according to the color and edge features of the broiler target, making the channels related to the chicken body more activated and focusing on the position area where the chicken exists in the image. For the thermal branch, the CBAM module highlights the features of the high-temperature area according to the temperature intensity and weakens the interference of the cold background area.
[0009] The step of A4 specifically includes the following: A41. On the one hand, use the RGB feature as the query, and the thermal feature For keys and values, calculate heat-guided RGB features. On the other hand, using the heat features as queries and the RGB features as keys and values, calculate vision-guided heat features; A42. Introduce a deformable attention mechanism for the query position in the RGB image to learn a set of reference points and their offsets to calculate the enhanced features that fuse the semantic information of the heat features at the query position in the RGB image as where is the number of sampled keys for each query, represents the previous offset sampling position on the heat features, is the corresponding attention weight. Through the deformable attention mechanism, the RGB features selectively focus on a few points adjacent to or semantically related to their own positions on the thermal image to achieve cross-modal alignment; Similarly, for the query position in the thermal map to learn a set of reference points and offsets and calculate the attention only at these specific positions to obtain the enhanced features that fuse the semantic information of the RGB features at the query position in the thermal image as and fuse the two to obtain the fused feature map ; A43. And make full use of the information of both the RGB and thermal imaging modalities through multi-source position encoding to enhance the representation ability of the feature map, provide the combination of spatial coordinates and temperature information, and enable the network model to more accurately match and aggregate the corresponding features based on the dual clues of space and temperature; A44. Feed the fused feature map into the mask generation Transformer decoder of Mask2Former to obtain the final segmentation result through the decoder.
[0010] The steps of A43 specifically include the following: For each pixel position in the heat feature image the original two-dimensional sine position encoding is denoted as and the heat intensity value at this position is encoded into a vector and through multi-source position encoding, is added element-wise or concatenated with and then projected to form a new position encoding When calculating the attention, add the new position encoding to the corresponding key and value representations. In this way, when calculating the correlation, the attention mechanism not only considers the pure relative position but also perceives the distribution of the heat source intensity; Similarly, for the RGB feature map, first obtain a standard two-dimensional sine position encoding PE xᵧ is used to represent the spatial information of this position. At the same time, the RGB color value of this position is extended into a vector with the same dimension as the model through linear transformation to form the visual information encoding PE x ᵧ^RGB, and then for PE x ᵧ and PE x ᵧ^RGB are element-wise added or concatenated and then projected to generate the mixed position encoding PE x ᵧ^multi. This encoding is added to the key and value representations in the attention calculation of the Transformer, so that when the attention mechanism calculates the feature correlation, it can not only capture the accurate spatial position information but also obtain rich visual details.
[0011] The steps of A44 specifically include the following content: The fused feature map is used as the input of the key and value of the Transformer decoder, and a global multi-source position encoding is given to the query. In the decoder, each query draws information from the fused feature map through multi-head attention and continuously updates its own representation vector. Since the fused feature contains RGB and thermal information, the query will automatically focus on the regions with both visual and thermal consistency in the attention, that is, the positions where the broiler chickens exist. After several layers of iterative cross-attention and self-attention, each query will finally condense into a feature vector representing an instance; The decoder outputs a mask vector with the same size as the feature map for each query, takes the dot product of it with the fused feature map to generate mask predictions, and outputs the corresponding categories. The finally obtained query masks are processed through thresholding and non-maximum suppression to remove redundancy and retain the segmentation results of each broiler chicken.
[0012] The training and optimization of the network model include: The input includes aligned RGB image and thermal image pairs, as well as the ground truth mask of instance segmentation corresponding to each image into the constructed network model, and the network parameters are initialized. The Adam optimizer or SGD is used for iterative optimization, the initial learning rate is set, and the cosine annealing or multi-step descent strategy is used to adjust the learning rate to promote convergence. During the training process, the total loss function of the network model is constructed by the segmentation loss function, the temperature-guided consistency loss function, and the edge refinement loss function. The weights of the three loss functions are set respectively according to the training performance during the training process, so that the converged network model has the segmentation performance of high pixel accuracy, thermal consistency, and good edge contours at the same time.
[0013] After the network model is trained and reaches the set accuracy, it is also necessary to perform compression and acceleration optimization on the network model, which specifically includes the following content: Calculate the importance scores of the convolutional filters or attention heads of each layer of the MobileViT backbone, remove channels with small contributions, and reduce the number of queries in the Transformer decoder; Perform quantization-aware training or post-training quantization on the trained network model to improve the execution efficiency of the network model.
[0014] The RGB visible light camera and the thermal imager are used as acquisition devices to acquire image data in two modalities, RGB and thermal imaging. The two-modal annotation of the same instance obtained by annotating the acquired image data according to the annotation strategy includes: Use the RGB visible light camera and the thermal imager as acquisition devices. The two sensors are fixedly installed on the bracket above the broiler chicken house. The RGB visible light camera faces the ground and covers the chicken activity area from a top-down perspective to ensure that the whole body contour of the broiler chicken can be completely captured in the field of view. The RGB visible light camera is used to obtain visible light color images, and the thermal imager is used to synchronously obtain the thermal radiation intensity map of the same scene. The two acquisition devices ensure the field of view overlap through external calibration, and their relative positions and angles are calibrated so that the obtained RGB image and thermal image correspond spatially; Data acquisition is carried out in the form of video streams or continuous image frames, and the acquisition frequency is set according to needs to obtain a large number of images containing individual broiler chickens. The RGB visible light camera outputs high-resolution color images, and the thermal imager outputs lower-resolution infrared intensity maps. The thermal map is adjusted to match the size of the RGB image through interpolation. The acquisition process covers different time periods and lighting conditions, as well as different distances, postures, and occlusion situations in the chicken flock to obtain diverse data; The acquired two-modal images need to be annotated to generate the ground truth required for training. For each frame of image, a bounding box is marked for each broiler chicken on the RGB image to determine the position range of each chicken. At the same time, the corresponding high-temperature area is determined in combination with the thermal image, and the contour of the high-temperature area corresponding to the broiler chicken body is outlined on the thermal map. Each bounding box in the RGB image is corresponded to the high-temperature area in the thermal map through ID matching to form the two-modal annotation of the same instance and generate the instance segmentation ground truth mask.
[0015] The data preprocessing includes the following contents: Projectively transform the thermal image to the same coordinate system and resolution as the RGB image. For each pair of RGB image and thermal image, crop according to the field of view overlap area to ensure that the sizes of the two output images are the same and the pixels correspond one by one. Then, perform RGB modality processing on the RGB image, perform thermal imaging modality processing on the thermal image, and then perform data augmentation processing; For each labeled broiler instance, within the range of its RGB image bounding box, the normalized temperature values are extracted from the corresponding heat map region. A binary mask is extracted according to the set threshold, and this mask is cropped with the range of the bounding box to obtain an approximate broiler contour. The binary mask region of each chicken thus generated will serve as the main ground truth for instance segmentation.
[0016] The present invention has the following advantages: 1. By deeply fusing the thermal image with the RGB image, the complementary nature of the two modal data can be fully utilized, improving the adaptability of the system in different environments, thereby enhancing the overall performance of instance segmentation in complex scenarios.
[0017] 2. By introducing a lightweight vision Transformer (MobileViT) network and a deformable attention mechanism (Deformable Attention), while effectively reducing the computational complexity, the segmentation accuracy and adaptability of the model are improved. It can achieve real-time individual broiler segmentation and health monitoring on low-power devices, meeting the high requirements of agricultural intelligent monitoring systems for computational efficiency and real-time performance.
[0018] 3. The method for broiler instance segmentation based on the fusion of thermal imaging and RGB images can not only solve the segmentation accuracy problem of current technologies in high-density chicken flocks and complex background environments, but also further broaden its application scope in agricultural intelligent management. By real-time monitoring and analyzing the behavior of chicken flocks, this technology can be used for various applications such as broiler health monitoring, individual tracking, and behavior analysis, promoting the development of agricultural breeding towards a more intelligent and automated direction. Description of the Drawings
[0019] Figure 1 is a schematic flow chart of the present invention; Figure 2 is a schematic flow chart of the network structure design of the present invention. Detailed Embodiments
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application provided below in conjunction with the drawings in this application is not intended to limit the protection scope of the present application claimed, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present application. The present invention will be further described below in conjunction with the drawings.
[0021] As Figure 1 shown, the present invention specifically relates to a broiler instance segmentation method based on thermal imaging and RGB fusion, which solves the limitations of existing broiler instance segmentation technologies in complex agricultural environments, especially the recognition accuracy and real-time processing capabilities in the face of high density, occlusion, and complex background conditions. The specific content is as follows.
[0022] Step 1: Data acquisition; First, multi-modal image data is acquired. The present invention uses an RGB visible light camera and a thermal imager as the acquisition devices. The two sensors are fixedly installed on a bracket above the broiler breeding house. The camera faces the ground and covers the chicken flock activity area from a top-down perspective (e.g., vertically downward or slightly tilted at about 45 degrees) to ensure that the full body contours of the broilers can be completely captured within the field of view. The RGB camera is used to obtain visible light color images, and the thermal imager is used to synchronously obtain the thermal radiation intensity map (temperature heat map) of the same scene. The two imaging devices ensure field of view overlap through external calibration, and their relative positions and angles are calibrated so that the obtained RGB images and heat maps correspond spatially (e.g., by photographing a calibration board to calculate the mapping matrix and reprojection of the heat map into the RGB image coordinate system to achieve bimodal alignment).
[0023] The data acquisition is carried out in the form of video streams or continuous image frames, and the acquisition frequency can be set as needed, such as 5 to 30 frames per second, so as to obtain a large number of images containing individual broilers. Typically, the RGB camera outputs high-resolution color images (e.g., 1280×720 pixels), and the thermal imager outputs lower-resolution infrared intensity maps (e.g., 640×480 pixels). The heat map can be adjusted to match the size of the RGB image through interpolation. To obtain diverse data, the acquisition process should cover different time periods and lighting conditions (daytime, nighttime with infrared supplementary lighting), as well as different distances, postures, and partial occlusion situations among the chicken flock.
[0024] Annotation Strategy: The collected multi-modal images need to be manually annotated to generate the ground truth required for training. For each frame of the image, the annotator marks the bounding box for each broiler on the RGB image to roughly determine the position range of each chicken, and at the same time combines the thermal image to determine its corresponding high-temperature area. On the thermal map, the body of each chicken usually shows a temperature higher than the ambient background, so the temperature hot spot area of the broiler can be marked manually or semi-automatically. Specifically, the annotator outlines the high-temperature area corresponding to the body of the broiler on the thermal map (such as the hot areas of the head and trunk). By ID matching, each bounding box in the RGB image is corresponded to the high-temperature area in the thermal map to form a multi-modal annotation of the same instance. Based on the above annotation, a more refined instance segmentation ground truth mask can be further generated: for example, the pixels within the bounding box and belonging to the hot area are determined as chicken body pixels, so as to obtain the approximate binary mask contour of each chicken. This annotation strategy uses the position range provided by the RGB image and the shape clues provided by the thermal map to provide multi-modal supervision signals for subsequent model training.
[0025] Step 2. Data Preprocessing; Before using the collected data for model training, it is necessary to preprocess and augment the data of both RGB and thermal imaging modalities to improve the training effect and the generalization ability of the model.
[0026] First, align and crop the dual-modal images. Using the calibration results in Step 1, project the thermal image to the same coordinate system and resolution as the RGB image. For each pair of RGB-thermal images, crop according to the field of view overlap area to ensure that the sizes of the two output images are the same and the pixels correspond one by one. The pixels at the same position now contain the RGB three-channel values and the corresponding thermal intensity values, achieving precise spatial alignment.
[0027] Then, perform respective format transformation and normalization processing on the RGB image and the thermal image, specifically including: RGB Modal Processing: Scale or crop the RGB image to the unified size required by the neural network (such as 512×512 pixels). Perform color space conversion (such as BGR to RGB) and normalization. The pixel values of each channel are subtracted by the training set mean and divided by the standard deviation, or simply normalized to the [0,1] interval. For RGB images, traditional data augmentation strategies can be applied, such as color jitter (randomly adjusting brightness, contrast, saturation), random blur, etc., but care must be taken to ensure that the correspondence with the thermal map is not damaged.
[0028] Thermal Imaging Modal Processing: Convert the original temperature values of the thermal image into an intensity map available for the model. Such as adopting a linear normalization strategy: assuming the original thermal image gray scale range is (corresponding to a certain temperature range), then for each pixel temperature value Normalize: , and map it to the range of 0 to 1. The pixel values of the thermal image after such processing reflect the relative temperature levels and are similar to the pixel value range of the RGB image, facilitating the network to process them simultaneously. In addition, the single-channel thermal intensity map can be copied or extended to three channels (grayscale three channels) to adapt to certain network input formats, or a single-channel input layer can be specifically designed in the subsequent model structure. Appropriate filtering and denoising (such as median filtering to smooth thermal noise) and histogram equalization (to enhance temperature contrast) can also be considered in the thermal map preprocessing to highlight the temperature difference area between the broiler and the background.
[0029] In terms of data augmentation, it is particularly important to maintain the multimodal correspondence. Any spatial transformation-based augmentation operation must be applied to both the RGB and thermal maps simultaneously to ensure that they are still aligned. Common spatial augmentations include the following methods: Random flipping and rotation: Flip the image pair (horizontally) with a certain probability or rotate it within a small angle range, and transform the corresponding annotation masks simultaneously. This helps to eliminate directional biases.
[0030] Random translation and scaling: Slightly translate or scale the image (fill in the blank areas), and adjust the annotations accordingly. Enhance the robustness under different distances and positions.
[0031] Mosaic augmentation: Stitch multiple images together to form a new training sample. For example, randomly select 4 RGB-thermal image pairs from different scenarios, and after scaling and cropping, tile them into a 2×2 grid large image. In a typical Mosaic operation, first randomly determine the stitching center point position (relative to the size of the output composite image), and then: Place the first image in the upper left area of the composite image , place the second image in the upper right area , place the third image in the lower left area , and place the fourth image in the lower right area (where is the size of the output composite image). For adaptation, each original image is scaled according to the size of the divided area. After copying the image pixels to the corresponding quadrants of the composite image, the corresponding thermal map quadrants are synthesized in the same way. The coordinates of the annotations (bounding boxes, masks, etc.) of the original images are also shifted by the corresponding offsets and mapped to the coordinate system of the new image. Through Mosaic augmentation, multiple broiler instances from different scenarios can be seen simultaneously in a single frame, enriching the background and target combinations and improving the model's adaptability to crowded scenarios.
[0032] Other enhancements: On the premise of ensuring bimodal alignment, perturbations such as blur and noise are additionally applied to the RGB images, and temperature offsets or noise interferences can be simulated for the thermal images (such as adding or subtracting a small random amount to the heat values or local blur), enhancing the model's robustness to sensor noise.
[0033] Finally, training labels also need to be generated in the preprocessing stage for the model to learn. According to the annotation strategy in Step 1, the RGB bounding boxes are combined with the thermal region annotations to generate the ground truth of the segmentation mask for each instance.
[0034] Furthermore, for each annotated broiler instance, within the range of its RGB bounding box, the normalized temperature values are extracted from the corresponding thermal map region, and a binary mask is extracted according to a set threshold (for example, pixels with normalized temperature > 0.5 are regarded as the chicken body region). This mask is cropped to the range of the bounding box to obtain an approximate broiler contour. When necessary, the thermal region mask can be dilated through morphological operations (such as dilating by a certain number of pixels) to cover the entire chicken body contour. The binary mask region of each chicken thus generated will be used as the main ground truth for instance segmentation. These preprocessed and enhanced data (aligned RGB images, thermal images, and corresponding instance mask labels) will be used for subsequent model training.
[0035] Step 3: Network structure design; As Figure 2 shown, in this step, a segmentation network structure for fusing thermal imaging and RGB information is designed. The overall architecture is improved based on the instance segmentation framework of Mask2Former (an image segmentation model), adopting a lightweight backbone network and multiple attention mechanisms, and constructing a two-stream lightweight backbone + attention + multi-scale feature extraction network, laying a solid foundation for subsequent multi-modal fusion and instance segmentation. While keeping the whole network structure concise, it ensures that the extracted features are highly sensitive and fully expressive for the target (broiler). The main design is as follows: (1) Design a two-branch MobileViT (a lightweight general vision transformer for mobile devices) backbone: To extract the features of RGB and thermal maps respectively and ensure the model is lightweight and efficient, a two-stream branch MobileViT is used as the visual backbone structure, which specifically includes the following: The RGB image and the thermal image are respectively input into two parallel MobileViT backbones. The RGB branch receives the input of a three-channel color image, and the thermal branch receives the input of a single-channel thermal image. To adapt MobileViT to the processing of thermal maps, we modified the input layer of MobileViT - taking the average of the convolutional kernels of the RGB pre-trained weights on three channels for the initial weights of the thermal channel, so that it can directly accept the input of a 1-channel thermal map. The MobileViT network structures of the two branches are similar and can share the overall architecture hyperparameters, but do not share weights, so as to learn optimal features for different modalities respectively. Each MobileViT backbone extracts multi-scale feature representations through hierarchical convolutional and Transformer modules, making the final output feature map have global perception ability while the model is still very compact. By using the MobileViT backbone, the number of model parameters and the computational amount are significantly reduced compared with traditional large backbones (such as ResNet), laying a foundation for real-time operation on edge devices.
[0036] (2) Multi-path CBAM (an attention module that combines channel and spatial attention) attention fusion: Introduce the multi-path CBAM attention mechanism in the dual-branch backbone to strengthen the features of each modality. After extracting the intermediate feature maps from each MobileViT branch, connect the CBAM module for feature recalibration (this constitutes the "multi-path" - that is, insert CBAM at multiple stages of each branch of RGB and thermal).
[0037] For the RGB branch, CBAM adaptively assigns weights according to features such as the color and edges of the broiler target, making the channels related to the chicken body more activated and focusing on the position area where there are chickens in the image; for the thermal branch, CBAM highlights the features of the high-temperature area according to the temperature intensity and weakens the interference of the cold background area. Through the multi-path (multi-level, dual-branch) CBAM processing, the feature maps of RGB and thermal are purified and strengthened at each scale. This attention mechanism ensures that before entering the modality fusion, the features of each modality have highlighted the information related to the chicken body (such as: RGB features highlight the outline and texture of the chicken, and thermal features highlight the temperature hotspots of the chicken), providing high signal-to-noise ratio feature inputs for subsequent fusion.
[0038] (3) Multi-scale feature extraction and ASPP (a deep learning technology for extracting multi-scale features) module: Considering that the size and distance of broilers in the image may be different, in order to improve the recognition of targets at different scales, the present invention incorporates a multi-scale feature extraction strategy at the backbone output stage.
[0039] Specifically, the MobileViT backbone outputs several layers of feature maps (e.g., at 1 / 4, 1 / 8, and 1 / 16 of the original image resolution respectively), constructs a simple FPN (Feature Pyramid Network) structure, upsamples the high-level semantic features and fuses them with the low-level detailed features to obtain rich and high-resolution feature maps for segmentation. Additionally, an ASPP (Atrous Spatial Pyramid Pooling) module is introduced on the highest-level semantic features to obtain a feature representation that combines local details and global context, better distinguishing adjacent broiler chickens and handling complex backgrounds. ASPP is applied to the top-level features of both the RGB and thermal branches to enhance the robustness of each modality to scale changes. The feature maps output by ASPP will serve as the main input features for subsequent cross-modal feature fusion and Transformer decoding.
[0040] Step 4: Introduce two cross-modal feature fusion mechanisms; After extracting the features of RGB and thermal imaging respectively, it is necessary to effectively fuse the information of these two modalities to utilize their complementary advantages. In this step, a cross-modal attention fusion module is designed to achieve deep fusion at the feature level. The fusion mechanism is based on the attention structure of Transformer (a deep learning model based on the attention mechanism), combined with deformable attention and multi-source positional encoding to optimize cross-modal information interaction, enabling the model to accurately distinguish each broiler chicken from the background and other individuals in various complex scenarios.
[0041] (1) Overview of feature alignment and fusion: Since after the processing in Step 3, the high-level feature maps output by the RGB branch and the thermal branch are consistent in spatial size (both restored to near the original image resolution or aligned through upsampling), and the pixel coordinates correspond. Therefore, assume that the RGB feature map and the thermal feature map are obtained, where is the feature map resolution and is the number of channels. The goal is to fuse these two feature maps into a comprehensive feature map that contains more discriminative information for broiler chicken instance recognition. Direct fusion methods such as concatenation or element-wise addition are difficult to fully utilize the complementarity of the two modalities. Therefore, a cross-modal attention fusion module is designed to adaptively combine the RGB and thermal features through learned attention weights.
[0042] (2) Cross-modal attention mechanism: The cross-modal attention adopts the query-key-value operation idea of Transformer, using the features of one modality as the query to retrieve the relevant information in the features of the other modality as the key / value, thereby achieving information interaction.
[0043] In the specific implementation, bidirectional cross-modal attention is adopted. On the one hand, taking the RGB features as the queries, the thermal features as the keys and values, the heat-guided RGB features are calculated. On the other hand, taking the thermal features as the queries and the RGB features as the keys and values, the vision-guided thermal features are calculated. Then the two are combined to form the final fused features. The specific contents are as follows: Heat-guided RGB attention: Let , first obtain the query matrix, key matrix, and value matrix through linear mapping: , where is a learnable projection matrix.
[0044] When focusing on the RGB query at the th spatial position, its attention weight is calculated by the dot product of the query vector at this position and the key vectors at all positions of the thermal features: (1), where represents the calculated attention weight, is the query vector extracted from the RGB features, representing the RGB features at the i-th position (e.g., the feature representation of a certain pixel point in the RGB image), is the key vector extracted from the thermal image features, representing the thermal map features at the thermal pixel position j, j represents the current thermal image position to be calculated, represents traversing all positions of the thermal image. Then calculate the enhanced features of the RGB image at the query position : (2), where is the value at the j th position in the thermal map, that is, the value of the thermal map features, which is to weight and converge the part with the highest correlation with the position in the thermal features . The obtained can be understood as "RGB features combined with thermal information".
[0045] Vision-guided thermal attention: Similarly, by swapping roles and taking the thermal features as the queries and the RGB features as the key values, the enhanced features after fusing the RGB semantics for each thermal map position can be obtained: (3), where is based on the thermal query and the RGB key The calculated attention weights are the features at the j th position in the RGB feature map.
[0046] After obtaining and , the feature maps obtained by fusing the two are . The fusion method can adopt channel concatenation + convolution or weighted summation, etc. For example, the and vectors at the corresponding positions can be concatenated and then mapped back to dimensions through convolution, or simply added ( is a learnable weight, initially set to 0.5). In this way, each position contains clues from both RGB and thermal: visual texture and contours as well as thermal intensity information.
[0047] (3) Introduce deformable attention: To improve the efficiency and robustness of attention fusion, a deformable attention mechanism is introduced in the above cross-modal attention. Standard global attention needs to traverse all keys for each query, with a large computational amount and strict assumptions of spatial alignment. Deformable attention allows each query to only focus on a limited number of sampled points and can learn the offsets of these points, thus flexibly corresponding to the position deviations of cross-modal features. Specifically, for a query position of the RGB feature, instead of summing over all positions of the heat map, a set of reference points (e.g., 5) and their offsets are learned, and attention is only calculated at these sampled positions. Then formula (2) is modified to: (4), where is the number of sampled keys for each query, represents the offset sampled position on the thermal feature, is the corresponding attention weight (obtained by calculating softmax normalization at these positions according to the aforementioned method). Through this mechanism, the RGB feature can selectively focus on a few points on the heat map that are adjacent or semantically related to its own position, achieving cross-modal alignment. For example, when the contour of a certain chicken in the RGB image is unclear due to visible light, its query can flexibly "find" the local high-temperature points on the heat map, even if they are not at the exact same pixel position due to parallax. With deformable attention, cross-modal fusion is both efficient (reducing computational complexity) and robust (allowing a certain position offset).
[0048] Similarly, for each query position j of the heat map, a set of reference points and offsets , calculate the attention only at these specific positions. Then formula (3) is modified to: (5), In this way, both the heatmap features and the RGB features can be flexibly aligned through the deformable attention, improving the robustness of cross-modal fusion.
[0049] The fused feature map is obtained after the cross-modal attention mechanism and the deformable attention mechanism. Specifically, the RGB and heatmap features are first fused through the cross-modal attention mechanism, and then further aligned and adjusted through the deformable attention. This feature map will be used as the input of the subsequent mask generation decoder for broiler instance segmentation.
[0050] (4) Multi-source position encoding: mainly used to enhance the representation ability of the feature map after feature map fusion, providing the combination of spatial coordinates and temperature information; Transformer attention usually needs to add position encoding to inject spatial position information. In this scheme, we design multi-source position encoding to make full use of the information of RGB and thermal dual modalities. Specifically, a hybrid position encoding that considers both spatial coordinates and temperature values is introduced into the fusion module. For example, for each pixel position of the heat feature image, the original two-dimensional sine position encoding is denoted as (including the information of the coordinates); at the same time, we can encode the heat intensity value at this position into a vector (for example, by linearly transforming the scalar temperature value to a vector with the same dimension as the model, or using the high-frequency sine encoding of the heat value). The multi-source position encoding fuses the two, such as adding or concatenating and element-wise and then projecting to form a new position encoding . During attention calculation, this encoding is added to the corresponding key and value representations. In this way, the attention mechanism not only considers the pure relative position when calculating the correlation, but also perceives the distribution of the heat source intensity.
[0051] For the RGB feature map, each pixel position (x, y) first obtains a standard two-dimensional sine position encoding PE x ᵧ to represent the spatial information of this position; at the same time, the RGB color value (or other visual features) at this position is linearly transformed into a vector with the same dimension as the model to form the visual information encoding PE x ᵧ^RGB; then, perform element-wise addition or concatenation on PE x ᵧ and PE x ᵧ^RGB and project to generate the hybrid position encoding PEx ᵧ^multi, which is added to the key and value representations in the attention calculation of the Transformer, enabling the attention mechanism to capture both precise spatial location information and rich visual details when calculating feature correlations.
[0052] For example, even if two pixels are far apart in the image, if they both have high-temperature features, the thermal encoding part will make them more relevant in the attention. This extension of the positional encoding incorporating temperature information is similar to the idea of embedding depth into the positional encoding in the DepthFormer method for multi-modal Transformers, enabling the Transformer attention to naturally fuse bimodal information without adding extra parameters. Therefore, the multi-source positional encoding works in tandem with cross-modal attention, enabling the model to more accurately match and aggregate corresponding features based on dual cues of space and temperature.
[0053] (5) Mask Transformer decoding: The fused features are then fed into the mask generation Transformer decoder of Mask2Former. The decoder consists of a set of learnable query embeddings (each query is intended to correspond to a potential broiler instance) and multiple layers of mask attention blocks.
[0054] The fused feature map is used as the input for the key and value of the Transformer decoder, and global multi-source positional encoding is assigned to the queries. In the decoder, each query extracts information from the fused features through multi-head attention and continuously updates its own representation vector. Since the fused features contain RGB and thermal information, the queries will automatically focus on regions with both visual and thermal consistency in the attention - which is exactly where the broilers are located. After several layers of iterative cross-attention and self-attention, each query will eventually converge into a feature vector representing an instance. Subsequently, these queries are mapped to mask segmentation results and class outputs through a linear layer.
[0055] Specifically, the decoder outputs a mask vector of the same size as the feature map for each query, which is dot-producted with the fused feature map to generate a mask prediction (the segmentation probability map corresponding to this query), and outputs the corresponding class (in this application, the class may only be "broiler"). Since we adopt the mask attention strategy, the attention mechanism is also utilized during the decoding process to directly generate masks on the feature map without the need for prior detection and then segmentation. Finally, the resulting query masks are processed through thresholding and non-maximum suppression (NMS) to remove redundancy and retain the segmentation results of each broiler.
[0056] Step 5, Model training and optimization; After completing the model architecture design, it is necessary to effectively train and optimize the model. For this purpose, a new multi-modal fusion loss function is constructed, which combines the objectives of segmentation accuracy, temperature consistency, and edge refinement to perform end-to-end training on the model. This step will detail the training strategy and loss function design.
[0057] (1) Use the annotated training data preprocessed in step 2 as the training set and input it into the model. The training is carried out in an end-to-end manner, and the input includes the aligned RGB image and thermal image pair, as well as the instance segmentation ground truth mask (and class label) corresponding to each image. Initialize the network parameters: The MobileViT backbone can use the weights pre-trained on ImageNet (RGB branch). Since there is no pre-training for the thermal branch, it can be randomly initialized or partially borrow the RGB pre-trained parameters; the Transformer decoder and fusion module are randomly initialized. Use the Adam optimizer or SGD for iterative optimization. The initial learning rate is set according to the model size (e.g., 1e-4), and the cosine annealing or multi-step descent strategy is used to adjust the learning rate to promote convergence. The entire training is performed on the GPU for several epochs until the segmentation performance on the validation set converges and improves.
[0058] (2) Multi-task loss function: During the training process, each batch of input passes through the model to obtain the predicted broiler instance mask and the corresponding class. The total loss function designed in the present invention consists of the following three parts.
[0059] (a) Segmentation main loss : Measures the difference between the predicted instance mask and the ground truth mask, and is the most basic supervision signal. Since this task belongs to instance segmentation, it can be regarded as a pixel-level binary classification (foreground: chicken, background) problem. A combination of binary cross-entropy loss and Dice loss is selected to measure the segmentation effect, so as to take into account both pixel accuracy and overall contour matching. For a single predicted mask (representing the probability that pixel is the foreground) and the corresponding binary ground truth mask , the cross-entropy loss function is defined as: (6), is the total number of pixels in the image, and cross-entropy emphasizes the accurate classification of each pixel.
[0060] The Dice loss function is defined as: (7), This loss is based on the Dice coefficient and measures the overlap degree between the predicted mask and the ground truth mask. The smoothing term prevents the denominator from being zero. When the prediction and the ground truth highly overlap, the Dice loss approaches zero.
[0061] Furthermore, the main segmentation loss function is expressed as the weighted sum of the two: (8), where is the weight coefficient (which can be taken as about 0.5 for balance respectively). Designed in this way, it not only ensures pixel-level accuracy but also pays attention to the overall contour quality, which is commonly used and effective in the instance segmentation task.
[0062] (b) Temperature-guided consistency loss : This loss term utilizes the thermal imaging information to ensure the consistency between the model prediction and the heat map signal. Intuitively, it is expected that the predicted chicken body mask covers the high-temperature area on the heat map, and there should be no false alarms in the background (low-temperature area). The temperature distribution provided by the thermal imaging is transformed into a soft target to supervise the segmentation output. The specific approach is as follows: Threshold the normalized value of the heat map to obtain the binary ground truth mask of the hot area as: (9), where is the temperature threshold (which can be set according to the environmental temperature experience, such as 0.5 or adaptively adjusted according to the average background temperature). denotes that pixel belongs to the high-temperature area (most likely part of the chicken), denotes the low-temperature background. Then, the temperature consistency loss is defined as the pixel binary classification error between the predicted probability map and , and the temperature-guided consistency loss is obtained as: (10), This is actually treating as an additional supervision mask and calculating the cross-entropy loss with the prediction. If is directly used for soft labels (non-binary), the KL divergence can also be used to measure the closeness of the and distributions. Through , the model will be penalized for the error of predicting the background in the high-temperature area (miss detection), and also for the misdetection of the foreground in the cold background area. In other words, the temperature consistency loss guides the model to learn the rule of "where there is heat, there is a chicken; where it is cold, there is no chicken", making the prediction result match the infrared thermal pattern. This ensures that the model does not miss real broiler chickens even under lighting changes or color camouflage conditions (when RGB is unreliable).
[0063] (c) Edge refinement loss : To further improve the contour quality of the segmentation mask, an edge refinement loss is introduced to specifically handle the accuracy of the object boundary. Usually, in the segmentation model, the mask edge may not be smooth enough or misaligned with the ground truth boundary. This loss is defined by comparing the edges of the ground truth mask and the predicted mask. First, calculate the edge map of the ground truth mask through an edge extraction operator (such as Sobel or Canny operator) and the edge map of the predicted probability map and . It can be represented as a 0-1 binary map, where 1 indicates that the pixel is on the mask boundary. Then, define the edge refinement loss as the average of the absolute values of the difference between the two, and the edge refinement loss is obtained as: (11), When the boundary of the predicted mask perfectly matches the ground truth, the difference between the two is zero and the loss is the lowest; if there are deviations, misalignments or jaggedness in the boundary, the loss increases. By minimizing , the model is encouraged to generate predictions with the same shape as the ground truth boundary. For example, if the edge of the ground truth mask is a smooth curve while the prediction has jaggedness, the loss will drive the model to adjust to make its edge smoother and more continuous; if the predicted edge is overall shifted outward or inward, the loss will also drive the model to correct the shift. It should be noted that focuses on shape consistency and is insensitive to minor errors inside the pixel level, which complements (pixel accuracy) to jointly improve the quality of the final mask.
[0064] (3) Total loss and optimization: Combining the above three parts, the total training loss function is constructed as: (12), where is an adjustable weight coefficient used to balance the contributions of the three losses to the overall objective. Initially, the three can be set to 1 or tuned according to the performance on the validation set. For example, if it is found that the segmentation accuracy is sufficient but the edges are not fine enough, the weight can be appropriately increased. During the training process, taking as the objective, update the model parameters through backpropagation to gradually reduce each sub-loss. It is worth mentioning that due to the addition of and , the model may require more iterations to simultaneously optimize the multi-task objective in the initial stage of training. A staged training strategy can be adopted: first, train for several epochs only with to obtain the basic segmentation ability, and then gradually introduce and Enhance details. Of course, the full loss function can also be used simultaneously from the start of training. Through reasonable hyperparameters and training schedules, after convergence, the model will have segmentation performance with high pixel accuracy, thermal consistency, and good edge contours.
[0065] Step 6: Model compression and optimization; After the model training is completed and satisfactory accuracy is achieved, in order to deploy it to resource-constrained edge devices, the model needs to be compressed and accelerated. This step includes techniques such as model pruning and quantization to streamline the complex deep model into a more efficient form.
[0066] (1) Model pruning: Channel pruning is performed on the convolutional layers and attention layers of the MobileViT backbone: Calculate the importance scores of each layer's convolutional filters or attention heads (which can be based on the norm of the weights or sensitivity analysis to the loss), and then remove those channels with less contribution. Especially in the dual-branch structure, there may be feature redundancy between the RGB branch and the thermal branch, and some redundant channels can be removed to reduce the model size. In addition, appropriate pruning is also performed on the number of queries in the Transformer decoder to reduce the number of queries and thus reduce the computation. The pruning adopts a progressive pruning strategy: First, fine-tune the model with a small pruning rate, and gradually increase the pruning intensity through multiple iterations. After each pruning, fine-tune the model (fine-tune) to recover the accuracy.
[0067] (2) Model quantization: The 8-bit quantization (INT8 quantization) method is used to compress the model from 32-bit floating-point representation. The specific process includes: performing quantization-aware training (QAT) or post-training quantization (PTQ) on the trained model. Quantization-aware training is to continue training for several epochs after training, but simulate operations such as convolution and matrix multiplication in the model as 8-bit quantization (simulating quantization errors through scaling factors and zero points), so that the model parameters gradually adapt to the quantization errors to avoid accuracy loss. During the quantization process, special attention should be paid to calibrating the attention operations and normalization layers in MobileViT to ensure an appropriate numerical dynamic range. For the dual-branch, it is necessary to statistically analyze the activation value distributions of each layer in the RGB and thermal branches to select appropriate quantization scales. After completing QAT, the weights are solidified into the INT8 format. Through quantization, the model volume will be reduced by about 4 times, and the inference calculation can utilize the 8-bit integer instructions of the CPU / special accelerator, greatly improving the execution efficiency. At the same time, it is verified that the segmentation accuracy of the quantized model on the validation set only drops slightly and still meets the requirements.
[0068] (3)Other optimizations: Use the original model as the teacher network to train a smaller student network, enabling the student to learn the teacher's behavior in bimodal fusion, thereby further reducing the model complexity. Additionally, optimizing the convolution operations in MobileViT using grouped convolutions and optimizing the operators in the fusion layer (such as decomposing the multi-head attention into multiple small matrix multiplications) can also improve efficiency.
[0069] Step 7: Edge deployment and accelerated inference; Deploy the optimized model to the edge computing device, which is composed of a front-end bimodal camera sensor and a back-end embedded AI module as a whole. It can operate stably in the actual breeding environment, accurately identify and segment each broiler, providing a powerful tool for intelligent breeding management.
[0070] (1)Deployment platform selection: According to the on-site situation of the farm, select an edge device with sufficient performance to run the model in real time, such as a small embedded system equipped with NVIDIA Jetson series modules (TX2, Xavier, Orin, etc.), or other embedded AI devices that support GPU acceleration. The deployment platform installs a camera interface for accessing the RGB camera and thermal imager signals in Step 1, and has a model inference framework environment (such as TensorRT acceleration library, CUDA support, etc.).
[0071] (2)Model conversion and acceleration: Convert the pruned and quantized model into a deployment format, such as ONNX format, and then use TensorRT to optimize and compile the model. TensorRT is deeply optimized for the NVIDIA GPU platform. It fuses the operators in the computation graph, optimizes the memory, and uses FP16 / INT8 operations to improve the inference speed. We use TensorRT to optimize the convolution and attention layers of the model into efficient kernels and enable INT8 precision (combining the calibration table quantized in Step 6). In particular, TensorRT will expand and optimize the convolution and Transformer sub-modules of MobileViT to complete the calculation on the GPU with the least memory access. Additionally, optimize the input and output: Configure the inference with a fixed batch size (batch = 1) to reduce the dynamic overhead; pre-allocate the video memory buffer to store the input RGB / thermal map and the output mask. The engine file compiled by TensorRT can be directly loaded and run on the device to achieve maximum throughput.
[0072] (3)Real-time inference: On the edge device side, the running program continuously reads synchronized frames from the dual cameras, performs preprocessing on each pair of RGB and thermal images (according to the process in step 2: alignment, normalization, etc.), and then sends the processed tensors into the TensorRT engine for forward inference. The model outputs the segmentation masks (and class labels) of all broiler instances in the frame. Slight post-processing is performed on the output masks: removing small regions (removing noise fragments of a few pixels), smoothing the edges (correcting the burrs using morphological smoothing or boundary tracking algorithms), and assigning IDs to each instance mask for continuous frame tracking, etc. Finally, the masks are drawn on the original image or output in the form of coordinates to achieve real-time detection and contour marking of broilers.
[0073] (4)Performance metrics: This invention not only focuses on inference speed and accuracy but also needs to evaluate the actual performance of the system in the actual breeding environment. The following are the targeted performance evaluation metrics.
[0074] Average inference latency: Calculate the processing time for each input RGB and thermal image pair and report the average processing time per frame.
[0075] (13), L For the number of frames to be evaluated, taking Jetson Xavier as an example, through actual testing, the single-frame inference latency is about 40 milliseconds, which is equivalent to a processing speed of more than 25 frames per second. For the input dual-modal image with a resolution of , it can meet the real-time requirements of farm video monitoring (the general video frame rate is 25fps). Even on the relatively low-end TX2 platform, after INT8 optimization, it can reach about 15fps, basically meeting the real-time monitoring requirements. The model memory occupancy and power consumption are also controlled within an acceptable range: the size of the quantized model is only dozens of MB, the GPU occupancy rate during inference is moderate, and it can run stably for a long time. In actual deployment, the system can work continuously for 7×24 hours, perform broiler instance segmentation on the ingested image stream frame by frame, and obtain the position and contour information of each chicken in real time. This information can be further used for applications such as chicken counting, behavior analysis, and health status monitoring (combined with temperature anomaly detection).
[0076] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and improvements, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and alterations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A broiler instance segmentation method based on thermal imaging and RGB fusion, characterized by: The segmentation method comprises: Step 1: Use an RGB visible light camera and a thermal imager as acquisition devices to acquire image data in two modalities, RGB and thermal imaging. Annotate the acquired image data according to the annotation strategy to obtain bimodal annotations of the same instance, and perform data preprocessing to generate annotated training data for network model learning. Step 2: Input the labeled training data after data preprocessing as a training set into the constructed network model for training to obtain a trained and optimized network model; Step 3: Re-input the collected image data into the trained and optimized network model, extract the feature information of the two modalities of RGB and thermal imaging in the image, and fuse the feature information of the two modalities through the cross-modal attention mechanism, and combine the deformable attention mechanism and multi-source position encoding to optimize the cross-modal information interaction, and finally generate the final broiler instance segmentation mask through the decoder.
2. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 1 is characterized by: The construction of the network model in step 2 specifically includes the following contents: A1. Use the dual-branch MobileViT as the visual backbone structure to receive RGB images and thermal images respectively. The RGB branch receives three-channel RGB image input, and the thermal branch receives single-channel thermal image input. Multi-scale feature maps are extracted through hierarchical convolution and Transformer. A2. After extracting the feature map from each MobileViT branch, the CBAM module is connected to perform feature recalibration. The RGB and thermal feature maps are purified and enhanced at each scale to highlight the chicken body related information. A3. Build a feature pyramid structure, upsample high-level semantic features and fuse them with low-level detail features to obtain rich and high-resolution feature maps, and introduce the ASPP module on the highest-level semantic features to obtain feature representations that coexist with local details and global context, so as to better distinguish adjacent broiler individuals and handle complex backgrounds. A4. Realize deep fusion at the feature level through the cross-modal attention fusion module.
3. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 2 is characterized in that: In step A2, for the RGB branch, the CBAM module adaptively assigns weights according to the color and edge features of the broiler target, so that the channel activation related to the chicken body is stronger and focuses on the location area where the chicken exists in the image; For the thermal branch, the CBAM module highlights the characteristics of the high-temperature area according to the temperature intensity and weakens the interference of the background cold area. The high-temperature area refers to the red area in the thermal image.
4. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 2 is characterized in that: The step A4 specifically includes the following contents: A41. On the one hand, RGB features For query, hot features As keys and values, thermally guided RGB features are calculated. On the other hand, with thermal features as queries and RGB features as keys and values, visually guided thermal features are calculated; A42. Introducing a deformable attention mechanism for querying the position of RGB images , learn a set of reference points and its offset , calculate the RGB image at the query position The enhanced features of fusion thermal feature semantics are ,in is the number of sampled keys per query, represents the previous offset sampling position of the thermal feature, is the corresponding attention weight; Similarly, for the query location of the heat map , learn a set of reference points and offset , only calculate the attention at these locations, and get the thermal image query location The enhanced features of RGB feature semantics are: , the two are fused to obtain the fused feature map ; A43, and through multi-source position encoding to make full use of the information of RGB and thermal imaging modalities to enhance the representation ability of feature maps, provide a combination of spatial coordinates and temperature information, so that the network model can more accurately match and aggregate corresponding features based on dual clues of space and temperature; A44. Send the fused feature map to the mask of Mask2Former to generate the Transformer decoder, and get the final segmentation result through the decoder.
5. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 4 is characterized in that: The steps of A43 specifically include the following: For each thermal feature map pixel position , the original two-dimensional sinusoidal position coding is denoted as , the heat intensity value of this location Encoded as a vector , through multi-source position coding and Add or concatenate elements and project them to form a new position code ,When calculating the attention, the new position encoding is added to the corresponding key and value representations, so that the attention mechanism not only considers the pure relative position when calculating the relevance, but also perceives the distribution of the heat source intensity; Similarly, for the RGB feature map, first obtain a standard two-dimensional sinusoidal position encoding PE x ᵧ, used to represent the spatial information of the position. At the same time, the RGB color value of the position is expanded into a vector with the same dimension as the model through linear transformation to form the visual information encoding PE x ᵧ^RGB, then PE x ᵧ and PE x ᵧ^RGB is added or concatenated at the element level and then projected to generate a hybrid positional encoding PE x ᵧ^multi, this encoding is added to the key and value representations in the Transformer's attention calculation, so that the attention mechanism can capture both precise spatial location information and rich visual details when calculating feature relevance.
6. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 4 is characterized in that: The steps of A44 specifically include the following: The fused feature map is used as the key and value input of the Transformer decoder, and the query is given a global multi-source position encoding. In the decoder, each query draws information from the fused feature map through multi-head attention and continuously updates its own representation vector. Since the fused feature contains RGB and thermal information, the query will automatically focus on the area with both visual and thermal consistency, that is, the location of the broiler. After several layers of cross-attention and self-attention iterations, each query will eventually condense into a feature vector representing an instance. The decoder outputs a mask vector of the same size as the feature map for each query, performs a dot product with the fused feature map to generate a mask prediction, and outputs the corresponding category. The final query mask removes redundancy through threshold processing and non-maximum suppression, retaining the segmentation results of each broiler.
7. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 1 is characterized in that: The training and optimization of the network model includes: The input includes aligned RGB images and thermal image pairs, as well as the instance segmentation truth mask corresponding to each image into the constructed network model, and the network parameters are initialized. The Adam optimizer or SGD is used for iterative optimization, the initial learning rate is set, and the learning rate is adjusted using cosine fire or multi-step descent strategy to promote convergence. During the training process, the total loss function of the network model is constructed by the segmentation loss function, the temperature-guided consistency loss function, and the edge refinement loss function. The weights of the three loss functions are set separately according to the training performance during the training process, so that the converged network model has high pixel accuracy, thermal consistency, and good edge contour segmentation performance.
8. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 1 is characterized in that: After the network model is trained and reaches the set accuracy, it is also necessary to compress and accelerate the network model, including the following: Calculate the importance score of each convolution filter or attention head of the MobileViT backbone, remove channels with small contributions, and reduce the number of queries of the Transformer decoder; Perform quantization-aware training or post-training quantization on the trained network model to improve the execution efficiency of the network model.
9. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 1, characterized in that: The method of using an RGB visible light camera and a thermal imager as acquisition devices to acquire image data in two modes, RGB and thermal imaging, and annotating the acquired image data according to an annotation strategy to obtain a two-modal annotation of the same instance includes: RGB visible light camera and thermal imager are used as acquisition equipment. The two sensors are fixed on the bracket above the broiler breeding house. The RGB visible light camera faces the ground and covers the chicken activity area from a bird's-eye view to ensure that the whole body outline of the broiler can be fully captured in the field of view. The RGB visible light camera is used to obtain visible light color images, and the thermal imager is used to synchronously obtain the thermal radiation intensity map of the same scene. The two acquisition devices are externally calibrated to ensure the overlap of the field of view. Their relative positions and angles are calibrated so that the obtained RGB image corresponds to the thermal image in space. Data collection is carried out in the form of video streams or continuous image frames. The collection frequency is set as needed to obtain a large number of images containing individual broiler chickens. The RGB visible light camera outputs high-resolution color images, and the thermal imager outputs lower-resolution infrared intensity images. The thermal map is adjusted to match the RGB image size through interpolation. The collection process covers different time periods and lighting conditions, as well as different distances, postures and occlusions in the flock to obtain diverse data; The collected bimodal images need to be annotated to generate the true values required for training. For each frame of the image, a bounding box is marked for each broiler on the RGB image to determine the position range of each chicken. At the same time, the corresponding high-temperature area is determined in combination with the thermal image. The contour of the high-temperature area corresponding to the broiler body is outlined on the thermal map. Each bounding box in the RGB image is matched with the high-temperature area in the thermal map through ID matching to form a bimodal annotation of the same instance and generate an instance segmentation true value mask.
10. The broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 1, characterized in that: The data preprocessing includes the following contents: The thermal image is projected and transformed to the same coordinate system and resolution as the RGB image. Each pair of RGB images and thermal images is cropped according to the overlapping area of the field of view to ensure that the two output images are of the same size and have one-to-one pixel correspondence. Then, the RGB image is processed in RGB mode, and the thermal image is processed in thermal imaging mode, and then data enhancement is performed. For each annotated broiler instance, the normalized temperature value is extracted from the corresponding heat map area within the bounding box of its RGB image, and a binary mask is extracted according to the set threshold. This mask is clipped with the range of the bounding box to obtain the broiler outline. The resulting binary mask area of each chicken will be used as the main truth value for instance segmentation.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and medium
CN110910304A
Road scene semantic segmentation method based on attention residual learning
CN111860411A
Low-precision image semantic segmentation method based on improved attention mechanism
CN112508960A
RGB-T image semantic segmentation method based on modal difference reduction
CN112991350A
RGBT target tracking method based on cross-modal attention mechanism and twin structure
CN113628249A
Cited By
Interactive modal-free instance segmentation method based on click perception shape prior
CN120953622A
Interactive modal-free instance segmentation method based on click-aware shape priors
CN120953622B
Dynamic contour attention-driven cross-modal fusion method and system
CN121010869A
Multi-frequency information fusion double-trunk neural network fruit identification method
CN121121202A
Beef cattle identification method based on multi-scale segmentation optimization and multi-modal data fusion
CN121616931A