A Broiler Instance Segmentation Method Based on the Fusion of Thermal Imaging and RGB

By combining RGB and thermal imaging data, using a lightweight network model to perform instance segmentation of broiler chickens, the segmentation accuracy problem in high density and complex backgrounds is solved, real-time individual segmentation and health monitoring of broiler chickens on edge devices is realized.

CN120070897BActive Publication Date: 2025-07-11SICHUAN AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510527451.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-11
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing broiler instance segmentation technology lacks segmentation accuracy under high density, occlusion and complex backgrounds, and the computing resources of deep learning models are high, making it difficult to meet real-time processing requirements.

Method used

Data was collected using an RGB visible light camera and thermal imager, features were extracted through a dual-branch MobileViT network, RGB and thermal imager were fused with CBAM module and a cross-modal attention mechanism, broiler instance segmentation mask was generated using a Mask2Former decoder, and the model was optimized through lightweight and quantization to adapt to edge devices.

Benefits of technology

It improves the accuracy and robustness of broiler instance segmentation, meets the needs of real-time and low power consumption, and broadens the application scope of agricultural intelligent management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070897B_ABST
    Figure CN120070897B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for instance segmentation of broiler chickens based on the fusion of thermal imaging and RGB, belonging to the field of image processing. The segmentation method includes: collecting image data and performing preprocessing; using the labeled training data after data preprocessing as a training set and inputting it into a constructed network model for training to obtain a network model optimized after training; re-inputting the collected image data into the network model optimized after training, extracting the feature information of two modalities, RGB and thermal imaging, in the image, fusing the feature information of the two modalities through a cross-modal attention mechanism, and optimizing cross-modal information interaction by combining a deformable attention mechanism and multi-source position encoding, and finally generating a final instance segmentation mask of broiler chickens through a decoder. The present invention realizes accurate segmentation and tracking of individual broiler chickens, and further improves the robustness and recognition ability of the model in a complex agricultural environment through the temperature information of thermal imaging images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a method for instance segmentation of broiler chickens based on the fusion of thermal imaging and RGB. Background Art

[0002] With the continuous expansion of the scale of broiler chicken farming, traditional manual monitoring methods can no longer meet the requirements of modern production for accuracy and efficiency. The individual identification, behavior analysis, and health monitoring of broiler chickens have become important components of intelligent management. However, in high-density, complex background, and dynamic environments, existing deep learning technologies still face many challenges.

[0003] 1. Challenges in instance segmentation of broiler chickens: In a high-density chicken flock, occlusion and overlap between chickens, as well as complex backgrounds (such as feeding troughs, straw, etc.), make it difficult for existing object detection and instance segmentation algorithms to effectively distinguish individual chickens. Although new network architectures (such as Mask2Former) have advantages in processing global information, they still face difficulties in capturing detailed features. In addition, traditional RGB images provide limited information in low-light and occlusion scenarios, resulting in insufficient segmentation accuracy. 2. Potential of multimodal data fusion: Thermal imaging technology can provide additional data that cannot be obtained from RGB images by capturing temperature information, especially in low-light and occlusion environments, which helps to improve the accuracy of target recognition. However, in the current broiler chicken farming scenario, the technical research on combining thermal imaging and RGB images is still scarce, and their complementary advantages have not been fully utilized. 3. Edge computing and real-time processing requirements: Intelligent monitoring systems in agricultural fields usually need to run on edge devices, which requires the segmentation model to not only have high accuracy but also low latency and low power consumption. Existing deep learning models have high computational resource requirements and are difficult to meet the requirements of real-time inference and efficient processing.

[0004] Therefore, how to combine thermal imaging and RGB images, and improve the accuracy and robustness of broiler chicken instance segmentation through deep learning, especially in high-density, occlusion, and complex backgrounds, and solve the deficiencies of existing methods in environmental adaptability, real-time performance, and segmentation accuracy is an issue that needs to be considered currently. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art, and provides a method for instance segmentation of broiler chickens based on the fusion of thermal imaging and RGB, which solves the deficiencies existing in the prior art.

[0006] The purpose of the present invention is achieved through the following technical solutions: A method for instance segmentation of broiler chickens based on the fusion of thermal imaging and RGB, the segmentation method comprising:

[0007] Step 1: Use an RGB visible light camera and a thermal imager as acquisition devices to obtain image data in two modalities, RGB and thermal imaging. Label the acquired image data according to the labeling strategy to obtain bimodal labels for the same instance, and perform data preprocessing to generate labeled training data for network model learning.

[0008] Step 2: Use the labeled training data after data preprocessing as the training set and input it into the constructed network model for training to obtain an optimized network model.

[0009] Step 3: Re-enter the acquired image data into the optimized network model after training, extract the feature information of RGB and thermal imaging in the image, fuse the feature information of the two modalities through a cross-modal attention mechanism, optimize cross-modal information interaction by combining a deformable attention mechanism and multi-source position encoding, and finally generate the final broiler instance segmentation mask through a decoder.

[0010] The construction of the network model in Step 2 specifically includes the following:

[0011] A1. Use a dual-branch MobileViT as the visual backbone structure to receive RGB images and thermal images respectively. The RGB branch receives three-channel RGB image input, and the thermal branch receives single-channel thermal image input, and extracts multi-scale feature maps through hierarchical convolution and Transformer.

[0012] A2. After extracting the feature maps from each MobileViT branch, connect a CBAM module for feature recalibration to purify and strengthen the RGB and thermal feature maps at each scale to highlight chicken body-related information.

[0013] A3. Construct a feature pyramid structure, upsample the high-level semantic features and fuse them with the low-level detail features to obtain rich and high-resolution feature maps, and introduce an ASPP module on the highest-level semantic features to obtain a feature representation with both local details and global context, better distinguishing adjacent broiler individuals and handling complex backgrounds.

[0014] A4. Achieve deep fusion at the feature level through a cross-modal attention fusion module.

[0015] In the step of A2, for the RGB branch, the CBAM module adaptively assigns weights according to the color and edge features of the broiler target, making the channels related to the chicken body more activated and focusing on the position area where chickens exist in the image.

[0016] For the thermal branch, the CBAM module highlights the features of high-temperature regions according to the temperature intensity and weakens the interference of cold background areas.

[0017] The steps of A4 specifically include the following content:

[0018] A41. On the one hand, using the RGB feature as the query, and the thermal feature as the key and value, calculate the thermally-guided RGB feature. On the other hand, using the thermal feature as the query, and the RGB feature as the key and value, calculate the visually-guided thermal feature;

[0019] A42. Introduce the deformable attention mechanism. For the query position in the RGB image, learn a set of reference points and their offsets , calculate the enhanced feature that fuses the thermal feature semantics at the query position in the RGB image, where is the number of sampled keys for each query, represents the previous offset sampling position on the thermal feature, is the corresponding attention weight. Through the deformable attention mechanism, the RGB feature selectively focuses on a few points adjacent to or semantically related to its own position on the thermal image to achieve cross-modal alignment. Similarly, for the query position in the thermal map, learn a set of reference points and offsets , calculate the attention only at these specific positions, and obtain the enhanced feature that fuses the RGB feature semantics at the query position in the thermal image. Then fuse the two to get the fused feature map ;

[0020] A43. And make full use of the information of both the RGB and thermal imaging modalities through multi-source position encoding to enhance the representation ability of the feature map, provide the combination of spatial coordinates and temperature information, and enable the network model to more accurately match and aggregate the corresponding features based on dual clues of space and temperature;

[0021] A44. Feed the fused feature map into the mask generation Transformer decoder of Mask2Former, and obtain the final segmentation result through the decoder.

[0022] The steps of A43 specifically include the following content:

[0023] For each pixel position in the thermal feature image, the original two-dimensional sine position encoding is denoted as , encode the thermal intensity value at this position into a vector , and through multi-source position encoding, combine with Project after element-wise addition or concatenation to form a new position encoding When calculating attention, add the new position encoding to the corresponding key and value representations. In this way, when calculating the correlation, the attention mechanism not only considers the pure relative position but also perceives the distribution of the heat source intensity;

[0024] Similarly, for the RGB feature map, first obtain a standard two-dimensional sine position encoding PExᵧ to represent the spatial information of the position. At the same time, expand the RGB color value of the position into a vector with the same dimension as the model through linear transformation to form a visual information encoding PExᵧ^RGB. Then, perform element-wise addition or concatenation on PExᵧ and PExᵧ^RGB and project them to generate a mixed position encoding PExᵧ^multi. This encoding is added to the key and value representations in the attention calculation of the Transformer, so that when the attention mechanism calculates the feature correlation, it can not only capture the accurate spatial position information but also obtain rich visual details.

[0025] The steps of A44 specifically include the following content:

[0026] Take the fused feature map as the key and value input of the Transformer decoder, and assign a global multi-source position encoding to the query. In the decoder, each query extracts information from the fused feature map through multi-head attention and continuously updates its own representation vector. Since the fused feature contains RGB and thermal information, the query will automatically focus on the regions with both visual and thermal consistency in the attention, that is, the positions where the broilers are located. After several layers of cross-attention and self-attention iterations, each query will finally converge into a feature vector representing an instance;

[0027] The decoder outputs a mask vector with the same size as the feature map for each query, takes the dot product with the fused feature map to generate a mask prediction, and outputs the corresponding category. The finally obtained query mask removes redundancy through threshold processing and non-maximum suppression, and retains the segmentation results of each broiler.

[0028] The training and optimization of the network model include:

[0029] The input includes the aligned RGB image and thermal image pairs, as well as the instance segmentation ground truth masks corresponding to each image, which are input into the constructed network model. The network parameters are initialized, and iterative optimization is performed using the Adam optimizer or SGD. The initial learning rate is set, and the learning rate is adjusted using the cosine annealing or multi-step decay strategy to promote convergence. During the training process, the total loss function of the network model is constructed by the segmentation loss function, temperature-guided consistency loss function, and edge refinement loss function. The weights of the three loss functions are set respectively according to the training performance during the training process, so that the converged network model has the segmentation performance of high pixel accuracy, thermal consistency, and good edge contours at the same time.

[0030] After the network model is trained and reaches the set accuracy, it is also necessary to perform compression and acceleration optimization on the network model, which specifically includes the following:

[0031] Calculate the importance scores of the convolutional filters or attention heads of each layer of the MobileViT backbone, remove the channels with small contributions, and reduce the number of queries of the Transformer decoder;

[0032] Perform quantization-aware training or post-training quantization on the trained network model to improve the execution efficiency of the network model.

[0033] The RGB visible light camera and thermal imager are used as acquisition devices to acquire RGB and thermal image data of two modalities. According to the annotation strategy, the two-modal annotation of the same instance for the acquired image data includes:

[0034] The RGB visible light camera and thermal imager are used as acquisition devices. The two sensors are fixedly installed on the bracket above the broiler breeding house. The RGB visible light camera faces the ground and covers the chicken activity area from a top-down perspective to ensure that the whole body contours of the broilers can be completely captured in the field of view. The RGB visible light camera is used to obtain visible light color images, and the thermal imager is used to synchronously obtain the thermal radiation intensity map of the same scene. The two acquisition devices ensure the field of view overlap through external calibration, and their relative positions and angles are calibrated so that the obtained RGB image and thermal image correspond spatially;

[0035] Data acquisition is carried out in the form of video streams or continuous image frames. The acquisition frequency is set according to needs, so as to obtain a large number of images containing individual broilers. The RGB visible light camera outputs high-resolution color images, and the thermal imager outputs low-resolution infrared intensity maps. The thermal map is adjusted to match the size of the RGB image through interpolation. The acquisition process covers different time periods and lighting conditions, as well as different distances, postures, and occlusion situations in the chicken flock to obtain diverse data;

[0036] The collected bimodal images need to be annotated to generate the ground truth required for training. For each frame of the image, bounding boxes are marked for each broiler on the RGB image to determine the position range of each chicken. At the same time, the corresponding high-temperature area is determined by combining the thermal image, and the contour of the high-temperature area corresponding to the broiler body is outlined on the thermal map. Each bounding box in the RGB image is corresponded to the high-temperature area in the thermal map through ID matching to form the bimodal annotation of the same instance, and the instance segmentation ground truth mask is generated.

[0037] The data preprocessing includes the following contents:

[0038] The thermal image is projected and transformed into the same coordinate system and resolution as the RGB image. For each pair of RGB image and thermal image, they are cropped according to the field of view overlapping area to ensure that the sizes of the two output images are the same and the pixels correspond one by one. Then, the RGB image is processed in the RGB modality, the thermal image is processed in the thermal imaging modality, and then data augmentation processing is carried out;

[0039] For each annotated broiler instance, within the range of its RGB image bounding box, the normalized temperature value is extracted from the corresponding thermal map area, and a binary mask is extracted according to the set threshold. This mask is cropped with the range of the bounding box to obtain an approximate broiler contour. The binary mask area of each chicken thus generated will be used as the main ground truth for instance segmentation.

[0040] The present invention has the following advantages:

[0041] 1. By deeply fusing the thermal image and the RGB image, the complementary nature of the two modal data can be fully utilized, improving the adaptability of the system in different environments, and thus enhancing the overall performance of instance segmentation in complex scenarios.

[0042] 2. By introducing a lightweight vision Transformer (MobileViT) network and a deformable attention mechanism (Deformable Attention), while effectively reducing the computational complexity, the segmentation accuracy and adaptability of the model are improved, enabling real-time individual broiler segmentation and health monitoring on low-power devices, meeting the high requirements of the agricultural intelligent monitoring system for computational efficiency and real-time performance.

[0043] 3. The broiler instance segmentation method based on the fusion of thermal imaging and RGB images can not only solve the segmentation accuracy problem of the current technology in high-density chicken flocks and complex background environments, but also further broaden its application scope in agricultural intelligent management. By real-time monitoring and analyzing the behavior of the chicken flock, this technology can be used for various applications such as broiler health monitoring, individual tracking, and behavior analysis, promoting the development of agricultural breeding towards a more intelligent and automated direction. Description of the Drawings

[0044] Figure 1 Schematic diagram of the process of the present invention;

[0045] Figure 2 Schematic diagram of the process of the network structure design of the present invention. Specific embodiments

[0046] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the protection scope of the present application claimed, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the protection scope of the present application. The present invention will be further described below with reference to the accompanying drawings.

[0047] As Figure 1 shown, the present invention specifically relates to a broiler instance segmentation method based on thermal imaging and RGB fusion, which solves the limitations of the existing broiler instance segmentation technology in complex agricultural environments, especially the recognition accuracy and real-time processing ability problems under high density, occlusion and complex background conditions. The specific contents are as follows.

[0048] Step 1: Data acquisition;

[0049] First, multi-modal image data is acquired. The present invention uses an RGB visible light camera and a thermal imager as the acquisition devices. The two sensors are fixedly installed on the bracket above the broiler breeding house. The camera faces the ground and covers the chicken flock activity area from a top-down view (for example, vertically downward or slightly inclined by about 45 degrees) to ensure that the whole body contour of the broiler can be completely captured in the field of view. The RGB camera is used to obtain visible light color images, and the thermal imager is used to synchronously obtain the thermal radiation intensity map (temperature thermal map) of the same scene. The two imaging devices ensure that the fields of view overlap through external calibration, and their relative positions and angles are calibrated so that the obtained RGB image and the thermal map correspond spatially (for example, by photographing a calibration board to calculate the mapping matrix and reprojection the thermal map into the RGB image coordinate system to achieve bimodal alignment).

[0050] Data collection is carried out in the form of video streams or continuous image frames, and the collection frequency can be set as needed, for example, from 5 frames per second to 30 frames per second, so as to obtain a large number of images containing individual broiler chickens. Typically, an RGB camera outputs high-resolution color images (e.g., 1280×720 pixels), and a thermal imager outputs infrared intensity maps with lower resolution (e.g., 640×480 pixels). The thermal map can be adjusted to match the size of the RGB image through interpolation. To obtain diverse data, the collection process should cover different time periods and lighting conditions (using infrared fill light during the day and at night), as well as different distances, postures, and partially occluded situations in the chicken flock.

[0051] Annotation strategy: The collected multi-modal images need to be manually annotated to generate the ground truth required for training. For each frame of the image, the annotator marks the bounding box for each broiler chicken on the RGB image to roughly determine the position range of each chicken, and at the same time combines the thermal imaging map to determine its corresponding high-temperature area. On the thermal map, the body of each chicken usually shows a temperature higher than the ambient background, so the temperature hot spot area of the broiler chicken can be marked manually or semi-automatically. Specifically, the annotator outlines the high-temperature area corresponding to the body of the broiler chicken on the thermal map (such as the hot areas of the head and trunk). By matching the IDs, each bounding box in the RGB image is corresponded to the high-temperature area in the thermal map to form multi-modal annotations of the same instance. Based on the above annotations, a more refined instance segmentation ground truth mask can be further generated: for example, the pixels within the bounding box and belonging to the hot area are determined to be chicken body pixels, so as to obtain the approximate binary mask contour of each chicken. This annotation strategy uses the position range provided by the RGB image and the shape clues provided by the thermal map to provide multi-modal supervision signals for subsequent model training.

[0052] Step 2: Data preprocessing;

[0053] Before using the collected data for model training, it is necessary to preprocess and augment the data of the two modalities of RGB and thermal imaging to improve the training effect and the generalization ability of the model.

[0054] First, align and crop the dual-modal images. Using the calibration results in Step 1, projectively transform the thermal image to the same coordinate system and resolution as the RGB image. For each pair of RGB-thermal images, crop them according to the field of view overlap area to ensure that the two output images are of the same size and the pixels correspond one by one. The pixels at the same position now contain the RGB three-channel values and the corresponding thermal intensity values, achieving precise spatial alignment.

[0055] Then, perform respective format transformation and normalization processing on the RGB image and the thermal image, specifically including:

[0056] RGB Modal Processing: Scale or crop the RGB image to the unified size required by the neural network (e.g., 512×512 pixels). Perform color space conversion (such as BGR to RGB) and normalization, subtracting the mean value of the training set from the pixel values of each channel and dividing by the standard deviation, or simply normalizing to the range [0,1]. For RGB images, traditional data augmentation strategies can be applied, such as color jitter (randomly adjusting brightness, contrast, saturation), random blur, etc., but care must be taken to ensure that the correspondence with the heatmap is not destroyed.

[0057] Thermal Imaging Modal Processing: Convert the original temperature values of the thermal image into an intensity map that can be used by the model. For example, adopt a linear normalization strategy: assume that the grayscale range of the original thermal image is (corresponding to a certain temperature range), then normalize each pixel temperature value as follows: to map it to the range between 0 and 1. The pixel values of the thermal image after such processing reflect the relative temperature level and are similar to the pixel value range of the RGB image, facilitating the network to process them simultaneously. In addition, the single-channel thermal intensity map can be copied or extended to three channels (gray three channels) to adapt to the input format of some networks, or a single-channel input layer can be specifically designed in the subsequent model structure. Appropriate filtering and denoising (such as median filtering to smooth thermal noise) and histogram equalization (enhancing temperature contrast) can also be considered in the heatmap preprocessing to highlight the temperature difference area between the broiler and the background.

[0058] In terms of data augmentation, it is particularly important to maintain the multi-modal correspondence. Any spatial transformation-based augmentation operation must be applied to both the RGB and the heatmap simultaneously to ensure that they are still aligned. Common spatial augmentations include the following methods:

[0059] Random flipping and rotation: Flip the image pair (horizontally) with a certain probability, or rotate it within a small angle range, while transforming the corresponding annotation mask. This helps to eliminate directional bias.

[0060] Random translation and scaling: Slightly translate or scale the image (filling the blank areas), and adjust the annotation accordingly. Enhance the robustness in different distance and position situations.

[0061] Mosaic Augmentation: Stitch multiple images together to form a new training sample. For example, randomly select 4 RGB-thermal image pairs from different scenes, scale and crop them, and then tile them into a 2×2 grid large image. In a typical Mosaic operation, first randomly determine the stitching center point position (relative to the size of the output composite image), and then: place the first image in the upper left area of the composite image place the second image in the upper right area place the third image in the lower left area and place the fourth image in the lower right area (where is the size of the output composite image). For adaptation, each original image is scaled according to the size of the divided area. After copying the image pixels to the corresponding quadrants of the composite image, the corresponding heat map quadrants are synthesized in the same way. The coordinates of the annotations (bounding boxes, masks, etc.) of the original images are also offset accordingly and mapped to the coordinate system of the new image. Through Mosaic augmentation, multiple broiler instances from different scenes can be seen simultaneously in a single frame, enriching the background and target combinations and improving the model's adaptability to crowded scenes.

[0062] Other augmentations: On the premise of ensuring bimodal alignment, additional perturbations such as blur and noise are applied to the RGB images, and temperature offsets or noise interferences can be simulated for the thermal images (such as adding or subtracting a small random amount to the heat values or local blur), enhancing the model's robustness to sensor noise.

[0063] Finally, training labels also need to be generated in the preprocessing stage for the model to learn. According to the annotation strategy in step 1, the RGB bounding boxes are combined with the thermal region annotations to generate the ground truth of the segmentation mask for each instance.

[0064] Furthermore, for each annotated broiler instance, within the range of its RGB bounding box, the normalized temperature values are extracted from the corresponding thermal map region, and a binary mask is extracted according to a set threshold (for example, pixels with normalized temperature > 0.5 are regarded as the chicken body region). This mask is cropped with the range of the bounding box to obtain an approximate broiler contour. When necessary, the thermal region mask can be dilated through morphological operations (such as dilating by a certain number of pixels) to cover the entire chicken body contour. The binary mask region of each chicken thus generated will be used as the main ground truth for instance segmentation. These preprocessed and augmented data (aligned RGB images, thermal images, and corresponding instance mask labels) will be used for subsequent model training.

[0065] Step 3, Network structure design;

[0066] As Figure 2 shown, in this step, a segmentation network structure for fusing thermal imaging and RGB information is designed. The overall architecture is improved based on the instance segmentation framework of Mask2Former (an image segmentation model), adopting a lightweight backbone network and multiple attention mechanisms, and constructing a two-stream lightweight backbone + attention + multi-scale feature extraction network, laying a solid foundation for subsequent multi-modal fusion and instance segmentation. While keeping the entire network structure concise, it ensures that the extracted features are highly sensitive and fully expressive for the target (broiler). The main designs are as follows:

[0067] (1) Design a dual-branch MobileViT (a lightweight general-purpose vision transformer for mobile devices) backbone: To extract the features of RGB and thermal images respectively and ensure the model is lightweight and efficient, a two-stream branch MobileViT is used as the visual backbone structure, which specifically includes the following:

[0068] Input the RGB image and the thermal image into two parallel MobileViT backbones respectively. The RGB branch receives the input of a three-channel color image, and the thermal branch receives the input of a single-channel thermal image. To adapt the MobileViT to the processing of thermal images, we modified the input layer of the MobileViT - taking the average of the convolutional kernels of the RGB pre-trained weights on the three channels for the initial weights of the thermal channel, so that it can directly accept the input of a 1-channel thermal image. The network structures of the two branches of MobileViT are similar, and the overall architecture hyperparameters can be shared, but the weights are not shared, so as to learn the optimal features for different modalities respectively. Each MobileViT backbone extracts multi-scale feature representations through hierarchical convolutional and Transformer modules, making the final output feature map have global perception ability while the model is still very compact. By using the MobileViT backbone, the number of model parameters and the computational amount are significantly reduced compared with traditional large backbones (such as ResNet), laying a foundation for real-time operation on edge devices.

[0069] (2) Multi-path CBAM (an attention module that combines channel and spatial attention) attention fusion: Introduce the multi-path CBAM attention mechanism in the dual-branch backbone to strengthen the features of each modality. After extracting the intermediate feature maps from each MobileViT branch, the CBAM module is connected for feature recalibration (this constitutes the "multi-path" - that is, the CBAM is inserted at multiple stages of each branch of RGB and thermal).

[0070] For the RGB branch, CBAM adaptively assigns weights according to the features such as the color and edges of the broiler target, making the channels related to the chicken body more activated and focusing on the area where the chicken exists in the image; for the thermal branch, CBAM highlights the features of the high-temperature area according to the temperature intensity and weakens the interference of the cold background area. Through the multi-path (multi-level, dual-branch) CBAM processing, the feature maps of RGB and thermal are purified and strengthened at each scale. This attention mechanism ensures that before entering the modality fusion, the features of each modality have highlighted the information related to the chicken body (such as: the RGB features highlight the contour and texture of the chicken, and the thermal features highlight the temperature hotspots of the chicken), providing high signal-to-noise ratio feature inputs for subsequent fusion.

[0071] (3)Multi-scale Feature Extraction and ASPP (a deep learning technique for extracting multi-scale features) Module: Considering that the size and distance of broiler chickens in the image may vary, in order to improve the recognition of targets at different scales, the present invention incorporates a multi-scale feature extraction strategy in the backbone output stage.

[0072] Specifically, the MobileViT backbone outputs several feature maps (for example, at 1 / 4, 1 / 8, and 1 / 16 of the original image resolution respectively), constructs a simple FPN (Feature Pyramid Network) structure, upsamples the high-level semantic features and fuses them with the low-level detailed features, so as to obtain rich and high-resolution feature maps for segmentation. In addition, an ASPP (Atrous Spatial Pyramid Pooling) module is introduced on the highest-level semantic features to obtain a feature representation that combines local details and global context, better distinguishing adjacent broiler chicken individuals and handling complex backgrounds. ASPP is applied to the top-level features of both the RGB and thermal branches respectively to enhance the robustness of each modality to scale changes. The feature maps output by ASPP will serve as the main input features for subsequent cross-modal feature fusion and Transformer decoding.

[0073] Step 4: Introduce two modal feature fusion mechanisms;

[0074] After extracting the features of RGB and thermal imaging respectively, it is necessary to effectively fuse the information of these two modalities to utilize their complementary advantages. In this step, a cross-modal attention fusion module is designed to achieve deep fusion at the feature level. The fusion mechanism is based on the attention structure of Transformer (a deep learning model based on the attention mechanism), combines Deformable Attention and multi-source position encoding to optimize cross-modal information interaction, enabling the model to accurately distinguish each broiler chicken from the background and other individuals in various complex scenarios.

[0075] (1)Overview of Feature Alignment and Fusion: Since after the processing in Step 3, the high-level feature maps output by the RGB branch and the thermal branch are consistent in spatial size (both restored to be close to the original image resolution, or aligned through upsampling), and the pixel coordinates correspond. Therefore, assuming the RGB feature map and the thermal feature map , where is the feature map resolution, is the number of channels. The goal is to fuse these two feature maps into a comprehensive feature map , containing more discriminative information for broiler chicken instance recognition. Direct fusion methods such as concatenation or element-wise addition are difficult to fully utilize the complementarity of the two modalities. For this reason, a cross-modal attention fusion module is designed to adaptively combine the RGB and thermal features through the learned attention weights.

[0076] (2) Cross-modal attention mechanism: Cross-modal attention adopts the query-key-value operation idea of Transformer, allowing the features of one modality to be used as the query to retrieve the relevant information in the features of another modality as the key / value, thereby realizing information interaction.

[0077] Specifically, two-way cross-modal attention is adopted: on the one hand, using the RGB features as the query, and the thermal features as the key and value to calculate the heat-guided RGB features; on the other hand, using the thermal features as the query, and the RGB features as the key and value to calculate the vision-guided thermal features. Then the two are combined to form the final fused features. The specific content includes the following:

[0078] Heat-guided RGB attention: Let , first obtain the query matrix, key matrix, and value matrix through linear mapping: , where is the learnable projection matrix.

[0079] When focusing on the RGB query at the -th spatial position, its attention weight is calculated by the dot product of the query vector at this position and the key vectors at all positions of the thermal features:

[0080] (1),

[0081] where, represents the calculated attention weight, is the query vector extracted from the RGB features, representing the RGB features at the i-th position (e.g., the feature representation of a certain pixel point in the RGB image), is the key vector extracted from the thermal image features, representing the thermal map features at the thermal pixel position j, j represents the current thermal image position to be calculated, represents traversing all positions of the thermal image. Then calculate the enhanced features of the RGB image at the query position :

[0082] (2),

[0083] where, is the value at the j -th position in the thermal map, that is, the value of the thermal map features, which is the weighted aggregation of the part with the highest correlation in the thermal features related to the position . The obtained It can be understood as "RGB features combined with thermal information".

[0084] Visual gravitational thermal attention: Similarly, by swapping roles, using thermal features as queries and RGB features as key values, enhanced features after fusing RGB semantics at each heatmap position can be obtained :

[0085] (3),

[0086] where is the attention weight calculated according to the thermal query and the RGB key , is the feature at the j th position in the RGB feature map.

[0087] After obtaining and , the feature maps of the two are fused . The fusion method can adopt channel concatenation + convolution or weighted summation, etc. For example, the and vectors at the corresponding positions can be concatenated and then mapped back to dimensions through convolution, or simply added ( is a learnable weight, initially set to 0.5). In this way, contains clues from both RGB and thermal at each position: visual texture and contour as well as thermal intensity information.

[0088] (3) Introduce deformable attention: To improve the efficiency and robustness of attention fusion, a deformable attention mechanism is introduced into the above cross-modal attention. Standard global attention needs to traverse all keys for each query, with a large computational amount and strict assumptions of spatial alignment. Deformable attention allows each query to only focus on a limited number of sampled points and can learn the offsets of these points, thus flexibly corresponding to the position deviations of cross-modal features. Specifically, for a query position of the RGB feature, instead of summing over all positions of the heatmap, a set of reference points (such as 5) and their offsets are learned, and attention is only calculated at these sampled positions. Then formula (2) is modified to:

[0089] (4),

[0090] where is the number of sampled keys for each query, represents the offset sampled position on the thermal feature, corresponds to the attention weights (obtained by calculating softmax normalization at these positions according to the aforementioned method). Through this mechanism, the RGB features can selectively focus on a few points adjacent to or semantically related to their own positions on the heatmap, achieving cross-modal alignment. For example, when the contour of a certain chicken in the RGB image is unclear due to visible light, its query can flexibly "search" for local high-temperature points on the heatmap, even if they are not at the exact same pixel position due to parallax. With deformable attention, cross-modal fusion is both efficient (reducing computational complexity) and robust (allowing a certain position offset).

[0091] Similarly, for each query position on the heatmap j , learn a set of reference points and offsets , and calculate the attention only at these specific positions. Then formula (3) is modified as:

[0092] (5),

[0093] In this way, both the heatmap features and the RGB features can be flexibly aligned through deformable attention, improving the robustness of cross-modal fusion.

[0094] The fused feature map is obtained after the cross-modal attention mechanism and the deformable attention mechanism. Specifically, the RGB and heatmap features are first fused through the cross-modal attention mechanism, and then further aligned and adjusted through deformable attention. This feature map will be used as the input of the subsequent mask generation decoder for broiler instance segmentation.

[0095] (4) Multi-source position encoding: mainly used to enhance the representation ability of the feature map after feature map fusion, providing the combination of spatial coordinates and temperature information; Transformer attention usually needs to add position encoding to inject spatial position information. In this scheme, we design multi-source position encoding to make full use of the information of both RGB and thermal modalities. Specifically, a hybrid position encoding that simultaneously considers spatial coordinates and temperature values is introduced into the fusion module. For example, for each pixel position of the thermal feature image , the original two-dimensional sine position encoding is denoted as (including coordinate information); at the same time, we can encode the thermal intensity value at this position as a vector (for example, by linearly transforming the scalar temperature value to a vector with the same dimension as the model, or using high-frequency sine encoding of the heat value). The multi-source position encoding fuses the two, such as adding or concatenating and element-wise and then projecting to form a new position encoding 。When calculating attention, this encoding is added to the corresponding key and value representations. In this way, when calculating the correlation, the attention mechanism not only considers the pure relative position, but also perceives the distribution of the heat source intensity.

[0096] For the RGB feature map, at each pixel position (x, y), a standard two-dimensional sine position encoding PExᵧ is first obtained to represent the spatial information of that position; at the same time, the RGB color value (or other visual features) at that position is extended to a vector with the same dimension as the model through a linear transformation to form a visual information encoding PExᵧ^RGB; then, after element-wise addition or concatenation of PExᵧ and PExᵧ^RGB and projection, a mixed position encoding PExᵧ^multi is generated, which is added to the key and value representations in the attention calculation of the Transformer, so that when the attention mechanism calculates the feature correlation, it can not only capture accurate spatial position information, but also obtain rich visual details.

[0097] For example, even if two pixels are far apart in the image, but if they both have high-temperature features, the thermal encoding part will make them more relevant in the attention. This extension of the position encoding incorporating temperature information is similar to the idea of embedding depth into the position encoding in the DepthFormer method for multi-modal Transformers, enabling the Transformer attention to naturally fuse bimodal information without adding extra parameters. Therefore, the multi-source position encoding works in coordination with the cross-modal attention, enabling the model to more accurately match and aggregate the corresponding features based on both spatial and temperature cues.

[0098] (5) Mask Transformer decoding: The fused features are then fed into the mask generation Transformer decoder of Mask2Former. The decoder contains a set of learnable query embeddings (each query is designed to correspond to a potential broiler instance) and multiple layers of mask attention blocks.

[0099] The fused feature map is used as the key and value input to the Transformer decoder, and a global multi-source position encoding is assigned to the queries. In the decoder, each query extracts information from the fused features through multi-head attention and continuously updates its own representation vector. Since the fused features contain RGB and thermal information, the queries will automatically focus on the regions with both visual and thermal consistency in the attention - which is exactly where the broilers are located. After several layers of iterative cross-attention and self-attention, each query will finally converge into a feature vector representing an instance. Subsequently, these queries are mapped to mask segmentation results and class outputs through a linear layer.

[0100] Specifically, the decoder outputs a mask vector of the same size as the feature map for each query, takes the dot product of it with the fused feature map to generate mask predictions (the segmentation probability map corresponding to that query), and outputs the corresponding category (in this application, the category may only be "broiler chicken"). Since we adopt the masked attention strategy, the attention mechanism is also used in the decoding process to directly generate masks on the feature map, eliminating the need for prior detection and then segmentation. Finally, the obtained query masks are processed through thresholding and non-maximum suppression (NMS) to remove redundancy, retaining the segmentation results of each broiler chicken.

[0101] Step 5: Model training and optimization;

[0102] After completing the model architecture design, the model needs to be effectively trained and optimized. To this end, a new multi-modal fusion loss function is constructed, combining the objectives of segmentation accuracy, temperature consistency, and edge refinement, and performing end-to-end training on the model. This step will detail the training strategy and loss function design.

[0103] (1) Use the pre-processed annotated training data in Step 2 as the training set and input it into the model. The training is carried out in an end-to-end manner, with the input including aligned RGB images and thermal image pairs, as well as the ground truth masks (and class labels) corresponding to each image. Initialize the network parameters: The MobileViT backbone can use the weights pre-trained on ImageNet (RGB branch). Since there is no pre-training for the thermal branch, it can be randomly initialized or partially borrow the RGB pre-training parameters; the Transformer decoder and fusion module are randomly initialized. Use the Adam optimizer or SGD for iterative optimization. The initial learning rate is set according to the model size (e.g., 1e-4), and the learning rate is adjusted using the cosine annealing or multi-step descent strategy to promote convergence. The entire training is carried out on the GPU for several epochs until the segmentation performance on the validation set converges and improves.

[0104] (2) Multi-task loss function: During the training process, each batch of input passes through the model to obtain the predicted broiler chicken instance masks and corresponding categories. The total loss function designed in the present invention consists of the following three parts.

[0105] (a) Segmentation main loss : Measures the difference between the predicted instance mask and the ground truth mask, which is the most basic supervision signal. Since this task belongs to instance segmentation, it can be regarded as a pixel-level binary classification (foreground: chicken, background) problem. A combination of binary cross-entropy loss and Dice loss is selected to measure the segmentation effect, thus taking into account both pixel accuracy and overall contour matching. For a single predicted mask (representing the probability that pixel is the foreground) and the corresponding binary ground truth mask , the cross - entropy loss function is defined as:

[0106] (6),

[0107] is the total number of image pixels, and cross - entropy emphasizes the accurate classification of each pixel.

[0108] The Dice loss function is defined as:

[0109] (7),

[0110] This loss is based on the Dice coefficient, which measures the overlap between the predicted mask and the ground - truth mask. is a smoothing term to prevent the denominator from being zero. When the prediction and the ground - truth highly overlap, the Dice loss approaches 0.

[0111] Furthermore, the main segmentation loss function is expressed as a weighted sum of the two:

[0112] (8),

[0113] where is the weight coefficient (which can be taken as about 0.5 for balance). Designed in this way, it not only ensures pixel - level accuracy but also pays attention to the overall contour quality, which is commonly used and effective in instance segmentation tasks.

[0114] (b) Temperature - guided consistency loss : This loss utilizes thermal imaging information to ensure the consistency between the model prediction and the heat - map signal. Intuitively, it is expected that the predicted chicken body mask covers the high - temperature regions on the heat - map, and there should be no false alarms in the background (low - temperature regions). The temperature distribution provided by thermal imaging is transformed into a soft target to supervise the segmentation output. The specific approach is as follows: Threshold the normalized value of the heat - map to obtain a binary heat - zone ground - truth mask as:

[0115] (9),

[0116] where is the temperature threshold (which can be set according to environmental temperature experience, such as 0.5 or adaptively adjusted according to the background average temperature). represents that pixel belongs to the high - temperature region (most likely part of the chicken), represents the low - temperature background. Then, the temperature consistency loss is defined as the pixel binary - classification error between the predicted probability map and to obtain the temperature - guided consistency loss as:

[0117] (10),

[0118] This actually treats as an additional supervision mask and calculates the cross-entropy loss with the prediction. If is directly used for soft labels (non-binary), the KL divergence can also be used to measure the closeness of the distribution. By , the model is penalized for misclassifying as background (false negatives) in high-temperature regions and for misclassifying as foreground in cold background regions. In other words, the temperature consistency loss guides the model to learn the rule of "chickens in hot areas, no chickens in cold areas", making the prediction results match the infrared thermal pattern. This ensures that the model does not miss real broiler chickens even under lighting changes or color camouflage (when RGB is unreliable).

[0119] (c) Edge refinement loss : To further improve the contour quality of the segmentation mask, an edge refinement loss is introduced to specifically handle the accuracy of the object boundary. Usually, the segmentation model may have cases where the mask edge is not smooth enough or misaligned with the ground truth boundary. This loss is defined by comparing the edges of the ground truth mask and the predicted probability map . . can be represented as a 0-1 binary map, where 1 indicates that the pixel is on the mask boundary. Then the edge refinement loss is defined as the average of the absolute values of the difference between the two, resulting in the edge refinement loss as:

[0120] (11),

[0121] When the boundary of the predicted mask exactly matches the ground truth, the difference between the two is zero and the loss is the lowest; if the boundary has deviations, misalignments, or jaggedness, the loss increases. By minimizing , the model is encouraged to produce predictions with the same boundary shape as the ground truth. For example, if the ground truth mask edge is a smooth curve and the prediction has jaggedness, the loss will drive the model to adjust to make its edge smoother and more continuous; if the predicted edge is overall shifted outward or inward, the loss will also drive the model to correct the shift. It should be noted that focuses on shape consistency and is insensitive to minor errors within the pixel level, which complements (pixel accuracy) to jointly improve the final mask quality.

[0122] (3) Total Loss and Optimization: Combining the above three parts, the total training loss function is constructed as follows:

[0123] (12),

[0124] where are adjustable weight coefficients used to balance the contributions of the three losses to the overall objective. Initially, all three can be set to 1 or tuned according to the performance on the validation set. For example, if it is found that the segmentation accuracy is sufficient but the edges are not fine, the weight of can be appropriately increased. During the training process, with as the objective, the model parameters are updated through backpropagation to gradually reduce each sub-loss. It is worth mentioning that due to the addition of and , the model may require more iterations to simultaneously optimize the multi-task objective at the initial stage of training. A staged training strategy can be adopted: first, train with only for several epochs to obtain the basic segmentation ability, and then gradually introduce and to enhance the details. Of course, the full loss function can also be used simultaneously from the beginning of training. Through reasonable hyperparameters and training scheduling, the model will have a segmentation performance with high pixel accuracy, thermal consistency, and good edge contours after convergence.

[0125] Step 6: Model Compression and Optimization;

[0126] After the model training is completed and satisfactory accuracy is achieved, in order to deploy it to resource-constrained edge devices, the model needs to be compressed and optimized for acceleration. This step includes techniques such as model pruning and quantization to streamline the complex deep model into a more efficient form.

[0127] (1) Model Pruning: Channel pruning is performed on the convolutional layers and attention layers of the MobileViT backbone: Calculate the importance scores of each convolutional filter or attention head (which can be based on the norm of the weights or sensitivity analysis of the loss), and then remove those channels with less contribution. Especially in the dual-branch structure, there may be feature redundancy between the RGB branch and the thermal branch, and some redundant channels can be removed to reduce the model size. In addition, the number of queries in the Transformer decoder is also appropriately pruned to reduce the number of queries and thus reduce the computation. The pruning adopts a progressive pruning strategy: First, fine-tune the model with a small pruning rate, and gradually increase the pruning intensity through multiple iterations. After each pruning, the model is fine-tuned and trained (fine-tune) to recover the accuracy.

[0128] (2)Model Quantization: The 8-bit quantization (INT8 quantization) method is adopted to compress the model from 32-bit floating-point representation. The specific process includes: performing quantization-aware training (QAT) or post-training quantization (PTQ) on the trained model. Quantization-aware training continues to train for several epochs after training, but simulates operations such as convolution and matrix multiplication in the model as 8-bit quantization (simulating quantization errors through scaling factors and zero points), enabling the model parameters to gradually adapt to quantization errors to avoid accuracy loss. During the quantization process, special attention should be paid to calibrating the attention operations and normalization layers in MobileViT to ensure an appropriate numerical dynamic range. For the dual-branch, it is necessary to statistically analyze the activation value distributions of each layer in the RGB and thermal branches to select appropriate quantization scales. After completing QAT, the weights are solidified into the INT8 format. Through quantization, the model volume will be reduced by approximately 4 times, and during inference, the calculation can utilize the 8-bit integer instructions of the CPU / special accelerator, significantly improving the execution efficiency. At the same time, it is verified that the segmentation accuracy of the quantized model on the validation set only drops slightly and still meets the requirements.

[0129] (3)Other Optimizations: Use the original model as the teacher network to train a smaller student network, enabling the student to learn the teacher's behavior in bimodal fusion, thereby further reducing the model complexity. In addition, using grouped convolution to optimize the convolution operations in MobileViT and optimizing the operators in the fusion layer (such as decomposing multi-head attention into multiple small matrix multiplications) can also improve the efficiency.

[0130] Step 7, Edge Deployment and Accelerated Inference;

[0131] Deploy the optimized model to the edge computing device, which is composed of a dual-modal camera sensor at the front end and an embedded AI module at the back end, and can operate stably in the actual breeding environment, providing a powerful tool for accurate identification and segmentation of each broiler chicken for intelligent breeding management.

[0132] (1)Deployment Platform Selection: According to the on-site situation of the farm, select an edge device with sufficient performance to run the model in real time, such as a small embedded system equipped with NVIDIA Jetson series modules (TX2, Xavier, Orin, etc.), or other embedded AI devices supporting GPU acceleration. The deployment platform installs a camera interface for accessing the signals of the RGB camera and thermal imager in Step 1, and has a model inference framework environment (such as TensorRT acceleration library, CUDA support, etc.).

[0133] (2) Model Conversion and Acceleration: Convert the pruned and quantized model into a deployment format, such as ONNX format, and then use TensorRT to optimize and compile the model. TensorRT is deeply optimized for the NVIDIA GPU platform. It fuses operators in the computation graph, optimizes memory, and uses FP16 / INT8 operations to improve the inference speed. We use TensorRT to optimize the convolutional and attention layers of the model into efficient kernels and enable INT8 precision (combining with the calibration table quantized in step 6). In particular, TensorRT unfolds and optimizes the convolutional and Transformer sub-modules of MobileViT to complete the computation with the least memory access on the GPU. In addition, optimize the input and output: Configure the inference with a fixed batch size (batch = 1) to reduce dynamic overhead; Pre-allocate the video memory buffer to store the input RGB / heatmap and the output mask. The engine file compiled by TensorRT can be directly loaded and run on the device to achieve maximum throughput.

[0134] (3) Real-time Inference: At the edge device side, the running program continuously reads synchronized frames from the dual cameras, performs preprocessing on each pair of RGB and heatmap (according to the process in step 2: alignment, normalization, etc.), and then sends the processed tensors into the TensorRT engine for forward inference. The model outputs the segmentation masks (and class labels) of all broiler instances in the frame. Perform minor post-processing on the output masks: Remove small areas (remove noise fragments of a few pixels), smooth the edges (correct the burrs using morphological smoothing or boundary tracking algorithms), and assign IDs to each instance mask for continuous frame tracking, etc. Finally, draw the masks on the original image or output them in coordinate form to achieve real-time detection and contour marking of broilers.

[0135] (4) Performance Metrics: The present invention not only focuses on the inference speed and accuracy but also needs to evaluate the actual performance of the system in the actual breeding environment. The following are the targeted performance evaluation metrics.

[0136] Average Inference Latency: Calculate the processing time for each pair of input RGB and heatmap images and report the average processing time per frame.

[0137] (13),

[0138] L For the number of frames evaluated, taking Jetson Xavier as an example, through actual testing, the single-frame inference latency is about 40 milliseconds, which is equivalent to a processing speed of more than 25 frames per second. For a resolution of The input dual-modal images can meet the real-time requirements of farm video monitoring (the general video frame rate is 25fps). Even on the relatively low-end TX2 platform, after INT8 optimization, it can reach about 15fps, basically meeting the real-time monitoring needs. The model memory occupancy and power consumption are also controlled within an acceptable range: the size of the quantized model is only dozens of MB, the GPU occupancy rate during inference is moderate, and it can run stably for a long time. In actual deployment, the system can work continuously for 7×24 hours, perform instance segmentation of broiler chickens frame by frame on the ingested image stream, and obtain the position and contour information of each chicken in real time. This information can be further used for applications such as chicken counting, behavior analysis, and health status monitoring (combined with temperature anomaly detection).

[0139] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and improvements, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A broiler instance segmentation method based on the fusion of thermal imaging and RGB, characterized in that: The segmentation method includes the following: Step 1: Use an RGB visible light camera and a thermal imager as acquisition devices to acquire image data in two modalities, RGB and thermal imaging. Label the acquired image data according to the labeling strategy to obtain bimodal labels of the same instance, and perform data preprocessing to generate labeled training data for network model learning. Step 2: Input the labeled training data after data preprocessing into the constructed network model for training to obtain an optimized network model. Step 3: Re-input the acquired image data into the optimized network model, extract the feature information of the RGB and thermal imaging modalities in the image, fuse the feature information of the two modalities through a cross-modal attention mechanism, optimize cross-modal information interaction by combining a deformable attention mechanism and multi-source position encoding, and finally generate the final broiler instance segmentation mask through a decoder. The construction of the network model in Step 2 specifically includes the following: A1: Use a dual-branch MobileViT as the visual backbone structure to receive RGB images and thermal images respectively. The RGB branch receives the input of a three-channel RGB image, and the thermal branch receives the input of a single-channel thermal image, and extracts multi-scale feature maps through hierarchical convolution and Transformer. A2: After extracting the feature maps from each MobileViT branch, connect a CBAM module for feature recalibration to purify and strengthen the RGB and thermal feature maps at each scale to highlight chicken body-related information. A3: Construct a feature pyramid structure, upsample the high-level semantic features and fuse them with the low-level detail features to obtain rich and high-resolution feature maps, and introduce an ASPP module on the highest-level semantic features to obtain a feature representation with both local details and global context, better distinguishing adjacent broiler individuals and handling complex backgrounds. A4: Achieve deep fusion at the feature level through a cross-modal attention fusion module. The steps of A4 specifically include the following: A41. On the one hand, using the RGB features as the query, and the thermal features as the key and value, calculate the heat-guided RGB features. On the other hand, using the thermal features as the query, and the RGB features as the key and value, calculate the vision-guided thermal features; A42. Introduce a deformable attention mechanism to query positions in RGB images , learn a set of reference points and their offsets , calculate the enhanced feature that fuses the thermal feature semantics at the query position in the RGB image as , where is the number of sampled keys for each query, represents the previous offset sampling position on the thermal feature, is the corresponding attention weight; similarly, for the query position in the heat map , learn a set of reference points and offsets , calculate attention only at these positions to obtain the enhanced feature that fuses the RGB feature semantics at the query position in the heat image as , and fuse the two to obtain the fused feature map ; A43: And use multi-source position encoding to make full use of the information of the RGB and thermal imaging modalities to enhance the representation ability of the feature maps, provide the combination of spatial coordinates and temperature information, and enable the network model to more accurately match and aggregate corresponding features based on dual clues of space and temperature. A44: Send the fused feature maps into the mask generation Transformer decoder of Mask2Former to obtain the final segmentation result through the decoder.

2. The method for instance segmentation of broilers based on the fusion of thermal imaging and RGB according to claim 1, characterized in that: In the steps of A2, for the RGB branch, the CBAM module adaptively assigns weights according to the color and edge features of the broiler target, making the channels related to the chicken body more activated and focusing on the position area where chickens exist in the image. For the thermal branch, the CBAM module highlights the features of the high-temperature area according to the temperature intensity and weakens the interference of the cold background area. The high-temperature area refers to the red area in the thermal image.

3. A method for instance segmentation of broilers based on the fusion of thermal imaging and RGB according to claim 1, characterized in that: The steps of A43 specifically include the following: For each thermal feature image pixel position , the original two-dimensional sine position encoding is denoted as , and the thermal intensity value at this position is encoded into a vector . Through multi-source position encoding, is added element-wise or concatenated with and then projected to form a new position encoding . When calculating attention, the new position encoding is added to the corresponding key and value representations. In this way, when calculating the correlation, the attention mechanism not only considers the pure relative position but also perceives the distribution of the heat source intensity; Similarly, for the RGB feature map, first, a standard two-dimensional sine positional encoding PExᵧ is obtained to represent the spatial information of the position. At the same time, the RGB color value at this position is linearly transformed into a vector with the same dimension as the model to form a visual information encoding PExᵧ^RGB. Then, PExᵧ and PExᵧ^RGB are element-wise added or concatenated and projected to generate a mixed positional encoding PExᵧ^multi, which is added to the key and value representations in the attention calculation of the Transformer. Thus, when the attention mechanism calculates the feature correlation, it can capture both precise spatial position information and rich visual details.

4. A method for instance segmentation of broilers based on the fusion of thermal imaging and RGB, as claimed in claim 1, wherein: The steps of A44 specifically include the following: The fused feature map is input as the key and value of the Transformer decoder, and a global multi-source positional encoding is assigned to the query. In the decoder, each query extracts information from the fused feature map through multi-head attention and continuously updates its own representation vector. Since the fused feature contains RGB and thermal information, the query will automatically focus on the regions with both visual and thermal consistency in the attention, that is, the positions where the broiler chickens are located. After several layers of iterative cross-attention and self-attention, each query will finally converge into a feature vector representing an instance. The decoder outputs a mask vector with the same size as the feature map for each query, takes the dot product of it and the fused feature map to generate a mask prediction, and outputs the corresponding category. The finally obtained query mask is processed through thresholding and non-maximum suppression to remove redundancy and retain the segmentation results of each broiler chicken.

5. A broiler instance segmentation method based on thermal imaging and RGB fusion according to claim 1, characterized in that: The training and optimization of the network model include: The input includes aligned RGB images and thermal image pairs, as well as the ground truth mask of instance segmentation corresponding to each image, into the constructed network model, and the network parameters are initialized. The Adam optimizer or SGD is used for iterative optimization, the initial learning rate is set, and the cosine annealing or multi-step descent strategy is used to adjust the learning rate to promote convergence. During the training process, the total loss function of the network model is constructed by the segmentation loss function, the temperature-guided consistency loss function, and the edge refinement loss function. The weights of the three loss functions are set respectively according to the training performance during the training process, so that the converged network model has the segmentation performance of high pixel accuracy, thermal consistency, and good edge contours at the same time.

6. The method for instance segmentation of broilers based on thermal imaging and RGB fusion according to claim 1, characterized in that: After the network model is trained and reaches the set accuracy, it is also necessary to perform compression and acceleration optimization on the network model, which specifically includes the following: Calculate the importance scores of each convolutional filter or attention head in each layer of the MobileViT backbone, remove the channels with small contributions, and reduce the number of queries in the Transformer decoder; Perform quantization-aware training or post-training quantization on the trained network model to improve the execution efficiency of the network model.

7. A method for instance segmentation of broilers based on the fusion of thermal imaging and RGB according to claim 1, characterized in that: The RGB visible light camera and the thermal imager are used as acquisition devices to acquire RGB and thermal imaging two-modal image data. The two-modal annotation of the same instance obtained by annotating the acquired image data according to the annotation strategy includes: An RGB visible light camera and a thermal imager are used as acquisition devices. The two sensors are fixedly installed on a bracket above the broiler chicken house. The RGB visible light camera faces the ground and covers the chicken activity area from a top-down perspective to ensure that the whole body contour of the broiler chickens can be completely captured in the field of view. The RGB visible light camera is used to obtain visible light color images, and the thermal imager is used to synchronously obtain the thermal radiation intensity map of the same scene. The two acquisition devices ensure field of view overlap through external calibration, and their relative positions and angles are calibrated so that the obtained RGB images and thermal images correspond spatially; Data acquisition is carried out in the form of video streams or continuous image frames. The acquisition frequency is set according to needs to obtain a large number of images containing individual broiler chickens. The RGB visible light camera outputs high-resolution color images, and the thermal imager outputs infrared intensity maps with lower resolution. The thermal map is adjusted to match the size of the RGB image through interpolation. The acquisition process covers different time periods and lighting conditions, as well as different distances, postures, and occlusion situations in the chicken flock to obtain diverse data; The acquired bimodal images need to be annotated to generate the ground truth required for training. For each frame of image, a bounding box is marked for each broiler chicken on the RGB image to determine the position range of each chicken. At the same time, the corresponding high-temperature area is determined in combination with the thermal image, and the contour of the high-temperature area corresponding to the broiler chicken body is outlined on the thermal map. Each bounding box in the RGB image is corresponded to the high-temperature area in the thermal map through ID matching to form a bimodal annotation of the same instance and generate an instance segmentation ground truth mask.

8. A method for instance segmentation of broilers based on the fusion of thermal imaging and RGB as claimed in claim 1, wherein: The data preprocessing includes the following: The thermal image is projected and transformed into the same coordinate system and resolution as the RGB image. For each pair of RGB image and thermal image, they are cropped according to the field of view overlap area to ensure that the two output images are of the same size and the pixels correspond one by one. Then, the RGB image is processed in the RGB modality, the thermal image is processed in the thermal imaging modality, and then data augmentation processing is carried out; For each annotated broiler chicken instance, within the range of its RGB image bounding box, the normalized temperature value is extracted from the corresponding thermal map area. A binary mask is extracted according to the set threshold, and this mask is cropped with the range of the bounding box to obtain the broiler chicken contour. The binary mask area of each chicken thus generated will be used as the main ground truth for instance segmentation.

Citation Information

Patent Citations

  • RGB-T multi-mode image instance segmentation method based on deep learning

    CN117934843A

  • A salient object detection method based on lightweight neural network

    CN119741580A