Non-contact evaluation method and system for health state of calf in lactation period
By combining multispectral imaging, depth sensing, and directional audio acquisition technologies with a multi-channel spatiotemporal feature fusion neural network model, the problems of subjectivity and insufficient monitoring in existing calf health management technologies have been solved. This enables accurate health assessment and early abnormality identification around the clock and without stress, forming a self-optimizing intelligent management system.
Patent Information
- Application Number
- CN202511633691.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for the health management of lactating calves are subject to strong subjectivity, cannot monitor around the clock, are prone to missing early abnormalities, and pose risks of stress and cross-infection due to contact operations. Furthermore, the monitoring dimensions are limited and the information is one-sided, making it difficult to accurately identify sub-health states or the early stages of diseases.
Employing multispectral imaging, depth sensing, and high-sensitivity directional audio acquisition technologies, combined with a multi-channel spatiotemporal feature fusion neural network model, non-contact multimodal data acquisition and intelligent comprehensive analysis are achieved. Health status assessment is performed through the multi-channel spatiotemporal feature fusion neural network model, integrating visual, thermodynamic, geometric motion, and acoustic features. Information is dynamically integrated using an attention mechanism to generate a health status assessment vector, which is then combined with a multivariate threshold model for determination.
It achieves all-weather, undisturbed, and accurate identification of early abnormalities in calves, reduces stress response, improves the accuracy and timeliness of health assessment, and forms an intelligent management system with self-optimization capabilities.
Smart Images

Figure CN121506480A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of animal husbandry technology, and in particular to a non-contact assessment method and system for the health status of suckling calves. Background Technology
[0002] In existing animal husbandry, especially in dairy farming, the health management of suckling calves is crucial, as it directly affects calf survival rates, future production performance, and the economic benefits of the farm. However, traditional health assessment methods mainly rely on daily pen patrols and manual physical examinations by farm workers. This approach has significant limitations. First, manual observation is highly subjective, and differences in experience can lead to inconsistent judgment standards. Furthermore, it cannot provide continuous, 24 / 7 monitoring, easily missing early abnormal symptoms in calves at night or when unattended. Second, direct contact procedures such as temperature measurement, oral examination, or auscultation of calves can cause significant harm. Significant stress responses can affect growth and development and may also increase the risk of cross-infection. In recent years, although some automated monitoring technologies based on computer vision or single sensors have emerged, such as monitoring activity levels through cameras or using ear tag thermometers, these methods often suffer from problems such as single monitoring dimensions and incomplete information. A simple decrease in activity levels may be caused by a variety of reasons, making it difficult to distinguish between normal rest and the precursor to disease. A single body temperature data cannot fully reflect the comprehensive state of multiple systems such as respiratory, digestive, and behavioral systems, resulting in insufficient accuracy and timeliness of early warnings. This makes it impossible to achieve early, accurate, and stress-free identification of sub-health conditions or the budding stage of diseases in calves.
[0003] Therefore, the livestock industry urgently needs a health status assessment solution that can integrate multi-dimensional information, achieve 24 / 7 uninterrupted operation, and perform intelligent comprehensive analysis. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this invention provides a non-contact assessment method and system for the health status of lactating calves.
[0005] The technical solution of this invention is: a non-contact assessment method for the health status of lactating calves, comprising: Step 1: Using a multispectral imaging unit, a depth sensing unit, and a high-sensitivity directional audio acquisition unit deployed above the calf's living area, non-contact multimodal raw data streams of the target calf are collected synchronously and continuously within a preset monitoring period. The multispectral imaging unit covers at least the visible light band and the long-wave infrared thermal imaging band to obtain the calf's body surface color texture morphology information and body surface temperature distribution information. The depth sensing unit is used to obtain the calf's three-dimensional point cloud data to reconstruct its accurate body contour and spatial movement trajectory. The high-sensitivity directional audio acquisition unit is used to separate and collect the calf's unique calls, breathing sounds, and sucking sounds after noise reduction. Step 2: Transmit the multimodal raw data stream to the edge computing node for timestamp alignment and data preprocessing. The preprocessing includes illumination normalization and background subtraction of image data to highlight the calf target, environmental temperature compensation calibration of thermal imaging data, point cloud filtering and registration of depth data to construct a coherent three-dimensional motion model, and spectrum analysis and characteristic sound event detection of audio data. Step 3: Input the preprocessed multimodal data into a pre-trained multi-channel spatiotemporal feature fusion neural network model. This model includes parallel visual feature extraction branches, thermodynamic feature extraction branches, geometric motion feature extraction branches, and acoustic feature extraction branches. The visual feature extraction branch uses a deep convolutional network to extract static and dynamic phenotypic features related to health status from visible light image sequences, including but not limited to the state of periorbital secretions, nasal mirror moisture, coat gloss and smoothness, anal area cleanliness, and postural stability when standing or lying down. The thermodynamic feature extraction branch extracts the temperature gradient distribution, temperature change trend, and specific patterns of temperature difference with the environment from the thermal imaging sequence of the calf's core body surface area and edge area. The geometric motion feature extraction branch extracts the calf's gait parameters, activity temporal distribution, frequency and smoothness of standing and lying down transitions, and quantitative indicators of interaction behavior with the mother calf or companions from the three-dimensional motion model. The acoustic feature extraction branch extracts the fundamental frequency, harmonic structure, energy envelope, and rhythmic and intensity features of breathing and sucking sounds from the audio stream. Step 4: The fusion layer of the multi-channel spatiotemporal feature fusion neural network model adopts a weighted fusion strategy based on the attention mechanism, dynamically calculates the contribution weight of different modal features to the health status assessment at a specific time point, and concatenates the weighted high-dimensional feature vectors, and then maps them to a multi-dimensional health status assessment vector through a fully connected layer. This vector contains at least three core dimensions of quantitative scores: physiological state index, behavioral vitality index, and stress response index. Step 5: Based on the health status assessment vector, combined with the preset multivariate threshold model established based on a large amount of historical healthy calf data, the health status of the target calf is comprehensively judged and classified, and the classification results including four levels are output: healthy, sub-healthy, need attention, and high disease risk. A structured health assessment report containing specific abnormal feature descriptions, risk probabilities and warning levels is generated. Step 6: Push the health assessment report to the administrator's terminal device in real time via the wireless communication module, and trigger a multi-level early warning mechanism when the assessment result indicates that attention is needed or the disease risk is high.
[0006] As a preferred embodiment of the present invention, the pre-training process of the multi-channel spatiotemporal feature fusion neural network model specifically includes: constructing a large-scale, high-quality, finely labeled multimodal dataset of calves, which contains synchronous multispectral images, depth images, thermal imaging, and audio data of thousands of healthy and sick calves under different seasons, different lighting conditions, and different feeding and management models; and manually annotating key health indicators in the data by experts, with the annotation information covering specific clinical symptoms, behavioral abnormalities, and final diagnostic conclusions; adopting an end-to-end deep learning training paradigm, using the multimodal data as input and the expert-annotated health status as the supervision signal, and using a combination of focal length and depth information. A point loss function and a label smoothing regularization optimizer are used for model training to address class imbalance in the data and improve the model's generalization ability. During training, an adversarial training strategy is introduced to enhance the model's robustness to common disturbances in real-world aquaculture environments by adding imperceptible perturbations to the input data. Simultaneously, a cross-modal self-supervised learning task is used as an auxiliary learning objective, forcing the model to learn deep semantic relationships between different modalities, thereby improving the quality of feature representation and fusion effects. The model performance is periodically evaluated on independent validation sets, and the model structure is pruned and optimized based on the evaluation results to ensure efficient operation on resource-constrained edge computing nodes.
[0007] As a preferred embodiment of the present invention, the specific implementation of the attention-based weighted fusion strategy is as follows: a temporal attention sub-network is trained for the output feature sequence of each feature branch. This sub-network can analyze the dynamic changes of the feature sequence in the time dimension and calculate the importance weight of the feature vector at each time step. Simultaneously, a cross-modal attention mechanism is set up. This mechanism takes the feature vectors of all modalities at the current time as input and calculates an inter-modal attention weight vector through a lightweight neural network. This weight vector reflects the relative importance of the visual modality, thermodynamic modality, geometric motion modality, and acoustic modality for the overall health status assessment at the current time. The temporal attention weight and the cross-modal attention weight are multiplied element-wise to obtain the final fusion weight of each modality at each time step. The original features are weighted and summed using these weights to generate a fusion feature representation that contains rich temporal information and highlights the contribution of key modalities. This fusion feature representation is then input into the subsequent fully connected layer for classification and regression.
[0008] As a preferred embodiment of the present invention, the method further includes a continuous online learning and model adaptive optimization step: after system deployment, follow-up data of calves assessed as "requiring attention" or "high disease risk" are continuously collected, including records of human intervention, veterinary diagnosis results, and recovery status, and these data are used as new labeled samples; an incremental learning framework is established, and the deployed multi-channel spatiotemporal feature fusion neural network model is fine-tuned periodically using newly collected high-value samples, while knowledge distillation technology is used to prevent the model from forgetting old knowledge when learning new knowledge, so that the assessment model can adapt to the breeding environment, population characteristics, and disease trends of a specific cattle farm, and achieve continuous improvement in assessment accuracy and personalized customization.
[0009] As a preferred embodiment of the present invention, the method, when assessing the health status of calves, adds a specific analysis module for suckling behavior quality, taking into account the physiological characteristics of lactating calves: by analyzing the suckling sounds collected by the high-sensitivity directional audio acquisition unit, combined with the calf's head posture data relative to the mother calf's udder obtained by the depth sensing unit, the start and end of each suckling event are accurately identified, and the duration, suckling rate, sucking force intensity index, and sucking interval pattern of a single suckling are calculated; the extracted sucking behavior characteristics are compared with the baseline pattern of healthy calves of the same age, and if weak sucking, excessively short sucking time, abnormally prolonged sucking interval, or significantly reduced sucking behavior frequency is detected, it is used as an important negative indicator, significantly increasing the probability of determining a high level of concern or disease risk, thereby achieving early and sensitive assessment of the calf's nutritional intake status and vitality.
[0010] As a preferred embodiment of the present invention, it further includes a contactless assessment system for the health status of suckling calves. The system comprises: a data acquisition module, physically consisting of a rigid or flexible support, and a multispectral imaging unit, a depth sensing unit, and a high-sensitivity directional audio acquisition unit mounted on the support. The support is designed to be fixed above or to the side of the calf island, feeding pen, or suckling area, ensuring comprehensive coverage of the calf's main activity area. All sensing units are dustproof, waterproof, and corrosion-resistant to adapt to the breeding environment. The data acquisition module integrates a first-level data processing chip for initial data compression and time synchronization. An edge computing and analysis module, the core of which is the edge computing node, employs a system-on-a-chip with high-performance AI inference capabilities, encapsulated in an industrial-grade protective shell, and connected via wired or... The system is wirelessly connected to the data acquisition module, which stores and runs the multi-channel spatiotemporal feature fusion neural network model and the multivariate threshold model. It is responsible for receiving preprocessed data streams, performing complex feature extraction and fusion calculations, and completing comprehensive health status assessment and report generation. The cloud platform and user interaction module includes a cloud server cluster and user terminal applications. The edge computing and analysis module uploads the generated health assessment report and key compressed data to the cloud platform via a cellular network or local area network. The cloud platform is responsible for long-term data storage, big data analysis, visualization of group health trends, and generation of group management suggestion reports. Users receive real-time alerts, view individual and group calves' health details, historical data curves, and management guidance suggestions provided by the system through a dedicated application or web interface installed on their mobile phones, tablets, or computers.
[0011] As a preferred embodiment of the present invention, the support frame adopts an adjustable-height gantry or cantilever beam structure, on which a multi-degree-of-freedom gimbal is provided, allowing remote adjustment of the pitch, yaw, and roll angles of each sensing unit to achieve precise tracking of single or multiple calf targets; the high-sensitivity directional audio acquisition unit is not a single microphone, but a group of microphones arranged in a specific geometric array, combined with beamforming algorithms, which can effectively enhance the sound signal from the direction of the calf in the noisy background noise of the farm, and suppress interference noise from other directions such as feed mixers and fans; in addition, the multispectral imaging unit and the depth sensing unit are designed with a common optical path or near optical path in physical installation, and calibration ensures accurate spatial registration of different modal data, providing a solid foundation for subsequent multimodal feature fusion.
[0012] As a preferred embodiment of the present invention, the system also integrates a rule-based and data-driven dual-verification and decision support subsystem. This subsystem has a built-in expert knowledge base that stores typical multimodal feature pattern rules for common diseases of calves of different ages and breeds. When the evaluation result output by the multi-channel spatiotemporal feature fusion neural network model highly matches a rule in the expert knowledge base, the system will add a probability indication of the disease type corresponding to that rule to the evaluation report. At the same time, the system is connected to the farm's existing management information system, which can obtain static data such as birth weight, age, pedigree, and milk production of the calves. This static data is used as auxiliary features and input together with the health status evaluation vector into a lightweight decision tree or logistic regression model for secondary verification and comprehensive decision-making, so as to further improve the accuracy and reliability of the evaluation results and provide more targeted action suggestions for managers.
[0013] By adopting the above technical solution, the present invention has the following advantages: 1. This invention utilizes a non-contact fusion application of multispectral, depth sensing, and directional audio acquisition technologies. The system can acquire richer and more objective physiological and behavioral data than manual observation without disturbing the calf's natural behavior. This effectively avoids stress reactions caused by capture, restraint, and other operations, thereby creating a more favorable welfare environment for the calf's growth. At the same time, the collected data better reflects its true state, laying a solid foundation for accurate assessment.
[0014] 2. This invention utilizes a multi-channel spatiotemporal feature fusion neural network model, enabling the system to analyze multimodal information such as body surface morphology, body temperature distribution, three-dimensional kinetics, and acoustic features in parallel. By dynamically integrating the contributions of these information through an attention mechanism, the system can not only identify obvious clinical symptoms but also keenly capture extremely subtle early abnormal signals such as slight gait changes, weakened sucking force, and local abnormal body temperature. This allows for a shift from treating existing diseases to preventing future diseases, significantly advancing the window for health intervention.
[0015] 3. This invention, by processing data in real time through edge computing and automatically generating diagnostic reports and early warnings, frees managers from the heavy and inefficient daily inspections, allowing them to focus on handling high-risk situations that require intervention. This greatly improves labor productivity. At the same time, the system's continuous online learning capability enables it to adapt to the population and environmental characteristics of specific ranches, becoming more and more accurate with use. This forms a closed-loop intelligent management system with self-optimization capabilities, providing strong technical support for the digital upgrade of large-scale and intensive farming. Attached Figure Description
[0016] Figure 1 This invention is for the purpose of presenting a product. Detailed Implementation
[0017] References to embodiments herein mean that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0018] The non-contact assessment system of this invention is a closed-loop intelligent decision-making system integrating advanced sensing technology, edge computing, and artificial intelligence. Its core lies in the precise collaborative work of multiple modules. The data acquisition module captures the visual, thermodynamic, geometric spatial, and acoustic information of calves from the physical world simultaneously through multispectral cameras, depth sensors, and directional microphone arrays, forming a raw data stream. The edge computing and analysis module is deployed on-site at the farm. After receiving the data stream, it performs heavy computational tasks, including data alignment and preprocessing, as well as using a pre-trained multi-channel spatiotemporal feature fusion neural network model to extract deep features from multimodal data and perform intelligent fusion based on an attention mechanism, ultimately outputting a quantitative health status assessment vector and classification results. The cloud platform and user interaction module are responsible for aggregating data from multiple edge nodes, performing long-term storage, macro-trend analysis, and visualization. At the same time, it pushes personalized alerts, detailed reports, and management suggestions to users in real time, completing the last mile from data to decision-making. In addition, continuous online learning and optimization constitute the system's self-evolution mechanism, using user feedback and subsequent diagnostic results as new learning samples. Through incremental learning technology, the evaluation model on the edge side is continuously optimized, making the system smarter with use. These modules work closely together to achieve a fully automated and intelligent closed loop from perception, analysis, decision-making to feedback optimization, ensuring 24 / 7 uninterrupted and stress-free accurate monitoring of the health status of lactating calves.
[0019] Specifically, a non-contact assessment method for the health status of lactating calves, such as... Figure 1 As shown, it includes: Step 1: Using a multispectral imaging unit, a depth sensing unit, and a high-sensitivity directional audio acquisition unit deployed above the calf's living area, non-contact multimodal raw data streams of the target calf are collected synchronously and continuously within a preset monitoring period. The multispectral imaging unit covers at least the visible light band and the long-wave infrared thermal imaging band to obtain the calf's body surface color texture morphology information and body surface temperature distribution information. The depth sensing unit is used to obtain the calf's three-dimensional point cloud data to reconstruct its accurate body contour and spatial movement trajectory. The high-sensitivity directional audio acquisition unit is used to separate and collect the calf's unique calls, breathing sounds, and sucking sounds after noise reduction. Step 2: Transmit the multimodal raw data stream to the edge computing node for timestamp alignment and data preprocessing. Preprocessing includes illumination normalization and background subtraction of image data to highlight the calf target, environmental temperature compensation calibration of thermal imaging data, point cloud filtering and registration of depth data to construct a coherent 3D motion model, and spectrum analysis and characteristic sound event detection of audio data. Step 3: Input the preprocessed multimodal data into a pre-trained multi-channel spatiotemporal feature fusion neural network model. This model includes parallel visual feature extraction branches, thermodynamic feature extraction branches, geometric motion feature extraction branches, and acoustic feature extraction branches. The visual feature extraction branch uses a deep convolutional network to extract static and dynamic phenotypic features related to health status from visible light image sequences, including but not limited to the state of periorbital secretions, nasal mirror moisture, coat gloss and smoothness, anal area cleanliness, and postural stability when standing or lying down. The thermodynamic feature extraction branch extracts the temperature gradient distribution, temperature change trend, and specific patterns of temperature difference with the environment from the thermal imaging sequence of the calf's core body surface area and edge area. The geometric motion feature extraction branch extracts the calf's gait parameters, activity temporal distribution, frequency and smoothness of standing and lying down transitions, and quantitative indicators of interaction behavior with the mother calf or companions from the three-dimensional motion model. The acoustic feature extraction branch extracts the fundamental frequency, harmonic structure, energy envelope, and rhythmic and intensity features of breathing and sucking sounds from the audio stream. Step 4: The fusion layer of the multi-channel spatiotemporal feature fusion neural network model adopts a weighted fusion strategy based on the attention mechanism. It dynamically calculates the contribution weight of different modal features to the health status assessment at a specific time point, and concatenates the weighted high-dimensional feature vectors. Then, it maps them to a multi-dimensional health status assessment vector through a fully connected layer. This vector contains at least three core dimensions of quantitative scores: physiological state index, behavioral vitality index, and stress response index. Step 5: Based on the health status assessment vector, combined with the preset multivariate threshold model built on a large amount of historical healthy calf data, the health status of the target calf is comprehensively judged and classified, and the classification results include four levels: healthy, sub-healthy, need attention, and high disease risk. A structured health assessment report containing specific abnormal feature descriptions, risk probabilities and warning levels is generated. Step 6: Push the health assessment report to the administrator's terminal device in real time via the wireless communication module, and trigger a multi-level early warning mechanism when the assessment result indicates that attention is needed or the disease risk is high.
[0020] It should be noted that the pre-training process of the multi-channel spatiotemporal feature fusion neural network model specifically includes: constructing a large-scale, high-quality, finely labeled multimodal dataset of calves. This dataset contains synchronous multispectral images, depth images, thermal images, and audio data of thousands of healthy and sick calves under different seasons, lighting conditions, and feeding and management models. Key health indicators in the data are manually labeled by experts, with annotations covering specific clinical symptoms, behavioral abnormalities, and final diagnostic conclusions. An end-to-end deep learning training paradigm is adopted, using multimodal data as input and expert-labeled health status as supervision signals, combining... A focus loss function and a label smoothing regularization optimizer are used for model training to address class imbalance in the data and improve the model's generalization ability. During training, an adversarial training strategy is introduced to enhance the model's robustness to common disturbances in real-world aquaculture environments by adding subtle perturbations to the input data. Simultaneously, a cross-modal self-supervised learning task is used as an auxiliary learning objective, forcing the model to learn deep semantic relationships between different modalities, thereby improving the quality of feature representations and fusion effects. Model performance is periodically evaluated on independent validation sets, and the model structure is pruned and optimized based on the evaluation results to ensure its performance even under resource constraints. The efficient operation on computing nodes and the specific implementation of the attention-based weighted fusion strategy are as follows: A temporal attention sub-network is trained for the output feature sequence of each feature branch. This sub-network analyzes the dynamic changes of the feature sequence over time and calculates the importance weights of the feature vector at each time step. Simultaneously, a cross-modal attention mechanism is set up. This mechanism takes the feature vectors of all modalities at the current time as input and calculates an inter-modal attention weight vector through a lightweight neural network. This weight vector reflects the contribution of the visual modality, thermodynamic modality, geometric motion modality, and acoustic modality to the overall health status assessment at the current time. The relative importance is determined by multiplying the temporal attention weights and cross-modal attention weights element-wise to obtain the final fusion weights for each modality at each time step. These weights are then used to perform a weighted summation of the original features to generate a fusion feature representation that contains rich temporal information and highlights the contributions of key modalities. This fusion feature representation is then input into the subsequent fully connected layer for classification and regression. The method also includes a continuous online learning and model adaptive optimization step: after system deployment, follow-up tracking data of calves assessed as "needing attention" or "high disease risk" are continuously collected, including records of human intervention, veterinary diagnosis results, and rehabilitation status. This data is then used as new labeled samples.An incremental learning framework is established to periodically fine-tune the deployed multi-channel spatiotemporal feature fusion neural network model using newly collected high-value samples. Simultaneously, knowledge distillation technology is employed to prevent the model from forgetting old knowledge while learning new knowledge. This allows the evaluation model to adapt to the specific cattle farm's breeding environment, population characteristics, and disease trends, achieving continuous improvement in evaluation accuracy and personalized customization. When assessing the health status of calves, a specific analysis module for suckling behavior quality is added, taking into account the physiological characteristics of lactating calves: by analyzing suckling sounds collected by a high-sensitivity directional audio acquisition unit, combined with... The depth sensing unit acquires the calf's head posture data relative to the mother calf's udder, accurately identifying the start and end of each suckling event, and calculating the duration, rate, intensity index, and interval pattern of each suckling action. The extracted sucking behavior characteristics are compared with the baseline pattern of healthy calves of the same age. If weak sucking, excessively short sucking time, abnormally prolonged sucking intervals, or significantly reduced sucking frequency are detected, these are considered important negative indicators, significantly increasing the probability of determining a high level of concern or disease risk. This enables early and sensitive assessment of the calf's nutritional intake and vitality.
[0021] A contactless health assessment system for suckling calves includes: a data acquisition module, physically consisting of a rigid or flexible support frame and a multispectral imaging unit, a depth sensing unit, and a high-sensitivity directional audio acquisition unit mounted on the frame. The support frame is designed to be fixed above or to the side of the calf island, feeding pen, or suckling area, ensuring comprehensive coverage of the calf's main activity areas. All sensing units are dustproof, waterproof, and corrosion-resistant to adapt to the farming environment. The data acquisition module integrates a first-level data processing chip for initial data compression and time synchronization. An edge computing and analysis module, with an edge computing node at its core, uses a system-on-a-chip (SoC) with high-performance AI inference capabilities, encapsulated in an industrial-grade protective shell, and connects to the data acquisition module via wired or wireless means. This module stores and runs a multi-channel spatiotemporal feature fusion neural network model and a multivariate threshold model, responsible for receiving pre-processed data streams, performing complex feature extraction and fusion calculations, and completing comprehensive health status assessment and report generation. A cloud platform and user interface are also included. The interactive module, including a cloud server cluster and user terminal applications, and the edge computing and analysis module, upload the generated health assessment reports and key compressed data to the cloud platform via cellular networks or local area networks. The cloud platform is responsible for long-term data storage, big data analysis, visualization of group health trends, and generation of group management suggestion reports. Users can receive real-time alerts, view individual and group calf health details, historical data curves, and management guidance suggestions provided by the system through dedicated applications or web interfaces installed on mobile phones, tablets, or computers. The support frame adopts an adjustable-height gantry or cantilever beam structure, which is equipped with a multi-degree-of-freedom gimbal, allowing remote adjustment of the pitch, yaw, and roll angles of each sensing unit to achieve accurate tracking of single or multiple calf targets. The high-sensitivity directional audio acquisition unit is not a single microphone, but a group of microphones arranged in a specific geometric array. Combined with beamforming algorithms, it can effectively enhance the sound signal from the direction of the calf in the noisy background noise of the farm and suppress interference noise from other directions such as feed mixers and fans.Furthermore, the multispectral imaging unit and depth sensing unit are physically installed using a shared or near-optical path design, and calibration ensures accurate spatial registration of different modal data, providing a solid foundation for subsequent multimodal feature fusion. The system also integrates a rule-based and data-driven dual-verification and decision support subsystem: this subsystem has a built-in expert knowledge base storing typical multimodal feature pattern rules for common diseases in calves of different ages and breeds. When the evaluation result output by the multi-channel spatiotemporal feature fusion neural network model highly matches a rule in the expert knowledge base, the system adds a probability indication of the disease type corresponding to that rule to the evaluation report. Simultaneously, the system connects to the farm's existing management information system, enabling it to acquire static data such as birth weight, age, pedigree, and milk production of the calves. This static data is used as auxiliary features, input along with the health status assessment vector into a lightweight decision tree or logistic regression model for secondary verification and comprehensive decision-making, further improving the accuracy and reliability of the evaluation results and providing more targeted action suggestions for managers.
[0022] In summary, the system utilizes a data acquisition module deployed above the calf living areas (such as calf islands and nursing pens). The core of this module consists of a multispectral imaging unit, a depth sensing unit, and a high-sensitivity directional audio acquisition unit. These units are precisely calibrated and time-synchronized to ensure strict alignment of the captured data in both time and space. The multispectral imaging unit operates simultaneously in the visible light and long-wave infrared bands. The visible light camera continuously records high-definition video sequences under natural or supplemental lighting conditions. From these recordings, details of the calf's body surface morphology can be analyzed, such as whether there are abnormal secretions or redness on the eyelids, whether the dryness or wetness of the nose meets health standards, whether the coat's luster and smoothness have become rough and messy due to malnutrition or disease, and whether there is diarrhea or soiling in the anal and tail areas. The system detects traces of dyeing and assesses the coordination and stability of the calf's trunk and limbs when standing, lying down, or walking. A long-wave infrared thermal imaging camera passively receives the infrared radiation emitted from the calf's body surface and converts it into a high-resolution temperature distribution map. By analyzing this temperature map, the temperature of key areas such as the base of the ears, eyes, nose, and extremities can be monitored non-contactly, along with the temperature gradient distribution between the trunk and limbs. Any abnormal increase or decrease in local temperature, or a change in the overall temperature distribution pattern, could be an early signal of inflammation, circulatory disorders, or stress. The depth sensing unit actively projects invisible structured light or uses the time-of-flight principle to generate a real-time depth information map of the scene, thereby constructing a dynamic three-dimensional point cloud model of the calf. This model not only... This system can accurately calculate the calf's height, length, chest circumference, and other physical dimensions for growth and development assessment. It can also precisely track the calf's spatial location at any given moment, calculate its movement speed and gait characteristics such as stride length, cadence, and joint flexion / extension angles, and quantify its total activity level, frequency of standing and lying down transitions, and the smoothness of its movements over a certain period. Furthermore, it can accurately track the distance and frequency of social interactions with other calves or cows. The high-sensitivity directional audio acquisition unit typically uses microphone array technology, employing beamforming algorithms to create an acoustic focal point that precisely targets the calf. This effectively separates and enhances the calf's own sound signals amidst strong background noise, including the vocal characteristics of hunger, discomfort, or pain, as well as abnormal sounds such as panting and coughing that may occur during breathing. And most importantly, the rhythmic sucking and swallowing sounds produced when suckling at the mother's breast. All these multimodal raw data streams are transmitted in real time to the edge computing and analysis module, which is usually located in a waterproof and dustproof enclosure near the livestock shed. First, rigorous data preprocessing and time-series alignment are performed. The visible light images are automatically white-balanced and contrast-enhanced to eliminate the influence of lighting changes, and a background subtraction algorithm is used to accurately segment the calf target from the complex background. The thermal imaging data is calibrated based on real-time acquired environmental temperature reference points to eliminate interference from environmental heat radiation and ensure the accuracy of body surface temperature measurement. The depth data undergoes point cloud denoising, filtering, and inter-frame registration, thereby constructing a smooth and coherent three-dimensional movement trajectory and body posture change sequence of the calf.The audio data undergoes noise reduction, pre-emphasis, and frame segmentation. It is then converted into a time-spectrum image using a short-time Fourier transform, and an event detection algorithm is used to identify key acoustic events such as vocalizations, breathing sounds, and sucking sounds. The preprocessed multimodal data is then fed into the system's core, designed with multiple parallel branches to process different modalities: the vision branch is typically a 3D convolutional neural network or a combination of convolutional and recurrent neural networks, used to extract health-related spatiotemporal features from visible light image sequences; for example, it can learn that "a healthy calf always has clean eyes." This feature is distinguished from the pattern of "purulent discharge from the corners of the eyes of sick calves." The thermodynamics branch analyzes thermal imaging sequences to capture dynamic changes in the distribution of body surface temperature, such as identifying abnormal temperature hotspots in the ears caused by local infection. The geometric motion branch processes three-dimensional point cloud sequences to quantify kinematic parameters, such as identifying behavioral abnormalities like unsteady gait and difficulty getting up. The acoustics branch analyzes the time-spectrum characteristics of audio, extracting pitch and timbre features of vocalizations to determine emotional state, analyzing the clarity of breathing sounds to determine respiratory health, and precisely quantifying the rhythm and intensity of sucking behavior. The high-dimensional features extracted by these parallel branches are fed into a fusion layer based on an attention mechanism, which is key to the model's intelligence. This mechanism dynamically evaluates which modal information is more important at a specific moment for the current calf's specific performance. For example, when the calf is stationary, thermal imaging and breathing sounds may provide more information, while when it is moving, depth data may contribute more. The attention mechanism assigns appropriate weights to the features of each modality, thereby achieving focused and adaptive feature fusion. The fused high-level feature vector is then nonlinearly transformed through a fully connected layer, ultimately mapping to a comprehensive health status assessment vector. This vector typically includes quantitative scores across multiple dimensions, such as "physiological state index," "behavioral vitality index," and "stress response index." The system compares this assessment vector with a preset multivariate threshold model. This threshold model is based on a normal range benchmark established from massive historical healthy calf data. Through comparison, the system automatically classifies the calf's health status into one of four levels: "healthy," "sub-healthy," "requiring attention," or "high disease risk," and generates a structured assessment report. The report not only includes the level determination but also lists the main abnormal features leading to the determination and their confidence levels. Finally, this report is sent wirelessly to the cloud platform and user interaction module. The cloud platform is responsible for data storage, archiving, and group-level trend analysis. Users can receive alerts and detailed reports in real time via a mobile app or computer webpage, enabling immediate intervention for calves exhibiting health risks. The system also includes a continuous learning loop. After administrators diagnose and process calves alerted by the system, they can feed the final diagnostic results back to the system. The system utilizes this data with new labels.By employing incremental learning technology, the neural network model on the edge side is fine-tuned without forgetting old knowledge, allowing it to continuously adapt to the specific environment and population characteristics of the farm. This enables continuous self-evolution of assessment accuracy. The entire working principle chain starts with contactless data collection, proceeds to intelligent analysis, precise judgment, and real-time early warning, ultimately forming a feedback-based optimization loop. This achieves 24 / 7 automated, high-precision intelligent monitoring of the health status of nursing calves.
[0023] The above embodiments are provided for those skilled in the art to implement or use the present invention. Those skilled in the art can make various modifications or changes to the above embodiments without departing from the spirit of the present invention. Therefore, the scope of protection of the present invention is not limited to the above embodiments, but should be the maximum scope that conforms to the innovative features mentioned in the claims.
Claims
1. A non-contact assessment method for the health status of lactating calves, characterized in that, Including: Step 1: Using a multispectral imaging unit, a depth sensing unit, and a high-sensitivity directional audio acquisition unit deployed above the calf's living area, non-contact multimodal raw data streams of the target calf are collected synchronously and continuously within a preset monitoring period. The multispectral imaging unit covers at least the visible light band and the long-wave infrared thermal imaging band to obtain the calf's body surface color texture morphology information and body surface temperature distribution information. The depth sensing unit is used to obtain the calf's three-dimensional point cloud data to reconstruct its accurate body contour and spatial movement trajectory. The high-sensitivity directional audio acquisition unit is used to separate and collect the calf's unique calls, breathing sounds, and sucking sounds after noise reduction. Step 2: Transmit the multimodal raw data stream to the edge computing node for timestamp alignment and data preprocessing. The preprocessing includes illumination normalization and background subtraction of image data to highlight the calf target, environmental temperature compensation calibration of thermal imaging data, point cloud filtering and registration of depth data to construct a coherent three-dimensional motion model, and spectrum analysis and characteristic sound event detection of audio data. Step 3: Input the preprocessed multimodal data into a pre-trained multi-channel spatiotemporal feature fusion neural network model. This model includes parallel visual feature extraction branches, thermodynamic feature extraction branches, geometric motion feature extraction branches, and acoustic feature extraction branches. The visual feature extraction branch uses a deep convolutional network to extract static and dynamic phenotypic features related to health status from visible light image sequences, including but not limited to the state of periorbital secretions, nasal mirror moisture, coat gloss and smoothness, anal area cleanliness, and postural stability when standing or lying down. The thermodynamic feature extraction branch extracts the temperature gradient distribution, temperature change trend, and specific patterns of temperature difference with the environment from the thermal imaging sequence of the calf's core body surface area and edge area. The geometric motion feature extraction branch extracts the calf's gait parameters, activity temporal distribution, frequency and smoothness of standing and lying down transitions, and quantitative indicators of interaction behavior with the mother calf or companions from the three-dimensional motion model. The acoustic feature extraction branch extracts the fundamental frequency, harmonic structure, energy envelope, and rhythmic and intensity features of breathing and sucking sounds from the audio stream. Step 4: The fusion layer of the multi-channel spatiotemporal feature fusion neural network model adopts a weighted fusion strategy based on the attention mechanism, dynamically calculates the contribution weight of different modal features to the health status assessment at a specific time point, and concatenates the weighted high-dimensional feature vectors, and then maps them to a multi-dimensional health status assessment vector through a fully connected layer. This vector contains at least three core dimensions of quantitative scores: physiological state index, behavioral vitality index, and stress response index. Step 5: Based on the health status assessment vector, combined with the preset multivariate threshold model established based on a large amount of historical healthy calf data, the health status of the target calf is comprehensively judged and classified, and the classification results including four levels are output: healthy, sub-healthy, need attention, and high disease risk. A structured health assessment report containing specific abnormal feature descriptions, risk probabilities and warning levels is generated. Step 6: Push the health assessment report to the administrator's terminal device in real time via the wireless communication module, and trigger a multi-level early warning mechanism when the assessment result indicates that attention is needed or the disease risk is high.
2. A non-contact assessment method for the health status of lactating calves as described in claim 1, characterized in that, The pre-training process of the multi-channel spatiotemporal feature fusion neural network model specifically includes: constructing a large-scale, high-quality, finely labeled multimodal dataset of calves. This dataset contains synchronous multispectral images, depth images, thermal imaging, and audio data of thousands of healthy and sick calves under different seasons, lighting conditions, and feeding and management models. Key health indicators in the data are manually labeled by experts, with the labeling information covering specific clinical symptoms, behavioral abnormalities, and final diagnostic conclusions. An end-to-end deep learning training paradigm is adopted, using the multimodal data as input and the expert-labeled health status as the supervision signal, employing a combination of focus loss function and... A label smoothing regularization optimizer is used for model training to address class imbalance in the data and improve the model's generalization ability. During training, an adversarial training strategy is introduced to enhance the model's robustness to common disturbances in real-world aquaculture environments by adding imperceptible perturbations to the input data. Simultaneously, a cross-modal self-supervised learning task is used as an auxiliary learning objective, forcing the model to learn deep semantic relationships between different modalities, thereby improving the quality of feature representation and fusion effect. The model performance is evaluated regularly on independent validation sets, and the model structure is pruned and optimized based on the evaluation results to ensure that it can run efficiently on resource-constrained edge computing nodes.
3. A non-contact assessment method for the health status of lactating calves as described in claim 1, characterized in that, The specific implementation of the attention-based weighted fusion strategy is as follows: a temporal attention sub-network is trained for the output feature sequence of each feature branch. This sub-network can analyze the dynamic changes of the feature sequence in the time dimension and calculate the importance weight of the feature vector at each time step. Simultaneously, a cross-modal attention mechanism is set up. This mechanism takes the feature vectors of all modalities at the current time as input and calculates an inter-modal attention weight vector through a lightweight neural network. This weight vector reflects the relative importance of the visual modality, thermodynamic modality, geometric motion modality, and acoustic modality for the overall health status assessment at the current time. The temporal attention weight and the cross-modal attention weight are multiplied element-wise to obtain the final fusion weight of each modality at each time step. These weights are used to perform a weighted summation of the original features to generate a fusion feature representation that contains rich temporal information and highlights the contributions of key modalities. This fusion feature representation is then input into subsequent fully connected layers for classification and regression.
4. A non-contact assessment method for the health status of lactating calves as described in claim 1, characterized in that, The method also includes a continuous online learning and model adaptive optimization process: after system deployment, follow-up data of calves assessed as "needing attention" or "high disease risk" are continuously collected, including records of human intervention, veterinary diagnosis results, and recovery status, and these data are used as new labeled samples; an incremental learning framework is established, and the deployed multi-channel spatiotemporal feature fusion neural network model is fine-tuned periodically using newly collected high-value samples. At the same time, knowledge distillation technology is used to prevent the model from forgetting old knowledge when learning new knowledge, so that the assessment model can adapt to the breeding environment, population characteristics, and disease trends of a specific cattle farm, and achieve continuous improvement in assessment accuracy and personalized customization.
5. A non-contact assessment method for the health status of lactating calves according to claim 1, characterized in that, The method, when assessing the health status of calves, incorporates a specialized analysis module for suckling behavior quality, taking into account the physiological characteristics of lactating calves. By analyzing suckling sounds collected by the high-sensitivity directional audio acquisition unit and combining this with the calf's head posture data relative to the mother calf's udder obtained by the depth sensing unit, the method accurately identifies the start and end of each suckling event and calculates the duration, rate, intensity index, and interval pattern of each suckling action. The extracted sucking behavior characteristics are compared with the baseline pattern of healthy calves of the same age. If weak sucking, excessively short sucking time, abnormally prolonged sucking intervals, or significantly reduced sucking frequency are detected, these are considered important negative indicators, significantly increasing the probability of determining a high level of concern or disease risk. This enables early and sensitive assessment of the calf's nutritional intake and vitality.
6. A non-contact assessment method for the health status of lactating calves according to claim 1, characterized in that, It also includes a contactless assessment system for the health status of suckling calves. This system comprises: a data acquisition module, physically consisting of a rigid or flexible support, and a multispectral imaging unit, a depth sensing unit, and a high-sensitivity directional audio acquisition unit mounted on the support. The support is designed to be fixed above or to the side of the calf island, feeding pen, or suckling area, ensuring comprehensive coverage of the calf's main activity area. All sensing units are dustproof, waterproof, and corrosion-resistant to adapt to the breeding environment. The data acquisition module integrates a first-level data processing chip for initial data compression and time synchronization. An edge computing and analysis module, the core of which is the edge computing node, uses a system-on-a-chip with high-performance AI inference capabilities, encapsulated in an industrial-grade protective shell, and communicates with the system via wired or wireless means. The data acquisition module is connected, which stores and runs the multi-channel spatiotemporal feature fusion neural network model and the multivariate threshold model. It is responsible for receiving the preprocessed data stream, performing complex feature extraction and fusion calculations, and completing the comprehensive health status assessment and report generation. The cloud platform and user interaction module includes a cloud server cluster and user terminal applications. The edge computing and analysis module uploads the generated health assessment report and key compressed data to the cloud platform through a cellular network or local area network. The cloud platform is responsible for long-term data storage, big data analysis, visualization of group health trends, and generation of group management suggestion reports. Users can receive real-time alerts, view individual and group calves' health details, historical data curves, and management guidance suggestions provided by the system through a dedicated application or web interface installed on a mobile phone, tablet, or computer.
7. The contactless health status assessment system for lactating calves according to claim 6, characterized in that, The support frame adopts an adjustable-height gantry or cantilever beam structure, equipped with a multi-degree-of-freedom gimbal, allowing remote adjustment of the pitch, yaw, and roll angles of each sensing unit to achieve precise tracking of single or multiple calf targets. The high-sensitivity directional audio acquisition unit is not a single microphone, but a group of microphones arranged in a specific geometric array. Combined with beamforming algorithms, it can effectively enhance the sound signal from the direction of the calf in the noisy background noise of the farm, and suppress interference noise from other directions such as feed mixers and fans. In addition, the multispectral imaging unit and the depth sensing unit are designed with a common optical path or near-optical path in their physical installation, and calibration ensures accurate spatial registration of different modal data, providing a solid foundation for subsequent multimodal feature fusion.
8. The contactless assessment system for the health status of lactating calves according to claim 6, characterized in that, The system also integrates a rule-based and data-driven dual-verification and decision support subsystem. This subsystem has a built-in expert knowledge base that stores typical multimodal feature pattern rules for common diseases of calves of different ages and breeds. When the evaluation result output by the multi-channel spatiotemporal feature fusion neural network model highly matches a rule in the expert knowledge base, the system will add a probability indication of the disease type corresponding to that rule to the evaluation report. At the same time, the system connects to the farm's existing management information system and can obtain static data such as birth weight, age, pedigree, and milk production of mother calves. This static data is used as auxiliary features and input together with the health status evaluation vector into a lightweight decision tree or logistic regression model for secondary verification and comprehensive decision-making. This further improves the accuracy and reliability of the evaluation results and provides managers with more targeted action suggestions.