Milk cow mastitis multi-mode terminal cloud joint defense early warning system

The multimodal cloud-based joint prevention and early warning system for mastitis in dairy cows utilizes wearable devices and biometric visual recognition technology for early and accurate detection of mastitis, solving the problems of poor real-time performance and high energy consumption in existing technologies, and achieving efficient mastitis detection and management.

CN121393897APending Publication Date: 2026-01-23HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511661887.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing technologies for detecting mastitis in dairy cows suffer from poor real-time performance, high energy consumption, and limited detection modalities, resulting in low early detection rates and low farm management efficiency.

Method used

A multimodal edge-cloud joint prevention and early warning system for mastitis in dairy cows is adopted. Wearable devices are used for real-time monitoring of abnormal behavior and initial risk screening. Biometric visual recognition technology is combined for precise individual matching. Multimodal information such as movement data, body temperature indicators and visual features are integrated for cloud-based intelligent diagnosis, and a detection system based on the "edge-cloud" collaborative architecture is constructed.

Benefits of technology

It significantly improved the early detection rate of mastitis in dairy cows and the efficiency of farm management, reduced system energy consumption, and enabled remote monitoring and early warning through a multi-dimensional data visualization platform, thereby improving the robustness and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121393897A_ABST
    Figure CN121393897A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of dairy cow mastitis monitoring, and particularly relates to a dairy cow mastitis multi-mode end cloud joint defense early warning system. The system comprises a risk preliminary screening module used for analyzing collected data through a low-energy-consumption behavior detection method, and completing preliminary screening of mastitis risks and marking of high-risk individuals; the high-risk cattle identification and matching module is used for carrying out matching and identity confirmation on sick high-risk cattle through a cattle body identification method and collecting corresponding video image data; the multi-modal data fusion and diagnosis module is used for fusing and enhancing the motion signals, the body temperature and the video image data features to diagnose mastitis; and the cloud visual monitoring and early warning platform collects all monitoring data, analysis results and early warning information and is used for assisting a manager in early intervention. The method has the characteristics that the early detection rate of the cow mastitis and the pasture management efficiency can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of dairy cow mastitis monitoring, and particularly relates to a dairy cow mastitis multi-modal end-cloud joint defense early warning system. BACKGROUND

[0002] Dairy cow mastitis is the most common and high-incidence disease in dairy cow population, is affected by multiple factors, and is difficult to prevent and control and is prone to recurrence, which has caused great economic losses and management pressure to the dairy farming industry in China. The economic losses caused by the disease cover two levels of direct and indirect: the direct losses mainly manifest as significant reduction in milk production, raw milk waste and high treatment costs; the indirect losses reflect the increase in culling rate and the damage to reproductive performance, which seriously erodes the operating efficiency of the farm.

[0003] Although there are many types of existing dairy cow mastitis detection technologies, there are still the following three key challenges in the actual complex farming scene: (1) The traditional manual monitoring method highly depends on the participation of professional veterinarians, and has the problems of poor real-time performance and high labor cost. This method mainly detects clinical mastitis, which leads to a large number of potential diseased individuals that cannot be identified in time. In addition, the data recording in the traditional monitoring process completely depends on manual completion, and has the defects of easy data loss and difficult system integration, which cannot establish an effective big data analysis system to develop and implement long-term disease prevention and control strategies.

[0004] (2) The existing intelligent monitoring technologies generally have the problem of relying on single modal data, which leads to insufficient detection accuracy. For example, some technologies achieve diagnosis by analyzing the physicochemical properties of milk, but are easily disturbed by external factors such as the diet and drinking water of dairy cows, thereby producing misjudgments; although wearable devices can monitor physiological indicators such as body temperature and behavior in real time, their judgment mainly relies on a single indicator to infer the risk of mastitis, which lacks multi-dimensional collaborative analysis, thereby limiting the reliability of the detection results.

[0005] (3) The current advanced monitoring technologies face the dual challenges of equipment cost and operating cost in popularization and application. For example, the purchase price of special monitoring equipment such as infrared imaging instruments is high, and the sensors deployed in some systems have high energy consumption because they need to work continuously, which significantly increases the overall use cost, restricts the large-scale popularization of advanced monitoring technologies, and ultimately affects the improvement of the overall prevention and control effect of dairy cow mastitis.

[0006] Therefore, it is very important to design a dairy cow mastitis multi-modal end-cloud joint defense early warning system that can significantly improve the early detection rate of dairy cow mastitis and the management efficiency of the farm. SUMMARY

[0007] The present application is to overcome the problems of poor real-time performance, high energy consumption and single detection mode of the existing dairy cow mastitis detection technology in practical application, and provides a dairy cow mastitis multi-modal end-cloud joint defense early warning system which can realize real-time monitoring and risk preliminary screening of abnormal behaviors of dairy cows through wearable devices, complete individual accurate matching combined with biological feature visual recognition technology, and realize cloud intelligent diagnosis by fusing multi-modal information such as motion data, body temperature index and visual features, so as to realize early and accurate detection of dairy cow mastitis.

[0008] In order to achieve the above-mentioned application purposes, the present application adopts the following technical solutions: The dairy cow mastitis multi-modal end-cloud joint defense early warning system comprises: A risk preliminary screening module is used for analyzing the collected data by a low-energy-consumption behavior detection method, completing preliminary screening of mastitis risk and marking high-risk individuals; A high-risk cow identification and matching module is used for matching and identity confirmation of diseased high-risk cows by a cow body identification method, and collecting corresponding video image data; A multi-modal data fusion and diagnosis module is used for fusing and enhancing motion signals, body temperature and video image data features, and realizing diagnosis of mastitis; A cloud visual monitoring and early warning platform is used for collecting all monitoring data, analysis results and early warning information, and assisting managers to carry out early intervention.

[0009] As a preferred, the risk preliminary screening module specifically comprises the following: S11, continuously monitoring and collecting key physiological and behavioral indicators of dairy cows by wearing wearable devices with sensors; the sensors include accelerometers, gyroscopes and thermometers; the key physiological and behavioral indicators include activity amount, resting state, walking frequency and body temperature; S12, designing a low-energy-consumption behavior detection method based on the collected motion signals, for ensuring the robustness of the model under low sampling frequency. As a preferred, the design of the low-energy-consumption behavior detection method based on the collected motion signals specifically comprises the following steps: S121, training a teacher classification network and a reconstruction network using high sampling frequency data; the teacher classification network learns behavior classification ability through cross-entropy loss, and the reconstruction network reconstructs the input high sampling frequency image through mean square error loss, and the specific formula is as follows: (1) ; Wherein, is the cross-entropy loss function of the teacher classification network, represents the number of samples in the data set, indicates that the data set contains high sampling frequency samples and its corresponding label , is a one-hot label vector , th element of which is used to identify whether a sample belongs to the th activity, is the posterior probability of each activity; (2) ; wherein, is the reconstruction loss function of the teacher reconstruction network, is the feature vector obtained by inputting the high sampling frequency sample into the teacher feature extractor, is the generated reconstructed image; S122, input the low sampling frequency data into the student classification network to extract the feature vector , and generate the reconstructed image through the reconstruction network, the specific formula is as follows: (3) ; wherein, represents the information recovery RIR loss based on reconstruction; S123, the correlation matrix , , and , of the th layer feature map of the teacher classification network and the student classification network is calculated along the time dimension and the sensor axis dimension respectively; the correlation of the student feature and the teacher feature is constrained through the normalized cosine similarity, and the total loss is the mean of the loss of each layer, the specific formula is as follows: (4) ; (5) ; (6) ; wherein, and respectively represent the time correlation distillation loss and the inter-axis correlation distillation loss, represents the information recovery loss of the th layer of correlation distillation, represents the information recovery CDIR loss based on correlation distillation, represents the time dimension, respectively represent , the horizontal vector after normalization, represents the total number of layers of the network; S124, combined with the classification loss , the RIR loss and CDIR loss , and by joint loss function The training of the student classification network is implemented: (7) ; S125, statistical index calculation is performed on the obtained cow behavior data, a traditional machine learning model is constructed, and preliminary determination of the risk of cow mastitis is realized.

[0010] As preferred, in the high-risk cow identification and matching module, the cow body identification method specifically includes the following processes: S21, a cow body input picture is given , wherein H, W and C represent the height, width and channel number of the cow body image respectively, the ViT model divides the cow body input image into non-overlapping image blocks by using a sliding window mechanism, the step length of sliding is S, and the length of the image block is P, and the cow body input image with a resolution of HxW is divided into fixed-size image blocks, and the specific formula is as follows: (8) ; , wherein represents a floor operation; S22, the segmented patch blocks are embedded into the input sequence as the representation of the local features of the cow body, and an additional sequence is also embedded into the input sequence to learn the representation of the global features of the cow body; S23, after the input sequence is sent into a layer Transformer encoder, the output features of the first layer are obtained , and the local features and the global features of the cow body are separated according to the channels to obtain the local features and the global features ; S24, the cow body local features are sequentially processed by two convolutional blocks: (9) ; , wherein represents the output of the convolutional spatial local feature aggregation module; represents two convolutional normalization and Mish activation function operations; represents the channel-wise concatenation of the input features; S25, the network is optimized by using a triplet loss and a label smoothing cross-entropy loss (10) ;​ wherein, is a joint loss function for training the cow recognition network, is a hyperparameter, representing the weight relationship between the label smoothing cross-entropy loss and the triplet loss term; (11); wherein, represents a random noise coefficient; represents the number of cow species; represents a real label, 1 for a positive class and 0 for a negative class; represents a predicted probability output by the model; (12); wherein, represents the input of the current cow picture; represents a difficult sample image of the same batch and the same class as ; represents a difficult sample image of the same batch and different classes as ; , ), respectively represent the feature vectors output by the model corresponding to the pictures; represents the square of the 2-norm; represents the minimum interval between the distances between different cow images and the distances between the same cow images.

[0011] As a preferred, in the multi-modal data fusion and diagnosis module, a multi-modal interactive neural network CMI-Net model is used for multi-modal interactive fusion, which specifically includes the following steps: S31, for the input multi-modal data containing motion signals, body temperature data and RGB video, first, a plurality of independent processing branches are used to extract features of each single modal data; each branch realizes the feature extraction function of the corresponding modal based on convolution operation; a Res-LCB module is designed, which is defined as follows: (13); wherein, and respectively represent the feature maps of layer and layer, and respectively represent 1x1 and 1x3 convolution operations, represents a feature addition operation, and RELU(•) represents a rectified linear unit activation function.

[0012] S32, a joint cross-modal interaction module CMIM based on attention mechanism is designed, which uses multi-modal information to adaptively recalibrate the time and axial features of each modality.

[0013] As preferred, the joint cross-modal interaction module CMIM based on attention mechanism specifically includes the following processes: S321, setting and respectively represent the features of the specific layers of the two convolutional layer networks; 、 and respectively represent the channel number of the feature map, the spatial height dimension of the feature map and the spatial width dimension of the feature map; A and G are input features of the joint cross-modal interaction module; S322, average pooling operation is performed along the channel of the input feature to generate two spatial graphs; then the two spatial graphs are mapped and connected, and mapped into a joint representation , the operation is as follows: (14) wherein, represents the channel number of the joint representation , (•) represents the average pooling operation, and [•] represents the connection operation; S323, setting two spatial attention graphs and are generated by putting the joint representation into two independent convolutional layers, and the results are calculated by sigmoid function (•), the specific formula is as follows: (15) ; S324, recalibrating the input features using and to generate two final refined features and : (16) ; wherein, represents the feature multiplication operation.

[0014] As preferred, the cloud visualization monitoring and early warning platform is specifically as follows: The manager can view the health status, behavior record, body temperature curve and early warning notification of each cow in the pasture in real time through the mobile phone App or PC client; The platform provides detailed individual health reports, which cover the daily activities, abnormal behavior records and historical early warning information of the cows; Through the historical data backtracking function of the platform, the manager analyzes the long-term health trend of the cow by viewing a single early warning event.

[0015] Compared with the prior art, the present application has the following beneficial effects: (1) the present application proposes a dairy cow mastitis multi-modal end-cloud joint defense early warning system based on an "end-edge-cloud" collaborative architecture; the system realizes real-time monitoring and risk preliminary screening of abnormal behaviors of dairy cows through wearable devices, completes individual accurate matching in combination with biological feature visual recognition technology, and performs cloud intelligent diagnosis by fusing multi-modal information such as motion data, body temperature indicators and visual features, finally realizes remote monitoring and early warning through a multi-dimensional data visualization platform, and builds a full-chain solution from data perception to intelligent decision-making, which significantly improves the early detection rate of dairy cow mastitis and the efficiency of pasture management; (2) the present application is designed by using a low-power method, and high-performance abnormal behavior detection is realized at a low sampling frequency by optimizing the method, and the overall energy consumption of the system is greatly reduced in combination with a lightweight wearable sensor data acquisition scheme; (3) the present application innovatively fuses multi-modal information such as motion data, body temperature indicators and visual features, and effectively improves the detection robustness and accuracy in complex scenes through a cross-modal feature complementary mechanism; (4) the present application realizes early warning preliminary screening by capturing the change of behavior characteristics before the appearance of clinical symptoms of the disease based on the behavior abnormality mastitis detection triggering mechanism, and realizes early warning of mastitis in combination with multi-modal data detection and verification. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 It is a principle flow chart of the dairy cow mastitis multi-modal end-cloud joint defense early warning system. DETAILED DESCRIPTION

[0017] In order to more clearly illustrate the embodiments of the present application, the specific embodiments of the present application will be described below with reference to the drawings. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings and other embodiments can be obtained by those skilled in the art without creating any creative labor on the basis of these drawings.

[0018] As shown in Figure 1 The present application provides a dairy cow mastitis multi-modal end-cloud joint defense early warning system based on an "end-edge-cloud" collaborative architecture, which comprises: A risk preliminary screening module is used for analyzing the collected data by a low-energy behavior detection method, completing preliminary screening of mastitis risk and marking high-risk individuals; A high-risk cow identification and matching module is used for matching and identity confirmation of diseased high-risk cows by a cow body identification method, and collecting corresponding video image data for subsequent multi-modal fusion detection; A multi-modal data fusion and diagnosis module is used for fusing and enhancing motion signals, body temperature and video image data features to realize diagnosis of mastitis. A cloud visual monitoring and early warning platform is used for collecting all monitoring data, analysis results and early warning information to assist managers in early intervention.

[0019] Further, the risk screening module specifically includes the following: S11, first, through a wearable device integrated with an accelerometer, a gyroscope, a thermometer and the like, key physiological and behavioral indicators of the dairy cow such as activity, rest state, walking frequency and body temperature are continuously monitored and collected for 7x24 hours; S12, a low-energy-consumption behavior detection method is designed based on the collected motion signals (including three-axis acceleration and three-axis angular velocity) to ensure the robustness of the model at a low sampling frequency: The design of the low-energy-consumption behavior detection method based on the collected motion signals specifically includes the following steps: S121, a teacher classification network and a reconstruction network are trained using high sampling frequency data; the teacher classification network learns behavior classification ability through cross-entropy loss, and the reconstruction network reconstructs the input high sampling frequency image through mean square error loss, and the specific formula is as follows: (1) ; Wherein, is the cross-entropy loss function of the teacher classification network, represents the number of samples in the data set, represents a sample pair in the data set containing a high sampling frequency sample and its corresponding label , is the first element in the one-hot label vector , which is used to identify whether a certain sample belongs to the first activity class, is the posterior probability of each activity class; (2) ; Wherein, is the reconstruction loss function of the teacher reconstruction network, is the feature vector obtained by inputting the high sampling frequency sample into the teacher feature extractor, is the generated reconstructed image; S122, low sampling frequency data is input into the student classification network to extract a feature vector , and a reconstructed image is generated through the reconstruction network, and the specific formula is as follows: (3) ;​ wherein, denotes the reconstruction-based information recovery RIR loss; By constraining the similarity of the reconstructed image and the high sampling frequency image, the feature representation capability of the student network is enhanced.

[0020] S123, the first layer feature map, respectively along the time dimension and the sensor axis dimension, the correlation matrix , and , ; by normalizing the cosine similarity constraint, the correlation of the student feature and the teacher feature, the total loss is the mean of each layer loss, the specific formula is as follows: (4); (5); (6); wherein, and respectively represent the time correlation distillation loss and the inter-axis correlation distillation loss, denotes the information recovery loss of the correlation distillation of the layer, denotes the information recovery CDIR loss based on the correlation distillation, denotes the time dimension, respectively represent , the normalized horizontal vector, represent the total number of layers of the network; S124, combined with the classification loss , the RIR loss and the CDIR loss , and through the joint loss function to realize the training of the student classification network: (7); S125, statistical index calculation is performed on the obtained dairy cow behavior data, and a traditional machine learning model such as SVM (Support Vector Machine) is constructed, which is used to realize the preliminary judgment of the dairy cow mastitis risk.

[0021] Further, for the dairy cows exceeding the risk threshold, the system accurately locates the target through the cow body recognition method, and synchronously starts the video image data acquisition to serve the subsequent multi-modal data fusion analysis. In the high-risk cow identification and matching module, the cow body recognition method specifically includes the following processes: S21, given a picture of a cow body input , where H, W, C represent the height, width and channel number of the cow body image respectively, the ViT model divides the cow body input image into image blocks with non-overlapping pixels using a sliding window mechanism, the sliding step is S, and the side length of the image block is P, the cow body input image with resolution HxW is divided into fixed-size image blocks, the specific formula is as follows: (8) ; , where represents the floor operation; S22, the image blocks cut off are embedded into the input sequence as the representation of the local features of the cow body, in addition, an additional sequence is also embedded into the input sequence to learn the representation of the global features of the cow body; S23, after the input sequence is sent into a Transformer encoder, the output features of the first layer are obtained, and the local features and the global features of the cow body are separated according to the channel to obtain the local features and the global features ; S24, the local features of the cow body are processed by two convolutional blocks: (9) ; , where represents the output of the convolutional spatial local feature aggregation module; represents two convolutional operations plus normalization and Mish activation function operations; represents the concatenation of the input features according to the channel; S25, the network is optimized by using a triplet loss and a label smoothing cross-entropy loss : (10) ; , where is the joint loss function for training the cow body recognition network, is a hyperparameter, representing the weight relationship between the label smoothing cross-entropy loss and the triplet loss term, which is set to 2 in this method; (11) ; , where represents a random noise coefficient, which is set to 0.1 in this paper; represents the number of cow species; represents the real label, the positive class is 1 and the negative class is 0; probabilities representing predictions output by the model; (12); wherein, inputs representing current cow pictures; representing difficult sample images of the same batch of the same category; representing difficult sample pictures of the same batch of different categories; representing difficult sample pictures of the same batch of different categories; representing difficult sample pictures of the same batch of different categories; ), respectively represent feature vectors output by the model for corresponding pictures; representing the square of the 2-norm; representing the minimum interval between the distances between different cow images and the distances between the same cow images, which is set to 0.4 in the method.

[0022] Further, in the multi-modal data fusion and diagnosis module, a multi-modal interactive neural network CMI-Net model is used for multi-modal interactive fusion to realize accurate and robust diagnosis of mastitis. After locking the target cow, the system deeply fuses the preliminary screening motion signals and body temperature data with the visual data collected by the RGB camera in the cloud to realize the final accurate diagnosis, which includes the following steps: S31, for the input multi-modal data containing motion signals, body temperature data and RGB video, first, multiple independent processing branches are used to extract features of each single modal data; each branch realizes the feature extraction function of the corresponding modal based on convolution operation; the residual unit in the deep residual network has a behavior similar to a set, and the response amplitude is small, inspired by this, in order to improve the representation ability and robustness of the model, the present application designs a Res-LCB module, which is defined as follows: (13); wherein, and represent feature maps of the layer and the layer respectively, and represent 1x1 and 1x3 convolution operations respectively, RELU(•) represents the rectified linear unit activation function.

[0023] S32, a joint cross-modal interaction module CMIM based on attention mechanism is designed, which uses multi-modal information to adaptively recalibrate the time and axial features of each modal.

[0024] ​Further, the joint cross-modal interaction module CMIM based on the attention mechanism specifically includes the following processes: S321, setting and respectively represent the features of two convolutional network specific layers; 、 and respectively represent the channel number of the feature map, the spatial height dimension of the feature map and the spatial width dimension of the feature map; A and G are input features of the joint cross-modal interaction module; S322, average pooling operation is performed along the channel of the input feature to generate two spatial graphs; then the two spatial graphs are mapped and connected, and mapped into a joint representation , the operation is as follows: (14) wherein, indicates the channel number of the joint representation , (•) indicates the average pooling operation, and [•] indicates the connection operation; S323, setting two spatial attention maps and are generated by putting the joint representation into two independent convolutional layers, and the results are calculated by sigmoid function (•) to generate, the specific formula is as follows: (15) ; S324, using and to recalibrate the input feature to generate two final refined features and : (16) ; wherein, indicates the feature multiplication operation (i.e. the product of the corresponding position elements of two matrices). In particular, batch normalization operation is performed after each convolution operation in the method. The increase of the channel number and the reduction of the spatial dimension are realized by Res-LCB and maximum pooling operation respectively.

[0025] Further, the cloud visual monitoring and early warning platform specifically includes the following: The system will establish a cloud-based visualization platform, aggregating all monitoring data, analysis results, and early warning information to provide comprehensive decision support for farm managers. Managers can use a mobile app or PC client to view the health status, behavior records, body temperature curves, and early warning notifications of every cow on the farm in real time. The platform provides detailed individual health reports, covering the cow's daily activities, abnormal behavior records, and historical early warning information. Through the platform's powerful historical data backtracking function, managers can not only view individual early warning events but also analyze the long-term health trends of cows, enabling earlier prediction and intervention for potential diseases such as mastitis, effectively preventing the spread and deterioration of diseases. This remote, intelligent decision support function transforms farm management from a traditional model relying on manual experience and regular inspections to a proactive and refined management model based on real-time data, greatly improving management efficiency and the overall operational benefits of the farm.

[0026] This invention proposes a multimodal edge-cloud collaborative prevention and early warning system for bovine mastitis based on an "edge-cloud" collaborative architecture. The system utilizes wearable devices for real-time monitoring and initial risk screening of abnormal bovine behavior, combines biometric visual recognition technology for precise individual matching, and integrates multimodal information such as motion data, body temperature indicators, and visual features for intelligent cloud-based diagnosis. Ultimately, this system constructs a multimodal edge-cloud collaborative prevention and early warning system for bovine mastitis based on an "edge-cloud" collaborative architecture, enabling early and accurate detection of bovine mastitis.

[0027] The above description is merely a detailed explanation of preferred embodiments and principles of the present invention. For those skilled in the art, there may be changes in specific implementation methods based on the ideas provided by the present invention, and these changes should also be considered within the scope of protection of the present invention.

Claims

1. A multi-modal end-cloud joint defense early warning system for dairy cow mastitis, characterized in that, The application relates to a dairy cow mastitis risk early warning system based on multi-modal data fusion and diagnosis. The system comprises the following: a risk preliminary screening module for analyzing collected data through a low-energy-consumption behavior detection method to complete preliminary screening of mastitis risks and marking of high-risk individuals; a high-risk cow identification and matching module for matching and identity confirmation of diseased high-risk cows through a cow body identification method and collecting corresponding video image data; a multi-modal data fusion and diagnosis module for fusing and enhancing motion signals, body temperatures and video image data features to realize mastitis diagnosis; 2. The mastitis multi-modal end-cloud joint defense early warning system for dairy cows according to claim 1, characterized in that, and a cloud-based visual monitoring and early warning platform for collecting all monitoring data, analysis results and early warning information to assist managers in early intervention. The risk preliminary screening module is specifically as follows: S11, key physiological and behavioral indicators of the dairy cow are continuously monitored and collected through a wearable device provided with a sensor; the sensor comprises an accelerometer, a gyroscope and a thermometer; the key physiological and behavioral indicators comprise activity amount, resting state, walking frequency and body temperature; 3. The mastitis multi-modal end-cloud joint defense early warning system for dairy cows according to claim 2, characterized in that, S12, a low-energy-consumption behavior detection method is designed based on the collected motion signals to ensure the robustness of the model under low sampling frequency. The design of the low-energy-consumption behavior detection method based on the collected motion signals specifically comprises the following steps: (1); wherein, the cross-entropy loss function of the teacher classification network, represents the number of samples in the dataset, denotes the pair of a sample and its corresponding label in the dataset, is the th element in the one-hot label vector is the posterior probability of each class of activity;​​​​ (2); wherein, a reconstruction loss function for reconstructing the network for the teacher, inputting the high sampling frequency sample into the teacher feature extractor to obtain a feature vector, the generated reconstructed image; S122, input the low sampling frequency data into the student classification network, and extract a feature vector and generate a reconstructed image through the reconstruction network The specific formula is as follows: (3); wherein, denotes the reconstruction-based information recovery, RIR, loss; S123, the first layer feature maps, respectively, along the time dimension and the sensor axis dimension 、 and 、 ; by normalizing the cosine similarity constraint, the correlation of the student feature and the teacher feature, the total loss is the average of each layer loss, the specific formula is as follows: (4); (5); (6); wherein, and respectively represent the temporal correlation distillation loss and the inter-axis correlation distillation loss, represents the information recovery loss of the correlation distillation of the layer, represents the information recovery CDIR loss based on the correlation distillation, represents the time dimension, respectively represent , the normalized horizontal vector, represents the total number of layers of the network; S124, combined classification loss , RIR loss and CDIR loss , and through the joint loss function training of the student classification network: (7); S121, a teacher classification network and a reconstruction network are trained using high sampling frequency data; the teacher classification network learns behavior classification ability through cross-entropy loss, and the reconstruction network reconstructs the input high sampling frequency image through mean square error loss, and the specific formula is as follows:

4. The mastitis multi-modal end-cloud joint defense early warning system for dairy cows according to claim 3, characterized in that, S125, dairy cow behavior data are subjected to statistical index calculation to construct a traditional machine learning model for realizing preliminary determination of mastitis risks of the dairy cow. S21, a cow body input picture is given where H, W, and C represent the height, width, and channel number of the cow body image, respectively. The ViT model divides the cow body input image into non-overlapping image blocks using a sliding window mechanism with a step size of S and a block size of P. The cow body input image with a resolution of HxW is divided into fixed-size image blocks, and the specific formula is as follows: (8); wherein represents a floor operation; S22, the segmented image block embedded into the input sequence as a representation of local features of the cow body, in addition to an extra sequence also embedded into the input sequence for learning a representation of global features of the cow body; S23, send the input sequence into After the layer Transformer encoder, get the output feature of the first layer , and separate the local feature of the body of the cow from the global feature according to the channel, get the local feature and ; In the high-risk cow identification and matching module, the cow body identification method specifically comprises the following process: (9); wherein, represents an output of the convolution space local feature aggregation module; represents twice convolution plus normalization and Mish activation function operation; represents concatenation of input features by channel; S25, using triple loss and label smoothed cross entropy loss Optimizing network: (10); wherein, is a joint loss function for training the cow body recognition network, is a hyperparameter, representing the weight relationship between the label smoothing cross-entropy loss and the triplet loss term; (11); wherein, represents a random noise coefficient; represents the number of bovine species; represents the true label, positive class is 1, negative class is 0; represents the predicted probability output by the model; (12); wherein, an input representing a current cow picture; an input representing a difficult sample image of the same batch of the same category; an input representing a difficult sample image of the same batch of the same category; an input representing a difficult sample image of the same batch of the different category; an input representing a difficult sample image of the same batch of the different category; , ), respectively represent a feature vector output by the model corresponding to the picture; represent the square of the 2-norm; represent the minimum interval of the distance between different cow images and the distance between the same cow images.

5. The mastitis multi-modal end-cloud joint defense early warning system for dairy cows according to claim 4, characterized in that, S24, cow body local features are subjected to sequential processing through two convolution blocks: In the multi-modal data fusion and diagnosis module, multi-modal interactive fusion is carried out based on a multi-modal interactive neural network CMI-Net model, and the specific steps are as follows: (13); wherein, and respectively represent layer and characteristic map of the layer, and respectively represent 1x1 and 1x3 convolution operations, represents a feature addition operation, and RELU(•) represents a rectified linear unit (ReLU) activation function. S31, for the input multi-modal data comprising motion signals, body temperature data and RGB video, a plurality of independent processing branches are first adopted to respectively extract features of the single modal data; each branch is based on convolution operation to realize feature extraction function of the corresponding modal; a Res-LCB module is designed and defined as follows:

6. The mastitis multi-modal end-cloud joint defense early warning system for dairy cows according to claim 5, characterized in that, S32, a joint cross-modal interactive module CMIM based on an attention mechanism is designed to adaptively recalibrate time and axial features of each modal by using multi-modal information. S321, set and respectively represent the features of two convolutional layer network specific layers; , and respectively represent the channel number of the feature map, the spatial height dimension of the feature map and the spatial width dimension of the feature map; A and G are input features of the joint cross-modal interaction module; S322, average pooling operation is performed along the channel of the input feature to generate two spatial maps; then the two spatial maps are connected and mapped into a joint table , and the operation is as follows: (14); wherein, represents a joint table evidence the number of channels, (•) represents an average pooling operation, [•] represents a concatenation operation; S323, set two spatial attention maps and is generated by putting joint representation into two independent convolutional layers, and the result is calculated with sigmoid function (•), the specific formula is as follows: (15); S324, using and re-calibrating the input features, generating two final refined features and : (16); wherein represents a feature multiplication operation.

7. The mastitis multi-modal end-cloud joint defense early warning system for dairy cows according to claim 6, characterized in that, The joint cross-modal interactive module CMIM based on the attention mechanism specifically comprises the following process: The cloud-based visual monitoring and early warning platform is specifically as follows: Managers can view the health status, behavior records, body temperature curves and early warning notifications of each dairy cow in the farm in real time through a mobile phone APP or a PC client; The platform provides detailed individual health reports, and the content covers daily activities, abnormal behavior records and historical early warning information of the dairy cow; Through the historical data backtracking function of the platform, managers can analyze long-term health change trends of the dairy cow by viewing a single early warning event.