Systems and methods for detecting breeding receptivity in livestock and other four-legged mammals

The AI-based estrus detection system addresses manual detection inefficiencies by using AI models to analyze animal images, enhancing breeding program success and reducing costs through automated estrus detection.

WO2025199652A1PCT designated stage Publication Date: 2025-10-02ONE CUP PROD LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2025/050447
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-03-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Conventional methods for detecting estrus in livestock rely heavily on manual observation, which is resource-intensive and prone to errors, leading to inefficiencies in breeding programs and increased costs.

Method used

A computer-vision based Artificial Intelligence (AI) system using multiple AI models to analyze animal images for estrus detection, including bounding box detection and classification models to identify mounting or chin-resting states, providing automated notifications for estrus detection.

Benefits of technology

Enables efficient, automated estrus detection with reduced human resource requirements, improving breeding success rates and herd management by optimizing mating timing and reducing unnecessary expenses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025050447_02102025_PF_FP_ABST
    Figure CA2025050447_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems are described for management of animals exhibiting estrus related behaviour by processing one or more images captured by one or more imaging devices using one or more AI models for detecting and locating two animals with a pair of overlapping bounding boxes in a given image of the one or more images; determining if a current state of a first animal of the two detected animals matches a mounting state or a chin-resting state and when the first animal is in the mounting state or the chin-resting state identifying a second animal of the two detected animals as being mounted by or receiving the chin of the first animal which is indicative that the second animal is in an estrus state; and generating a notification when the second animal is in the estrus state or the first animal is in the mounting or chin-resting state.
Need to check novelty before this filing date? Find Prior Art

Description

TITLE: SYSTEMS AND METHODS FOR DETECTING BREEDING RECEPTIVITY IN LIVESTOCK AND OTHER FOUR-LEGGED MAMMALSCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 572,171 , filed on March 29, 2024. The entire content of U.S. Provisional Patent Application No. 63 / 572,171 is hereby incorporated by reference.FIELD

[0002] Various embodiments are described herein that generally relate to systems and methods for detecting breeding receptivity (estrus) in livestock and other animals such as four-legged mammalian species, and in particular to systems and methods for breeding receptivity detection in livestock and other animals such as four-legged mammalian species using computer vision and Artificial Intelligence (Al) technologies.BACKGROUND

[0003] Animal management has important economic and ecological values. For example, the worldwide population of cattle ranges from 1 billion to 1.5 billion. Encompassing over 250 breeds, the cattle industry is a massive component of the worldwide economy. Approximately 1.3 billion people depend on cattle for their livelihood.

[0004] Canada is home to 12 million cattle scattered across 65,000 ranches and farms. The average ranch size is about 200 animals. Up to five institutions may handle an animal in its lifetime - ranches, community pastures, auctions, feedlots, and processing plants. Cattle contribute $18 billion CDN to the Canadian economy.

[0005] Cattle are found in over 170 countries. Among these countries, Canada represents roughly 1 % of the worldwide market. The USA and Europe have a cattle market that is at 20 to 25 times the size of Canada. China, Brazil, and India also have a significant share of the cattle global market. In some countries, New Zealand and Uruguay, for example, cattle outnumber people.

[0006] Thus, there is a great need for efficient systems and methods for tracking and managing livestock. Moreover, scientists, ecological researchers and activists, and governments also require efficient wildlife management systems and methods for tracking wildlife such as bear, elk, wolves, endangered animal species, and the like.

[0007] However, cattle ranchers and operators are still largely relying on the traditional, manual methods such as paperwork for livestock tracking. Such manual methods cannot easily scale up (and in some cases may be impossible to scale up) and may pose significant burden to ranchers and operators for tracking and managing livestock of large numbers.

[0008] Animal management can involve animal assessments such as determining if an animal is receptive to be bred (i.e., in an estrus state). Detecting whether an animal is receptive to be bred (estrus detection) offers significant benefits in animal husbandry and agriculture. It allows for precise timing of mating or artificial insemination, optimizing breeding programs for improved genetic traits, and herd management. Efficient estrus detection can increase reproductive efficiency, resulting in higher conception rates, and in reduced costs that may be incurred as a result of unsuccessful breeding attempts. Identifying estrus can also aid in disease management, since changes in behavior and physiology can signal health issues in an animal.

[0009] Conventional methods of identifying animals in an estrus state typically involve direct human observation of the behavior of animals and typically require many human resources, particularly when monitoring large numbers of animals. Improper or imprecise estrus-related assessments (e.g., failing to detect estrus, incorrectly detecting estrus) can result in time windows for mating or artificial insemination being missed, resources being unnecessarily expended during non-fertile time windows, and breeding programs with low success rates, leading to lost revenue and / or additional expenses.

[0010] Therefore, there is a need for systems and methods of performing efficient animal assessments including determining estrus-related states of animals.SUMMARY OF VARIOUS EMBODIMENTS

[0011] In a broad aspect, in accordance with the teachings herein, there is provided at least one embodiment of a method for animal management comprising: receiving, at a computing device, one or more images captured by one or more imaging devices; processing the one or more images using one or more Al models for: detecting and locating two animals with a pair of overlapping bounding boxes in a given image of theone or more images; determining if a current state of a first animal of the two detected animals in the given image matches a mounting state and when the first animal is in the mounting state identifying a second animal of the two detected animals as being mounted by the first animal which is indicative that the second animal is in an estrus state; and generating a notification when the second animal is in the estrus state or the first animal is in the mounting state.

[0012] In at least one embodiment, determining if the current state of the first detected animal matches the mounting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model, the age classification model being trained to generate an age classification output, the age classification output being one of: adult or non-adult; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position, using an activity classification model, the activity classification model being trained to generate an activity classification output, the activity classification output being one of: standing or non-standing; and in response to the pair of overlapping bounding boxes corresponding to two adult animals being in a standing position, applying a mounting state detection model to an image portion encapsulated by the bounding box for the first animal, to generate a mounting classification output for the detected animal associated with the image portion, the mounting classification output indicating whether the first animal in the image portion is mounting the second animal which is indicative that the second animal is in the estrus state and the first animal will be in the estrus state in the future.

[0013] In at least one embodiment, determining if the current state of the first detected animal matches the mounting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model, the age classification model being trained to generate an age classification output, the age classification output being one of: adult or non-adult; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position, using an activity classification model, the activity classification model being trained to generate an activity classification output, the activity classification output being one of: standing or non-standing; in response to the pair of overlapping bounding boxes corresponding to two adult animals being in a standingposition, applying a mounting state detection model to an image portion encapsulated by the bounding boxes of the first detected animal, to generate a mounting classification output for the first detected animal associated with the image portion, the mounting classification output indicating whether the first animal in the image portion is mounting the second animal; and if the second animal being mounted stays stationary for a predefined period of time after the mounting has finished, identifying the second animal being mounted as being in a standing heat state and the estrus state.

[0014] In at least one embodiment, determining if the first detected animal matches the mounting state further comprises: in response to the pair of overlapping bounding boxes corresponding to two adult animals, each adult animal being in a standing position, identifying the bounding box from the pair of bounding boxes corresponding to a higher bounding box relative to a ground in the given image; and applying the mounting state detection model to the image portion corresponding to the higher bounding box to generate the mounting classification output.

[0015] In at least one embodiment, determining if the current state of the first detected animal matches the mounting state further comprises: in response to the pair of overlapping bounding boxes corresponding to two adult animals, determining a corresponding size of each of the bounding boxes of the pair of overlapping bounding boxes, and in response to the size of each the bounding boxes satisfying a predetermined minimum size, applying the mounting state detection model.

[0016] In at least one embodiment, the mounting state detection model is trained to: generate a mounting classification output distinguishing the mounting state from other non-mounting states; and generate a confidence value for the mounting classification output.

[0017] In at least one embodiment, the method further comprises determining if the confidence value for the mounting classification output satisfies a predetermined threshold; and in response to determining the confidence value for the mounting classification output satisfies the predetermined threshold, generating the notification.

[0018] In at least one embodiment, the one or more Al models comprise an animal identification model for identifying the one or more detected animals and associatinga unique identifier with each of the one or more detected animals in the given image and wherein providing the notification comprises including the unique identifier of the animals in the mounting state, the estrus state and / or the standing heat state.

[0019] In at least one embodiment, locating a given detected animal in the given image comprises generating a mask identifying each pixel of the given image associated with the given detected animal.

[0020] In at least one embodiment, the notification comprises information related to the current state of each of the detected animals, and one or more image or videos of each of the detected animals.

[0021] In at least one embodiment, the method further comprises annotating the one or more images to indicate the animal matching the estrus state, the mounting state and / or the standing heat state; and displaying, on a display, the annotated one or more images.

[0022] In at least one embodiment, the notification comprises an annotated image including the animals matching the estrus state, the mounting state and / or the standing heat state.

[0023] In a broad aspect, in accordance with the teachings herein, there is provided at least one embodiment of a method for animal management comprising: receiving, at a computing device, one or more images captured by one or more imaging devices; processing the one or more images using one or more Al models for: detecting and locating two animals with a pair of overlapping bounding boxes in a given image of the one or more images; determining if a current state of a first animal of the two detected animals in the given image matches a chin-resting state and when the first animal is in the chin-resting state identifying a second animal of the two detected animals as receiving a chin of an animal nears its rear end which is indicative that the second animal is in an estrus state; and generating a notification when the second animal is in the estrus state or the first animal is in the chin-resting state.

[0024] In at least one embodiment, determining if the current state of the first detected animal matches the chin-resting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an ageclassification model the age classification model being trained to generate an age classification output, the age classification output being one of: adult or non-adult; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position using an activity classification model, the activity classification model being trained to generate an activity classification output, the activity classification output being one of: standing or non-standing; in response to the pair of overlapping bounding boxes corresponding to two adult animals in a standing position, generating a merged bounding box, the merged bounding box encapsulating an image portion that includes the two standing adult animals; and generating a chin-resting classification output by using a chin-resting state classification model to analyze the merged bounding box, the chin-resting classification output indicating whether the first detected animal in the image portion is in a chin-resting state which is indicative that the second animal is in the estrus state and the first animal will be in the estrus state in the future.

[0025] In at least one embodiment, determining if the current state of the first detected animal matches the chin-resting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model, the age classification model being trained to generate an age classification output; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position using an activity classification model, the activity classification model being trained to generate an activity classification output, the activity classification output being one of: standing or nonstanding; in response to the pair of overlapping bounding boxes corresponding to two adult animals in a standing position, generating a merged bounding box, the merged bounding box encapsulating an image portion that includes the two standing adult animals; and generating a chin-resting classification output by using a chin-resting state classification model to analyze the merged bounding box, the chin-resting classification output indicating whether the first detected animal in the image portion is in a chin-resting state which is indicative of the second animal being in the estrus state.

[0026] In at least one embodiment, determining if the current state of the first detected animal matches the chin-resting state further comprises: in response to thepair of overlapping bounding boxes corresponding to two adult animals, determining a corresponding size of each of the bounding boxes of the pair of overlapping bounding boxes, and in response to the size of each the bounding boxes satisfying a predetermined minimum size, applying the chin-resting state classification model.

[0027] In at least one embodiment, the chin-resting state classification model is trained to: generate a chin-resting classification output distinguishing the chin-resting state from other non-chin-resting states; and generate a confidence value for the chinresting classification output.

[0028] In at least one embodiment, the method further comprises determining if the confidence value for the chin-resting classification output satisfies a predetermined threshold; and in response to determining the confidence value for the chin-resting classification output satisfies the predetermined threshold, generating the notification.

[0029] In at least one embodiment, locating a given detected animal in the given image comprises generating a mask identifying each pixel of the given image associated with the given detected animal.

[0030] In at least one embodiment, the notification comprises information related to the current state of each of the at least one detected animals, and one or more image or videos of each of the detected animal.

[0031] In at least one embodiment, the method further comprises annotating the one or more images to indicate each of the at least one detected animal matching the estrus state, and / or the chin-resting state; and displaying, on a display, the annotated one or more images.

[0032] In at least one embodiment, the notification comprises an annotated image including the at least one detected animal matching the estrus state, and / or the chinresting state.

[0033] In a broad aspect, in accordance with the teachings herein, there is provided at least one embodiment of a method for animal management comprising receiving, at a computing device, a plurality of images captured by one or more imaging devices, the plurality of images corresponding to images sequentially obtained during a predefined time window; processing the plurality of images using one or more Almodels for: detecting and locating two animals with overlapping bounding boxes in a given image of the one or more images; and determining if a current state of a first animal of the two detected animals in the given image matches a mounting state; when the first animal is in the mounting state, determining if a second animal of the two detected animals is in a standing heat state, wherein the second animal is in the standing heat state if the second animal remains stationary for at least several images sequentially obtained over a predefined period of time after the second animal is no longer being mounted; and providing a notification in response to when the second animal is determined to be in the standing heat state.

[0034] In at least one embodiment, determining if the current state of the first detected animal in the given image matches a mounting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model, the age classification model being trained to generate an age classification output, the age classification output being one of: adult or non-adult; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position, using an activity classification model the activity classification model being trained to generate an activity classification; in response to the pair of overlapping bounding boxes corresponding to two adult animals being in a standing position, applying a mounting classification image classification model to an image portion encapsulated by the bounding box of the first detected animal to generate a mounting classification output for the first detected animal, the mounting classification output indicating whether the first detected animal in the image portion is mounting the second detected animal.

[0035] In at least one embodiment, one of the bounding boxes is a rectangular bounding box, a square box or an oriented bounding box.

[0036] In at least one embodiment, the method comprises providing a candidate estrus image to a multi-modal deep learning Al model that is trained to analyze the candidate estrus image to generate a user notification including the candidate estrus image, image description text and a detected event indication where the image description text includes text describing what is shown in the candidate estrus image and the detected event indication includes text that describes a detected estrus event.

[0037] In at least one embodiment, the user notification further includes location and time information for indicating where and when the candidate estrus image was obtained.

[0038] In at least one embodiment, determining the candidate estrus image is based on detecting estrus according to any of the methods described herein.

[0039] In at least one embodiment, the multi-modal deep learning Al model is trained by providing to the multi-modal deep learning Al model: (a) an introduction description that is used to train the multi-modal deep learning Al model to perform certain functions, (b) an estrus detection signs description that is used to train the multi-modal deep learning Al model to learn estrus detection signs to look for in the candidate calving image and (c) a reporting estrus signs description that is used to train the multi-modal deep learning Al model to learn estrus detection reporting signs to use when generating the image description text and the detected event indication in the user notification.

[0040] In at least one embodiment, after the at least one animal is detected, a sex of the detected at least one animal is obtained using a sex classifier, and further processing of the given image is not performed when the at least one detected animal is male.

[0041] In at least one embodiment, the notification indicates that the detected animal in a mounting state or in a chin-resting state will be experiencing estrus within the near future.

[0042] In a broad aspect, in accordance with the teachings herein, there is provided at least one embodiment of a computing device comprising a memory storing program instructions and a processor that is coupled to the memory to read and execute the program instructions which configure the processor to perform a method for animal management including detecting at least one animal that is experiencing estrus, wherein the method is defined according to any of the methods described herein.

[0043] In a broad aspect, in accordance with the teachings herein, there is provided at least one embodiment of a non-transitory computer readable medium storing thereon program instructions that, when executed by a processor of a computingdevice configure the processing for performing a method for animal management including detecting at least one animal that is experiencing estrus, wherein the method is defined according to any of the methods described herein.

[0044] Other features and advantages of the present application will become apparent from the following detailed description taken together with the accompanying drawings. It should be understood, however, that the detailed description and the specific examples, while indicating preferred embodiments of the application, are given by way of illustration only, since various changes and modifications within the spirit and scope of the application will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE DRAWINGS

[0045] For a better understanding of the various embodiments described herein, and to show more clearly how these various embodiments may be carried into effect, reference will be made, by way of example, to the accompanying drawings which show at least one example embodiment, and which are now described. The drawings are not intended to limit the scope of the teachings described herein.

[0046] FIG. 1A is a schematic diagram showing an application of an animal assessment system, according to at least one example embodiment of this disclosure.

[0047] FIG. 1 B is a schematic diagram showing another example application of the animal assessment system of FIG. 1A.

[0048] FIG. 2 is a schematic diagram illustrating an example embodiment of the hardware structure of an imaging device that may be used with one of the embodiments of the animal assessment systems described herein.

[0049] FIG. 3A is a schematic diagram illustrating an example embodiment of the hardware structure of the client computing devices and server computer one of the animal assessment systems described herein.

[0050] FIG. 3B is schematic diagram illustrating an example embodiment of a simplified software architecture of the client computing devices and server of the at least one of the embodiments of the animal assessment system described herein.

[0051] FIG. 4A is a schematic diagram of a deep neural network (DNN) that may be used by the at least one of the embodiments of the animal assessment system described herein.

[0052] FIG. 4B is a schematic diagram of a Jigsaw model that may be used by the animal assessment system of FIG. 1A.

[0053] FIG. 5 is a flowchart showing an example embodiment of an animal assessment method that may be executed by a server program of at least one of the embodiments of the animal assessment system described herein, for detecting a mounting state in at least one animal at the site using an Artificial Intelligence (Al) pipeline.

[0054] FIG. 6 is a flowchart showing another example embodiment of an animal assessment method that may be executed by a server program of at least one of the embodiments of the animal assessment system described herein, for detecting a mounting state in at least one animal at the site using an Artificial Intelligence (Al) pipeline.

[0055] FIGS. 7A-7C are example synthetic images that may be generated for training a mounting state detection model of at least one of the embodiments of the animal assessment system described herein.

[0056] FIG. 8 is an example of an image that is processed according to the teachings herein wherein three animals are detected, and each is associated with a respective bounding box defined thereabout.

[0057] FIG. 9A shows an example image that may be processed according to the teachings herein.

[0058] FIG. 9B shows an output of an age classifier applied to the image of FIG. 9A.

[0059] FIG. 10A shows an example image that may be processed according to the teachings herein.

[0060] FIG. 10B shows an output of an activity classifier applied to the image of FIG. 10A.

[0061] FIG. 11 A shows an example image that may be processed according to the teachings herein.

[0062] FIGS. 11 B-11C show the image of FIG. 11A processed for detection of a mounting state.

[0063] FIGS. 12A-12H show images processed by the animal assessment system, showing a mounting state and associated heatmaps showing the accuracy of a mounting state detection model of the animal assessment system.

[0064] FIG. 13 is a flowchart showing an example embodiment of an animal assessment method that may be executed by a server program of at least one of the embodiments of the animal assessment system described herein, for detecting a standing heat in at least one animal at the site using an Artificial Intelligence (Al) pipeline.

[0065] FIGS. 14A-14F show images processed by the animal assessment method of FIG. 13 for standing heat detection.

[0066] FIG. 15 is a flowchart showing an example embodiment of an animal assessment method that may be executed by a server program of at least one of the embodiments of the animal assessment system described herein, for detecting chinresting in at least one detected animal at the site using an Artificial Intelligence (Al) pipeline.

[0067] FIG. 16 is a flowchart showing another example embodiment of an animal assessment method that may be executed by a server program of at least one of the embodiments of the animal assessment system described herein, for detecting chinresting in at least one detected animal at the site using an Artificial Intelligence (Al) pipeline.

[0068] FIGS. 17A-17B are example synthetic images that may be generated for training a chin-resting detection model of at least one of the embodiments of the animal assessment system described herein.

[0069] FIG. 18A shows an example image that may be processed according to the teachings herein.

[0070] FIGS. 18B-18D show the image of FIG. 18A processed for detection of a chin-resting state.

[0071] FIGS. 19A-19H show example images processed by one of the embodiments of the animal assessment system, showing a chin-resting state and associated heatmaps showing the accuracy of a chin-resting state detection model of the animal assessment system.

[0072] FIGS. 20A-20B show example images processed according to the teachings herein with bounding boxes added thereto in which a mounting state (FIG. 20A) and a chin-resting state (FIG. 20B) are detected.

[0073] FIG. 21 is a flowchart showing another example embodiment of an animal assessment method that may be executed by a server program of at least one of the embodiments of the animal assessment system described herein, for detecting a state indicative of breeding receptivity (e.g., estrus) in at least one detected animal at the site using an Artificial Intelligence (Al) pipeline.

[0074] FIGS. 22A-22B show example images of user notifications provided by at least one embodiment of an animal assessment system described herein.

[0075] FIGS. 23A and 23B show example embodiments of a Graphical User Interface (GUI) that may be used by at least one embodiment of an animal assessment system described herein to provide animal assessment data for a site installation.

[0076] FIGS. 24A and 24B show example images of an example user notification provided by at least one embodiment of an animal assessment system described herein.

[0077] Further aspects and features of the example embodiments described herein will appear from the following description taken together with the accompanying drawings.DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] Various embodiments in accordance with the teachings herein will be described below to provide an example of at least one embodiment of the claimed subject matter. No embodiment described herein limits any claimed subject matter. The claimed subject matter is not limited to devices, systems or methods having all ofthe features of any one of the devices, systems or methods described below or to features common to multiple or all of the devices, systems or methods described herein. It is possible that there may be a device, system or method described herein that is not an embodiment of any claimed subject matter. Any subject matter that is described herein that is not claimed in this document may be the subject matter of another protective instrument, for example, a continuing patent application, and the applicants, inventors or owners do not intend to abandon, disclaim or dedicate to the public any such subject matter by its disclosure in this document.

[0079] It will be appreciated that for simplicity and clarity of illustration, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements. In addition, numerous specific details are set forth in order to provide a thorough understanding of the embodiments described herein. However, it will be understood by those of ordinary skill in the art that the embodiments described herein may be practiced without these specific details. In other instances, well-known methods, procedures and components have not been described in detail so as not to obscure the embodiments described herein. Also, the description is not to be considered as limiting the scope of the embodiments described herein.

[0080] It should also be noted that the terms “coupled” or “coupling” as used herein can have several different meanings depending in the context in which these terms are used. For example, the terms coupled or coupling can have a mechanical or electrical connotation. For example, as used herein, the terms coupled or coupling can indicate that two elements or devices can be directly connected to one another or connected to one another through one or more intermediate elements or devices via an electrical signal, electrical connection, or a mechanical element, depending on the particular context.

[0081] Unless the context requires otherwise, throughout the specification and claims which follow, the word “comprise” and variations thereof, such as, “comprises” and “comprising” are to be construed in an open, inclusive sense, that is, as “including, but not limited to”.

[0082] Various terms used throughout the present description may be read and understood as follows, unless the context indicates otherwise: singular articles and pronouns as used throughout include their plural forms, and vice versa; similarly, gendered pronouns include their counterpart pronouns so that pronouns should not be understood as limiting anything described herein to use, implementation, performance, etc. by a single gender. Further definitions for terms may be set out herein; these may apply to prior and subsequent instances of those terms, as will be understood from a reading of the present description.

[0083] It should also be noted that, as used herein, the wording “and / or” is intended to represent an inclusive-or. That is, “X and / or Y” is intended to mean X or Y or both, for example. As a further example, “X, Y, and / or Z” is intended to mean any operative combination of X, Y or Z such as X, Y, Z, X and Y, X and Z, Y and Z, as well as X, Y and Z.

[0084] It should be noted that terms of degree such as "substantially", "about" and "approximately" as used herein mean a reasonable amount of deviation of the modified term such that the end result is not significantly changed. These terms of degree may also be construed as including a deviation of the modified term, such as by 1 %, 2%, 5% or 10%, for example, if this deviation does not negate the meaning of the term it modifies.

[0085] Furthermore, the recitation of numerical ranges by endpoints herein includes all numbers and fractions subsumed within that range (e.g., 1 to 5 includes 1 , 1.5, 2, 2.75, 3, 3.90, 4, and 5). It is also to be understood that all numbers and fractions thereof are presumed to be modified by the term "about" which means a variation of up to a certain amount of the number to which reference is being made if the end result is not significantly changed, such as 1 %, 2%, 5%, or 10%, for example.

[0086] At least a portion of the example embodiments of the systems or methods described in accordance with the teachings herein may be implemented as a combination of hardware or software. For example, a portion of the embodiments described herein may be implemented, at least in part, by using one or more computer programs, executing on one or more programmable devices comprising at least one processing element, and at least one data storage element (including volatile and non-volatile memory). These devices may also have at least one input device (e.g., a touchscreen, and the like) and at least one output device (e.g., a display screen, a printer, a wireless radio, and the like) depending on the nature of the device.

[0087] It should also be noted that some elements that are used to implement at least part of the embodiments described herein may be implemented via software that is written in a high-level procedural language such as object-oriented programming. The program code may be written in JAVA, PYTHON, C, C++, Javascript or any other suitable programming language and may comprise modules or classes, as is known to those skilled in object-oriented programming. Alternatively, or in addition thereto, some of these elements implemented via software may be written in assembly language, machine language, or firmware as needed.

[0088] At least some of the software programs used to implement at least one of the embodiments described herein may be stored on a storage medium (e.g., a computer readable medium such as, but not limited to, ROM, flash memory, magnetic disk, optical disc) ora device that is readable by a programmable device. The software program code, when read by the programmable device, configures the programmable device to operate in a new, specific and predefined manner in order to perform at least one of the methods described herein.

[0089] Furthermore, at least some of the programs associated with the systems and methods of the embodiments described herein may be capable of being distributed in a computer program product comprising a computer readable medium that bears computer usable instructions, such as program code, for one or more processors. The program code may be preinstalled and embedded during manufacture and / or may be later installed as an update for an already deployed computing system. The medium may be provided in various forms, including non- transitory forms such as, but not limited to, one or more diskettes, compact disks, DVD, tapes, chips, and magnetic, optical and electronic storage. In alternative embodiments, the medium may be transitory in nature such as, but not limited to, wire- line transmissions, satellite transmissions, internet transmissions (e.g., downloads), media, digital and analog signals, and the like. The computer useable instructions may also be in various formats, including compiled and non-compiled code.

[0090] Accordingly, any module, unit, component, server, computer, terminal or device described herein that executes software instructions may include or otherwise have access to computer readable media such as storage media, computer storage media, or data storage devices (removable and / or non-removable) such as, for example, magnetic disks, optical disks, or tape. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer readable instructions, data structures, program modules, or other data. Examples of computer storage media include RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information, and which can be accessed by an application, module, or both. Any such computer storage media may be part of the device or accessible or connectable thereto.

[0091] Detecting breeding receptivity (e.g., detecting estrus) in female animals is crucial for successful breeding and maintaining herd productivity. Accurate estrus detection allows farmers to time mating or artificial insemination when the animals are most fertile, increasing the likelihood of conception. This optimization reduces the time and resources required for each pregnancy, improving overall herd efficiency and profitability. Identifying estrus also aids in identifying potential reproductive problems, such as anestrous (lack of estrus), allowing for early intervention and veterinary care. Efficient estrus detection is essential for sustainable animal farming, ensuring consistent animal production and maintaining a healthy and economically viable herd.

[0092] Estrus indicators in female animals encompass a range of behavioral, physical, and physiological changes that signal the animal’s readiness for mating. By way of example, such indicators for mating readiness in animals can include any combination of restlessness, increased vocalization, tail flagging, reduced feed intake, change in behavior toward bulls, mounting, standing to be mounted (standing heat), and chin-resting. Being able to remotely and autonomously recognize animals that are in estrus can help with improving breeding success rates and overall animal management. For various animals such as, but not limited to, cattle, sheep, goats,pigs, horses, being able to remotely detect mounting behaviour will aid in identifying these animals that are experiencing estrus.

[0093] Traditional methods of detecting estrus rely on human monitoring of the behavior of animals and requires substantial human resources which can be costly and time consuming. New methods of detecting estrus include wearable tracking devices; however, the devices are expensive, require a lot of maintenance and cleaning, and may affect the animal’s behavior which can lead to inaccurate estrus detection. Therefore, there is a need for systems and methods of performing efficient estrus detection is animals.

[0094] Embodiments disclosed herein generally relate to a computer-vision based Artificial Intelligence (Al) system and method for tracking and managing animals. The Al system uses a plurality of Al models arranged according to an Al pipeline structure so that the embodiments disclosed herein may uniquely and consistently perform animal assessments at a site such as a ranch or feedlot. The disclosed embodiments allow at least some activities related to tracking and managing animals to be performed remotely (e.g., detection of behaviors, activities, and / or states), reducing human resources needed for tracking and managing animals.

[0095] At least one embodiment of the systems and methods disclosed herein may simultaneously monitor one or more animals in a herd for tracking and managing various aspects of the tracked animals such as, but not limited to, their health, activity, estrus behavior and / or nutrition, for example.

[0096] At least one embodiment of the systems and methods disclosed herein may continuously operate and / or proactively notify personnel at the site, such as ranch operators, for example, of any tracked animals that are exhibiting estrus indicators thereby allowing the ranch operator to artificially inseminate the animal at an optimized time.

[0097] Depending on the imaging devices or cameras used, embodiments disclosed herein may track and assess animals from various distances. For example, in at least one embodiment of the system disclosed herein, an animal may be assessed from up to 10 meters (m) away from an imaging device (for example, about 1 m, 2m, 3m, 4m, 5m, 6m, 7m, 8m, 9m, or 10m away) when the animal is still, andfurther tracked up to 20m, 30m, 40m, or 50m away from the imaging device with a visual lock. In at least one embodiment, where high-performance cameras such as optical-zoom cameras are used, for example, the system disclosed herein may identify animals from further distances. In one example, the system may include 25x optical- zoom cameras that are used to assess animals with sufficient confidence levels from more than 60m, more than 70m, more than 80m, more than 90m, or more than 100m away.

[0098] As will be described in more detail below, at least one embodiment of the systems and methods disclosed herein may assess animals in an image wherein the image may be obtained by an image acquisition device that is positioned from one or more angles relative to the animals. The images can be still images captured from the image acquisition device or can be image frames extracted from a video captured from the image acquisition device.

[0099] At least one of the embodiments described herein may allow one or more operators at a site to communicate with a server of the system via their handheld devices, such as smartphones or tablets, for example, and view assessed animals, and optionally data about the assessed animals, based on images of the animals that may be captured from one or more angles relative to the animals.

[0100] For example, at least one embodiment of the systems disclosed herein may process images of animals taken from one of several angles relative to the animals and assess animals in the images by detecting the animals in an image and processing the image of a detected animal to generate a plurality of sections. The sections may then be used by one or more Al models to make a final prediction of the animal’s current state and also generate the level of confidence that the final prediction is accurate. In at least one embodiment, the one or more Al models are only applied to animals that satisfy specific trigger conditions for detecting a particular activity, such as estrus, for example. The selection of which animals to run the one or more Al models on allows various embodiments of the systems and methods disclosed herein to save on computing resources by not running the one or more Al models on every animal detected in an image and only generating a final prediction of a detectedanimal’s current state and optionally, a confidence level of the final prediction for selected animals.

[0101] In at least one embodiment of the systems and methods disclosed herein, one or more assessments of the current state of the animals can be performed. The assessments may be performed continuously, periodically, or preferably in response to specific assessment conditions being met. In such embodiments, the disclosed systems and methods may provide various benefits to users, such as ranch operators, for any combination of the following conditions.(1) Mounting detection: On a ranch, animals may not be seen by a ranch operator, for long periods of time. Accordingly, in one aspect, at least one embodiment in accordance with the teachings herein may be used to autonomously detect when an animal is exhibiting estrus indicators such as being mounted and to notify one or more users (e.g., a ranch operator) to take one or more necessary action(s) on animals in heat, as necessary.(2) Standing heat detection: In another aspect, at least one embodiment in accordance with the teachings herein may be used to autonomously detect when an animal is exhibiting one or more estrus indicators such as standing heat and to notify one or more users (e.g., a ranch operator) to take one or more necessary action(s) on animals in heat, as necessary.(3) Chin-resting detection: In another aspect, at least one of the in accordance with the teachings herein may be used to autonomously detect when an animal is exhibiting one or more estrus indicators such as receiving a chinresting on its rear end and to notify one or more users (e.g., a ranch operator) to take one or more necessary action(s) on animals in heat, as necessary.

[0102] In at least one embodiment of the systems disclosed herein, the system may train the one or more Al models that are used continuously or periodically during operation, or periodically when the system is not being operated for performing various processes on images / videos containing animals.

[0103] In another aspect, in at least one embodiment of the systems and methods disclosed herein, one or more animal images are processed to generate bounding boxes for one or more sections of two animals in the animal images and optionally mask some sections thereof as needed. In some cases, key points may be determined for portions of the animal in the one or more sections. The key points and bounding boxes may be determined using one or more Al models such as was described in United States Patent No. 11 ,910,784, the entire contents of which are incorporated by reference herein.

[0104] At least one embodiment of the systems and methods disclosed herein may be trained with images and / or videos to learn certain characteristics of animals over time. For example, the collected temporal data may be used for learning how to accurately determine when an adult animal is displaying mounting behavior. For example, at least one embodiment of the systems and methods disclosed herein may use one or more Al models that are trained in detecting estrus behavior by using positive and negative training images and / or videos of animals, where positive training images / videos show animals experiencing estrus behavior and negative training images / videos show animals not experiencing estrus behavior. In some cases, it may be challenging to assemble a sufficient number of training images / videos for uncommon estrus events (e.g., training images of chin-resting or mounting from various angles with a variety of different animals). In some embodiments, a portion of the training images / videos may be simulated or synthetic. The simulated or synthetic training images / videos may be generated, for example, as described in PCT Patent Application No. PCT / CA2024 / 050321 , the entire contents of which are incorporated by reference herein.System Overview

[0105] Turning now to FIG. 1A, an animal assessment system, according to at least one embodiment of this disclosure, is shown and is generally identified using reference numeral 100. The animal assessment system 100 is generally associated with a site 102 that includes animals. The site 102 may be an outdoor site, an indoor site, or a site with indoor and outdoor spaces. Furthermore, the site may be an enclosed site, such as a barn, a pen, or a fenced lot, and / or the like, or the site may be open, suchas a wildlife park or a natural animal habitat. Examples of the site 102 may be a livestock ranch, a ranch facility, an animal habitat, and / or the like. The site 102 generally comprises a plurality of animals 104 such as livestock (e.g., cows, horses, pigs, sheep, etc.) and / orwild animals (e.g., bison, wolves, elk, bears, goats, elephants, birds, and / or the like) and other four-legged mammals. In at least one embodiment, one or more animals 104 may be also associated with an identification (ID) tag 106, such as a visible ID tag attached therewith.

[0106] The system 100 comprises one or more imaging devices 108 that are deployed or deployable at the site 102 and are adapted to communicate with a server 120 via a communication network. For example, the imaging devices 108 may communicate with a computer cloud 110 directly or via suitable intermediate communication nodes 112 such as access points, switches, gateways, routers, and / or the like that are deployed at the site 102 or at a site remote to the site 102. In an alternative embodiment, the imaging devices 108 may communicate with a site server 120s that is located at site 102, or remote to the site 102. The site server 120s may then communicate with the server 120, which may be located at the site 102 or remote to the site 102. One or more first users 114A may operate suitable client computingdevices 116 to communicate with the computer cloud 110 for monitoring the animals 104 at the site 102.

[0107] Although embodiments have been described above with reference to the accompanying drawings, those of skill in the art will appreciate that variations and modifications may be made.

[0108] For example, while the animal assessments are described primarily in terms of cows, it should be understood that these techniques can be applied to other animals, as indicated above, by training the various Al models in a somewhat similar fashion using images and other data from the other species.

[0109] FIG. 1 B shows another example use of the system 100 shown in FIG. 1A. In this example, the site 102 is a fenced site with a barn accessible to the animals 104. One or more fixed imaging devices 108B may be installed at the site 102 such as on the fence. Alternatively, a mobile imaging device, i.e., a drone 108A may be used to monitor the site 102.

[0110] The one or more imaging devices 108 may be any suitable imaging devices such as one or more drones 108A equipped with suitable cameras, one or more surveillance cameras 108B mounted at predefined anchor locations, one or more video cameras 108C movable with and operable by one or more second users 114B such as farmers, researchers, and / or the like, one or more cellphones or smartphones 108D equipped with suitable cameras and movable with and operable by the one or more second users 114B, and / or the like. In at least one embodiment, the first and second users 114A and 114B may be different users while in at least one other embodiment, the first and second users 114A and 114B may be the same users. Imaging devices 108 may be moveable with reference to the installation platform (e.g., fence, drone etc.). For example, imaging devices 108 may be point, tilt, zoom (PTZ) cameras.

[0111] Herein, the cameras 108 may be any suitable cameras for capturing images under suitable lighting conditions, such as any combination of one or more visible-light cameras, one or more infrared (IR) cameras, and / or the like. Depending on the types of the imaging devices 108, the deployment locations thereof, and / or the deployment / operation scenarios (e.g., deployed af fixed locations or movable with the user 114A during operation), the imaging devices 108 may be powered by any suitable power sources such as power grid, batteries, solar panels, Power of Ethernet (POE), and / or the like.

[0112] Moreover, the imaging devices 108 may communicate with other devices via any suitable interfaces such as the Camera Serial Interface (CSI) defined by the Mobile Industry Processor Interface (MIPI) Alliance, Internet Protocol (IP), Universal Serial Bus (USB), Gigabit Multimedia Serial Link (GMSL), and / or the like. The imaging devices 108 generally have a resolution (e.g. , measured by number of pixels) sufficient for the image processing described later. The imaging devices 108 that are used may be selected such that they are able to acquire images that preferably have a high resolution such as, but not limited to, 720p (1280x720 pixels), 1080p (1920x1080 pixels), 2K (2560x1440 pixels) or 4K (3840x2160 pixels), for example. Some or all imaging devices 108 may have a zoom function. Some or all imaging devices 108 may also have an image stabilization function and motion activated function.

[0113] In at least one embodiment, at least one of the imaging devices 108 may be controlled by the system 100 such that the angle, amount of zoom and / or any illumination of the imaging device 108 may be controlled with a server (e.g., server 120 or server 120s) that is associated with system 100. Alternatively, any mobile imaging devices 108b may be controlled to go from one site to another site or to move around larger sites in a controlled manner where the flight path of drone 108A is predefined or controlled by a user. In these embodiments, a server (e.g., server 120 or server 120s) of the site that provides these control signals to the imaging devices 108 may do so under program control or under control commands provided by system personnel.

[0114] FIG. 2 is a schematic diagram illustrating the hardware structure of the imaging device 108. As shown, the imaging device 108 comprises an image sensor 124, an image signal processing module 126, a memory module 128, an image / video encoder 130, a communication module 132, and other electrical modules 134 (e.g., lighting modules, phone modules when the imaging device 108 is a smartphone, motion sensing modules, and / or the like). The imaging device 108 also comprises a lens module 122 coupled to image sensor 124 for directing light to the image sensor 124 with proper focusing. The imaging device 108 may further comprise other mechanical, optical, and / or electrical modules such as an optical zoom module, panning / tilt module, propeller and flight control modules (if the imaging device 108 is a drone), and / or the like. The various elements of the imaging device 108 are interconnected by a system bus 136 to allow transfer of data, and control signals therebetween. Also included, but not shown, is a power supply and a power bus for providing power to various elements of the imaging device 108. It should be understood that by referring to various elements herein with the word module this implies that these elements include hardware.

[0115] The image sensor 124 may be a complementary metal oxide semiconductor (CMOS) image sensor, a charge coupled device (CCD) image sensor, or any other suitable image sensor.

[0116] The image signal processing module 126 receives captured images from the image sensor 124 and processes received images as needed, for example,automatically adjusting image parameters such as exposure, white balance, contrast, brightness, hue, and / or the like, digital zooming, digital stabilization, and the like. The image signal processing module 126 includes one or more processors with suitable processing power for performing the required processing.

[0117] The memory module 128 may comprise one or more non-transitory machine-readable memory units or storage devices accessible by other modules such as the image signal processing module 126, the image / video encoder module 130, and the communication module 132 for reading and executing machine-executable instructions stored therein (e.g., as a firmware), and for reading and / or storing data such as data of captured images and data generated during the operation of other modules. The memory module 128 may be volatile and / or non-volatile, non-removable or removable memory such as RAM, ROM, EEPROM, solid-state memory, hard disks, CD, DVD, flash memory, and / or the like.

[0118] The image / video encoder module 130 encodes captured images into a format suitable for storage and / or transmission such as jpeg, jpeg2000, png, and / or the like. As used herein, the term “captured images” may include stand-alone images captured, for example, by image sensor 124 and / or individual video frames of a captured video. The image / video encoder module 130 may also encode a series of captured images (where these images may be referred to as a captured “video stream” or “video clip”) and associated audio signals into a format suitable for storage and / or transmission using a customized or standard audio / video coding / decoding (codec) technology such as H.264, H.265, HVEC, MPEG-2, MPEG-4, AVI, VP3 to VP9 (i.e., webm format), MP3 (also called MPEG-1 Audio Layer III or MPEG-2 Audio Layer III), FTP, Secure FTP, RTSP, and / or the like. In at least one embodiment, the image / video encoder module 130 may create a video stream that has a lower frame rate that is comprised of standalone images combined into a video, such as for example 1 , 1 / 5thor 1 / 10thof a conventional frame rate, which may be considered to be more akin to stop motion video. This improves efficiency for downstream operations such as transmitting and / or processing the lower frame rate video stream. Regardless of frame rate, there is a sequence of image frames which may still be referred to as a video stream or video clip.

[0119] The communication module 132 is adapted to communicate with other devices of the animal assessment system 100, e.g., for receiving instructions from the computer cloud 110 and transmitting encoded images (generated by the image / video encoder module 130) thereto, via suitable wired or wireless communication technologies such as Ethernet, WI-FI® (WI-FI is a registered trademark of Wi-Fi Alliance, Austin, TX, USA), BLUETOOTH® (BLUETOOTH is a registered trademark of Bluetooth Sig Inc., Kirkland, WA, USA), ZIGBEE® (ZIGBEE is a registered trademark of ZigBee Alliance Corp., San Ramon, CA, USA), wireless broadband communication technologies such as Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Universal Mobile Telecommunications System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX), CDMA2000, Long Term Evolution (LTE), 3GPP 5G networks, and / or the like.

[0120] In at least one embodiment, the imaging devices 108 may be Internet of Things (loT) devices and the communication module 132 may communicate with other devices using a suitable low-power wide-area network protocol such as Long Range (LoRa), Bluetooth Low Energy (BLE), Z-Wave, and / or the like.

[0121] In at least one embodiment, parallel ports, serial ports, USB connections, optical connections, or the like may also be used for connecting the imaging devices 108 with other computing devices or networks although they are usually considered as input / output interfaces for connecting input / output devices. In at least one embodiment, removable storage devices such as portable hard drives, diskettes, secure digital (SD) cards, minSD cards, microSD cards, and / or the like may also be used for transmitting encoded images or video clips to other devices.

[0122] The computer cloud 110 generally comprises a network 118 and one or more server computers 120 connected thereto via suitable wired and wireless networking connections. The network 118 may be any suitable network such as a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), the Internet, and / or the like.

[0123] The server computer 120 may be a server-computing device, and / or a computing device acting as a server computer while also being used by a user. The client computing devices 116 may be any suitable portable or non-portable computingdevices such as desktop computers, laptop computers, tablets, smartphones, Personal Digital Assistants (PDAs), and / or the like. In each of these cases, the computing devices have one or more processors that are programmed to operate in a new specific manner when executing programs for executing various aspects of the processes described in accordance with the teachings herein.

[0124] Generally, the client computing devices 116 and the server computers 120, 120s have a similar hardware structure such as a hardware structure 140 shown in FIG. 3A. As shown, the computing device 116 / 120 comprises a processor module 142, a controlling module 144, a memory module 146, an input module 148, a display module 150, a communication / networking module 152, and other modules 154, with these various elements being interconnected by a system bus 156. It should be understood that by referring to various elements herein with the word module, this implies that these elements include hardware.

[0125] The processor module 142 may comprise one or more single-core or multiple-core processors such as INTEL® microprocessors (INTEL is a registered trademark of Intel Corp., Santa Clara, CA, USA), AMD® microprocessors (AMD is a registered trademark of Advanced Micro Devices Inc., Sunnyvale, CA, USA), ARM® microprocessors (ARM is a registered trademark of Arm Ltd., Cambridge, UK) manufactured by a variety of manufactures such as Qualcomm of San Diego, California, USA, under the ARM® architecture, or the like.

[0126] The controlling module 144 comprises one or more controlling circuits, such as graphic controllers, input / output chipsets and the like, for coordinating operations of various hardware components and modules of the computing device 116 / 120.

[0127] The memory module 146 includes one or more non-transitory machine- readable memory units or storage devices accessible by the processor module 142 and the controlling module 144 for reading and / or storing machine-executable instructions for the processor module 142 to execute, and for reading and / or storing data including input data and data generated by the processor module 142 and the controlling module 144. The memory module 146 may be volatile and / or non-volatile, non-removable or removable memory such as RAM, ROM, EEPROM, solid-state memory, hard disks, CD, DVD, flash memory, or the like. In use, the memory module146 is generally divided to a plurality of portions for different use purposes. For example, a portion of the memory module 146 (denoted as storage memory herein) may be used for long-term data storing, for example, for storing files or databases. Another portion of the memory module 146 may be used as the system memory for storing data during processing (denoted as working memory herein).

[0128] The memory module 146 may store images and / or videos received by the animal assessment system. In various embodiments, the memory module 146 may store one or more detected states of animals in association with the stored images and / or videos. In some embodiments, the memory module 146 may store additional data related the assessed animals. The additional data may be stored in association with the stored images / videos and / or detected states of animals. For example, the memory module 146 may store animal identification data (e.g., an identification (ID) number), animal size data (e.g., a height of the animal), animal weight data, animal birth data (e.g., an actual or estimated birth date of the animal), and / or animal progeny data (e.g., number, sex, birth date etc.).

[0129] The input module 148 comprises one or more input interfaces for coupling to various input devices such as a keyboard, a computer mouse, a trackball, a touchpad, a touch-sensitive screen, a touch-sensitive whiteboard, or other human interface devices (HID). The input device may be a physically integrated part of the computing device 116 / 120 (e.g., the touchpad of a laptop computer or the touch- sensitive screen of a tablet) or may be a device that is physically separate therefrom but functionally coupled thereto (e.g., a computer mouse). The input module 148 may also comprise other input interfaces and / or input devices such as microphones, scanners, cameras, Global Positioning System (GPS) components, and / or the like.

[0130] The display module 150 comprises one or more display interfaces for coupling to one or more displays such as monitors, LCD displays, LED displays, projectors, and the like. The display may be a physically integrated part of the computing device 116 / 120 (e.g., the display of a laptop computer or tablet), or may be a display device that is physically separate therefrom but functionally coupled thereto (e.g., the monitor of a desktop computer). In at least one embodiment, the displaymodule 150 may be integrated with a coordinate input to form a touch-sensitive screen or touch-sensitive whiteboard.

[0131] The communication / network module 152 is adapted to communicate with other devices of the animal assessment system 100 through the network 118 using suitable wired or wireless communication technologies such as those described above.

[0132] The computing device 116 / 120 may further comprise other modules 154 such as speakers, printer interfaces, phone components, and / or the like.

[0133] The system bus 156 interconnects various components 142 to 154 for enabling them to transmit and receive data and control signals to and from each other as required. Also included, but not shown, is a power supply and a power bus for providing power to these elements.

[0134] FIG. 3B shows a simplified software architecture 160 of the computing device 116 / 120. The software architecture 160 comprises an operating system (OS) 162, a plurality of application programs 164, an input interface 166 having a plurality of input-device drivers 168, an output interface 170 having a plurality of output-device drivers 172, and a logic memory 174. The operating system 162, application programs 164, input interface 166, and output interface 170 are generally implemented as computer-executable instructions or code in the form of software code or firmware code stored in the logic memory 174 which may be executed by the processor module 142.

[0135] The OS 162 manages various hardware components of the computing device 116 / 120 via the input interface 166 and the output interface 170, manages the logic memory 174, and manages and supports the application programs 164. The OS 162 is also in communication with other computing devices (not shown) via the network 118 to allow application programs 164 to communicate with those running on other computing devices. As those skilled in the art will appreciate, the OS 162 may be any suitable operating system such as MICROSOFT® WINDOWS® (MICROSOFT and WINDOWS are registered trademarks of the Microsoft Corp., Redmond, WA, USA), APPLE® OS X, APPLE® iOS (APPLE is a registered trademark of Apple Inc., Cupertino, CA, USA), Linux, ANDROID® (ANDROID is a registered trademark ofGoogle Inc., Mountain View, CA, USA), or the like. The computing devices 116 / 120 may all have the same OS or may have different OS’s.

[0136] The application programs 164 are executed or run by the processor module 142 (e.g., see: FIG. 3) for performing various tasks. Herein, the application programs 164 running on the server computers 120 (denoted “server programs”) perform tasks to provide server functions for managing network communication with client computing devices 104 and facilitating collaboration between the server computer 102 and the client computing devices 104. The server programs also perform one or more tasks for at least one embodiment of the animal assessment system 100 described herein based on images captured by the imaging devices 108. Herein, the term “server” may refer to a server computer 120 from a hardware point of view or a logical server from a software point of view, depending on the context.

[0137] The application programs 164 running on the client computing devices 116 (denoted “client programs” or so-called “apps”) perform tasks for communicating with the server programs for inquiry of animal assessment data and for displaying same on a display of the client computing devices 104.

[0138] The input interface 166 comprises one or more input device drivers 168 managed by the OS 162 for communicating with the input module 148 and respective input devices. The output interface 170 comprises one or more output device drivers 172 managed by the OS 162 for communicating with respective output devices including the display module 150. Input data received from the input devices via the input interface 166 is processed by the application programs 164 (directly or via the OS 162). The output generated by the application programs 164 is sent to respective output devices via the output interface 170 or sent to other computing devices via the network 118.

[0139] The logic memory 174 is a logical mapping of the physical memory module 146 for facilitating access to various data by the application programs 164. In at least one embodiment, the logic memory 174 comprises a storage memory area (174S) that may be mapped to a non-volatile physical memory such as hard disks, solid-state disks, flash drives, and the like, generally for long-term data storage (including storing data received from other devices, data generated by the application programs 164,the machine-executable code or software instructions of the application programs 164, and / or the like). The logic memory 174 also comprises a working memory area (174W) that is generally mapped to high-speed, and in some implementations, volatile, physical memory such as RAM, generally for application programs 164 to temporarily store data (including data received from other devices, data generated by the application programs 164, the machine-executable code or instruction of the application programs 164, and / or the like) during program execution. For example, a given application program 164 may load data from the storage memory area 174S into the working memory area 174W and may store data generated during its execution into the working memory area 174W. A given application program 164 may also store some data into the storage memory area 174S, as required or in response to a user’s command.

[0140] With the hardware and software structures described above, at least one embodiment of the animal assessment system 100 is configured to perform an assessment of at least one animal at the site 102. In particular, the various embodiments of the animal assessment system 100 described herein generally use a multi-layer artificial intelligence (Al) pipeline that is organized among several layers or levels that each have Al models where each level has a unique overarching purpose. The Al models may be implemented using various machine learning algorithms depending on the functionality of the Al models since some machine learning algorithms are better suited than others at performing certain functions. The Al models may be implemented as computer-executable instructions or code in the form of software code or firmware code stored in the logic memory 174 which may be executed by one or more processors of the processor module 142.

[0141] Examples of machine learning algorithms that may be used by the system 100 include, but are not limited, to one or more of Pre-Trained Neural Network (PTNN), Transfer Learning, Convolutional Neural Networks (CNN), Deep Neural Networks (DNN), Deep Convolutional Neural Networks (DCNN), Fully Connected Networks (FCN), Recurrent Neural Networks (RNN), Long Term Short Term (LSTM), Vision Transformer (ViT) Networks, Pyramid Networks, High-Resolution Networks (HR-Net), a DINO model (self-distillation with no labels), and / or Fine Grained Classification (FGC) such as Jigsaw Networks, for performing certain functions such as, but notlimited to, performing an animal assessment, such as mounting and / or chin-resting, for example, based on its appearance from one or more angles. The neural networks that are used by the system 100 may operate individually or together in a chaining tree. The machine learning algorithms may be trained. At least some of machinelearning models may be trained using a combination of authentic images (e.g., real world images) and synthetic images.

[0142] Referring now to FIG. 4A, shown therein is a schematic diagram of an example of a DNN Al model 400 which generally comprises an input layer 402, a plurality of hidden layers including a first hidden layer 404 and a final hidden layer 408, and an output layer 412 cascaded in series. Each hidden layer comprises a plurality of nodes. For example, first hidden layer 404 includes nodes 406a, 406b, 406c, 406d, etc., and hidden layer 408 includes nodes 410a, 410b, 410c, 41 Od, etc. with each node in a given hidden layer connecting to a plurality of nodes in a subsequent layer. Each node in the hidden layers is a computational unit having a plurality of weighted or unweighted inputs from the nodes of the previous layer that it is connected to (denoted “input nodes” thereof), a transfer function for processing these inputs to generate a result, and an output for outputting the result to one or more nodes of the subsequent layer (denoted “output nodes”) connected thereto.

[0143] Referring now to FIG. 4B, shown therein is a simplified schematic diagram of an architecture of a jigsaw Al model 450. In general, a jigsaw model may be used for fine-grained visual classification (FGCV) and may be configured for identifying subclasses of a given object category. The feature extraction layers of the jigsaw model can be implemented using any feature extractor known to those skilled in the art, including but not limited to Resnet. The jigsaw model can be used in combination with any neural network, such as a CNN, DNN, and ANN (Artificial Neural Network). The jigsaw model 450 employs a neural network to learn features progressively, from a finer granular level to a coarser granularity. A SAM optimize may be used to improve the training on JigSaw models.

[0144] In the jigsaw model 450, each of blocks 452, 454, 456 and 458 correspond to a feature extraction step that includes feature extraction layers, while block 460 corresponds to a convolution block and block 462 to a classification block. Eachfeature extraction block 452, 454, 456 and 458 is configured for extracting and learning features at a different level of granularity. For example, feature extraction blocks 452 and 454 correspond to shallower layers and are configured to learn fine-grained information, while feature extraction blocks 456 and 458 correspond to deeper layers and are accordingly configured to learn more abstract and coarse-grained information. This progressive learning allows the jigsaw model 450 to more accurately perform image classification. Although jigsaw model 450 is shown with four feature extraction blocks / stages, it will be understood that a jigsaw model architecture can have any number of features extraction stages, wherein each feature extraction stage learns at a coarser granularity level.

[0145] In the realm of computer vision, a Jigsaw model processes an image by cutting it into variously sized patches, which are then shuffled to form a puzzle with pieces of differing scales. This introduces complexity to the puzzle-solving task, as the neural network must learn to recognize and piece together relationships across both fine and broad details. The network's architecture typically includes convolutional layers that pick up on local visual cues, enhanced with layers such as recurrent or permutation-invariant ones to manage the shuffled layout, where Sharpness Aware Minimization (SAM) normalization helps the network focus on the most pertinent features. The training process involves coaxing the network to accurately predict the original arrangement of the patches, thereby imparting a nuanced comprehension of spatial relationships and contextual information within the image. Learning to solve puzzles with pieces of different sizes forces the model to understand diverse spatial relationships, enriching its ability to tackle advanced vision tasks such as image classification. This strategy underscores the pivotal role of spatial understanding in computer vision, particularly when deciphering the layout and interrelations of elements in images.

[0146] The Vision Transformer (ViT) represents a paradigm shift in image recognition by applying the principles of the Transformer architecture, traditionally used in natural language processing, to computer vision. Unlike conventional convolutional neural networks, ViT begins by partitioning an input image into a sequence of fixed-size patches and linearly embedding each of them, akin to words in a sentence. These patches are then processed through a series of Transformerencoder layers, which utilize self-attention mechanisms to weigh the importance of different patches relative to one another, allowing the model to focus on the most informative parts of the image. Positional encodings are added to retain the order of the patches. The culmination of this process is a final layer that aggregates the encoded information as an embedding vector. The embedding vector is a highdimensional representation that encodes the positional and contextual information about the image's patches as processed by the self-attention mechanisms within the Transformer layers. Preprocessing forViT involves data augmentation strategies such as random horizontal flipping, Color Jittering for dynamic color changes, random conversion to grayscale, pixel-wise normalization, resizing to a fixed resolution, application of Gaussian Blur for smoothness, and Solarization for photonegative effects. These techniques not only expand the diversity of the training set to prevent overfitting but also help the model to learn more general and robust features. The extensive training required for the Vision T ransformer is due to its reliance on attention mechanisms, which, while computationally intensive, enable the model to achieve state-of-the-art results, especially on large-scale image datasets.

[0147] The DINO model, standing for "Distillation with NO labels," is a selfsupervised learning technique that utilizes Vision Transformers (ViT) or ResNet (Residual Networks) as its backbone architecture to learn rich representations of visual data without the need for labeled datasets. The process is based on a teacherstudent paradigm, where the student model learns to replicate the output of the teacher model. The teacher is a static version of the Vision T ransformer that provides a target for the student to predict, and it is updated with the exponential moving average of the student's parameters over time. In DINO, multiple views of the same image, generated through augmentations such as cropping, resizing, and color distortions, are fed into both the student and teacher networks. The student network, which is a ViT or ResNet, attempts to predict the teacher network's output. Unlike conventional supervised learning, DINO does not rely on predefined labels but encourages the student to learn invariant features that are robust to the changes introduced by the augmentations. As the training progresses, the student gradually learns more generalized features of the images, effectively performing a form of knowledge distillation without labels. This results in the model being capable ofcapturing the essence of the visual patterns present in the dataset. The Vision Transformer's architecture, with its self-attention mechanisms, is particularly well- suited for this approach as it can focus on various parts of an image and understand the global context, making it an excellent backbone for DINO. The representations learned through DINO can then be utilized for a variety of downstream tasks. By adding a classification head, such as a fully connected layer, to the pre-trained student network, the model can be fine-tuned on a labeled dataset for tasks like image classification, object detection, or other tasks that require understanding of visual content. This hybrid approach leverages the unsupervised learning capabilities of DINO to handle complex, high-dimensional data and the structured approach of supervised learning for specific problem-solving, harnessing the power of ViT to achieve state-of-the-art performance on various vision tasks.

[0148] In various embodiments, the image classification models (using a JigSaw classifier, a DINO / FCN model, ViT, etc.) can be trained on 300 epochs with a learning rate of 5e-4. Several different learning rate decay techniques may be evaluated to determine the optimal learning rate decay (e.g., optimal in the sense of the most accurate training), including time-based decay, step decay, exponential decay or adaptive learning rate methods such as Adagrad, Adadelta, or RMSprop. T raining may be stopped early or continue for longer, based on the time taken for the loss graph to converge.

[0149] During training, data augmentations may be applied to the training data, to obtain more training data by applying transformations including but not limited to random horizontal flips and switching, rotation, random crops to upper and lower body, random scale and rotation, color jittering, noise injection, and normalization.

[0150] Training may be performed using loss graphs and accuracy to measure the model’s performance. For example, in various embodiments, for a suitable loss function with binary classification, binary cross-entropy can be used. Training accuracy can be measured using any suitable metric, for example, accuracy, precision, recall, ROC-AUC curve, etc.

[0151] After model training is complete, the model can be prepared for deployment into an Al-pipeline. The preparation for deployment can include a conversion process,for example, a native PyTorch model may be converted into ONNX format. Further, the preparation for deployment can include an optimization process for a target processor. For example, the model converted into ONNX format may be optimized for a target GPU using NVIDIA’S TensorRT toolset, where model optimizations are applied, including, but not limited to, converting all model weights to floating point 16 (FP-16) or 8-bit integers (INT-8) using precision calibration. Other optional optimizations include but are not limited to, layer and tensor fusion, kernel auto-tuning, dynamic tensor memory, sparse weight compression, and dynamic batching.

[0152] In various embodiments, deployment into an Al pipeline may involve certain processing. For example, the processing may include scaling input bounding box content to the required input dimensions of a given Al classification model. As other examples, the processing may reduce the output of the Al classification model to a binary classification and / or convert the output confidence scores to a tuple.

[0153] For training the various classification models described herein, a dataset that includes real and synthetic images that consists of positive and negative cases may be used. For example, synthetic training images may be generated as described in PCT International patent application serial no. PCT / CA2024 / 050321 titled “System and Method for Generation of Simulated Computer Vision Training Data Using a Virtual Reality Engine”, filed March 15, 2024.

[0154] The images can be obtained for various angles. For example, for the synthetic images, a virtual camera can be spun around a virtual animal from different heights, angles and positions to create a varied training set. Real world data is used to balance out the synthetic dataset with different lighting conditions, camera models, fields of view, and lens distortion. For example, a training image may consist of an animal, and a bounding box around the animal and optionally a body part of the animal. Typically, 10,000 training images may be generated, but more may be needed in some cases depending on the accuracy and robustness of the given Al classifier model. In various embodiments, diffusion models may be applied to images included in the training dataset to generate additional training images having greater variance and / or augmenting labelled real-world data using diffusion models. Greater variance may be generated by changing the appearance of the animal and / or the environment. Forexample, the additional training images may be generated by modifying animal fur patterns or colors, changing the background, modifying the weather, adjusting the time of day or lighting, changing the animal breed, etc., as explained earlier.

[0155] Training may be done following a standard training approach, using loss graphs and accuracy to measure the model’s performance. As described previously, the training may be divided into a training set of 70%, validation set of 20% and a testing set of 10%. The training set can include any suitable combination of real-world and synthetic training data. For example, the training set can include 90% synthetic data and 10% real-world data. In other examples, other combinations of synthetic and real-world data may be used. However, for it is preferable that the testing set will only consist of real-world data, and not include any synthetic data.Animal Assessments

[0156] Referring now to FIG. 5, shown therein is a flowchart showing an example embodiment of an animal assessment method 500 that may be implemented using one or more server programs 164 operating on one or more processors of the processor module 142 in at least one embodiment of the system 100 for assessing a state of at least one animal at the site 102 using an Al pipeline. Method 500 may be used for determining if an animal is in a mounting state. For ease of description, a cow (or cattle) is used hereinafter as an example of an animal. However, the methods described herein (including method 500) may be applied to other types of animals such as any livestock and other four-legged mammalian animal, as all mammal exhibit forms of mounting behavior.

[0157] In various embodiments, the method 500 may start automatically (e.g., periodically), manually under a user’s command, or when one or more imaging devices 108 send a request to the server program 164 for transmitting captured (and encoded) images thereto.

[0158] After the method 500 starts, the server program 164 commands the imaging devices 108 to capture images or video clips of at least one animal or an entire herd of animals (collectively denoted “images” hereinafter for ease of description) at the site 102 from one or more viewing angles such as front, side, top, rear, and the like, and starts to receive the captured images or video clips from one or more of the imagingdevices 108 (step 502). As those skilled in the art will appreciate, the images may be captured from one or more angles and over various distances depending on the orientation and camera angles of stationary imaging devices or the flightpath / movement and camera angle for mobile imaging devices. The camera angle of some imaging devices may also be controlled by one of the application programs 164 to obtain images of animals from one or more desired angles.

[0159] The images may be processed, for example, by automatically adjusting image acquisition parameters such as exposure, white balance, contrast, brightness, hue, and / or the like, digital zooming, digital stabilization, and the like. For example, functions that are available through the Vision Programming Interface (VPI) in the DeepStream SDK (offered by NVIDIA CORPORATION of Santa Clara, CA, U.S.A.) may be used for the above processing operations such as performing color balancing on images. The images may be processed, for example, using Al pipeline 3300 described in United States Patent No. 11 ,910,784. Software built into OpenCV or other libraries made available by Nvidia or other companies may also be used. The images may additionally be resized, for example, if the images used when training the machine learning algorithms described herein have a predetermined, standardized size. In such cases, the images may be resized to match the size of the training images. The RGB values of the images may additionally be normalized to transform the pixel intensity values to a standard scale. In at least one embodiment, some of the image processing is not needed such as when training uses images obtained using various image acquisition parameters and the machine learning algorithms described herein do not require images to be pre-processed. In some cases, image processing can involve dewarping an image such as, for example, when an image obtained using a 360-degree camera.

[0160] At step 504, the server program 164 uses one or more machine learning algorithms to process the images captured at step 502 and perform as an object detector to detect, locate and optionally identify animals in one or more of the acquired images. Various machine learning algorithms, such as, but not limited to, CNN or DNN based algorithms, may be used to detect animals from the acquired images. For example, in at least one embodiment, a single-pass bounding box algorithm may be used as the object detector and may be implemented using a Single Shot detector(SSD) algorithm, a DetectNet or DetectNet_V2 model, and / or the like. Alternatively, in at least one embodiment, a two-stage network may be used such as FasterRCNN, YOLOX or some other appropriate object detection and / or identification model. Alternatively, in at least one embodiment, a two-stage network may be used such as FasterRCNN or some other appropriate object detection model. In at least one embodiment, A mask model, such as MaskRCNN, that combines a mask with a detector may also be used. In some cases, some image pre-processing may be performed such as, for example, in cases where an image is de-warped such as for images that are obtained using a 360-degree camera. In at least one embodiment, images from acquired video may be scaled to be 720p for the Al models used in layer 0 and layer 1 of the Al pipelines described herein. In another alternative, in at least one embodiment, a combination of machine learning algorithms can be used to detect animals in the acquired image(s). For example, a first machine learning algorithm may be used to detect whether an object in an image is an animal, and a second machine learning algorithm may be used to identify the type of animal. The use of more than one machine learning algorithm for detecting and / or identifying animals in acquired image(s) can improve the accuracy of the object detector. For example, one or more of the techniques described in United States Patent No. 11 ,910,784 may be used.

[0161] The object detector detects one or more animals in a current image that is processed and also defines a bounding box to indicate the locations of the one or more detected animals in the image. The bounding box may be overlaid on the image to encapsulate or surround the detected animal and may be saved as a new image in at least one embodiment. For example, these images may be used for training in the future. The image may be saved in one of two forms 1) the entire image with the coordinates of the bounding box or 2) only the pixels for the animal defined by the bounding box. At this step, at least one animal in the image, including animals partially obscured or turned away from the camera, are detected and associated with a respective bounding box in at least one embodiment of the system 100. FIG. 8 shows an example image 342 in which three animals 104A, 104B, and 104C are detected, for which respective bounding boxes 344A, 344B, and 344C are generated and defined thereabout.

[0162] Once the animals are detected and the bounding boxes are generated, the coordinates of certain points (e.g., x1 , x2, y1 , y2) defining the corners of the bounding box or a combination of coordinates and dimensions of the bounding box (e.g., x, y, height, width) may be stored in a data store along with the image numberforthe image that is currently being processed and the animal ID (which is uniquely assigned to a given animal and may be done as described in United States Patent No. 11 ,910,784) so that the bounding box can be retrieved from the data store at a later point and used for various purposes such as certain types of measurements. The bounding boxes can also be used in animal assessments, as will be described in further detail with reference to steps 604-612 of method 600. It should be noted that a data store may be any of the memory elements described herein. In some cases, the server program 164 can determine the size of a bounding box (e.g., based on the coordinates or dimensions of the bounding box) and discard or not perform assessments on images where detected animals in those images are bounded by bounding boxes having dimensions below a predetermined threshold since bounding boxes that are too small may lead to inaccurate assessments.

[0163] In at least one embodiment, the bounding boxes may be generated such that at least one of the bounding boxes is a rectangular box, a square box or an oriented bounding box. The use of oriented bounding boxes may result in bounding boxes that are smaller but still encapsulate the detected animal since the oriented bounding box is rotated and may not include as many pixels compared to a nonoriented bounding box. The use of oriented bounding boxes may also reduce the amount of other items that are encapsulated by the oriented bounding box resulting in more accurate analysis of the image pixels encapsulated by the oriented bounding box.

[0164] One or more detected animals may then be identified and assessed at step 610, as will be described in further detail.

[0165] Optionally, masks may be generated for the one or more detected animals by identifying each image pixel associated with a detected animal. A mask may be generated, for example, by converting every image pixel not associated with a detected animal to black (e.g., a zero pixel value). Mask images may improve theaccuracy of models used to perform assessments, for example, image classification models used to detect one or more states of an animal. Generating the mask images may require additional computing resources and / or processing time.

[0166] At step 506, the server program 164 uses one or more machine learning algorithms to process the images captured at step 602 to determine if one or more of animals in a given image matches a mounting state. The server program 164 for example, uses a mounting state detection model, which is an Al model. The mounting state detection model can be a machine learning algorithm that is part of the Al pipeline of Al models and that is trained to generate a mounting classification output identifying whether an animal is in a mounting state or in a not mounting state. The machine learning algorithms can process images to identify all animals in a given image that are in a mounting state. As described, a mounting state for a first animal (e.g., first cow) can be indicative of an estrus state for a second animal (e.g., second cow) which is being mounted by the first cow and accordingly results of assessments determining whether an animal is in an estrus state or a mounting state can be used for animal management since th animal in the estrus state may be ready for artificial insemination and the animal in the mounting state may enter the estrus state within the foreseeable future such as in about 12 hours or so from when the mounting detection occurred. The mounting state detection model may be any suitable image classification model trained to generate classification output for an imaged animal such as, but not limited to, a CNN, DNN, jigsaw or ViT Al model. The classification output can correspond to a mounting state or a not mounting state from which it can be determined that another animal is or is not in the estrus state.

[0167] For both Jigsaw and DINO / FCN approaches, the animal bounding box may be scaled to 224x224 pixels or 256x256 pixels to ensure uniformity. In the preprocessing phase for training DINO / FCN and Jigsaw models, the learning rate may be configured to complement the prepared data, to pursue optimal convergence. A higher initial learning rate may be set, often between 0.0005 to 0.05 for DINO, and slightly lower for Jigsaw models, with both employing a cosine decay schedule for gradual reduction over 100 to 300 epochs. A warm-up period is included where the learning rate incrementally reaches the target value to stabilize early training dynamics. Adaptive methods like Adam or SGD with momentum may be used,particularly with Jigsaw, adjusting the learning rate per parameter. This strategy may be fine-tuned through hyperparameter optimization, tailored to the model’s needs and the nuances of the dataset, post standard resizing, normalization, and augmentation procedures in pre-processing. SAM may be used during the fine-tuning phase when the model is being adapted to a specific supervised task like classification. In the context of fine tuning a Jigsaw model, SAM may be employed after the self-supervised pre-training phase, where the model has learned to assemble the jigsaw puzzles and thus has gained an understanding of the data's structure. In at least one embodiment, performing mounting state detection may involve several processing steps and using several Al models, such as is described in the example embodiment shown in FIG. 6.

[0168] In some embodiments, prior to determining whether one or more animals in a given image match a mounting state, e.g., before running the mounting state detection model, the server program 164 may assesses the bounding boxes defined at step 504. For example, in such embodiments, the server program 164 may identify whether at least two animals are present in an image and if so whether there are one or more pairs of overlapping bounding boxes (e.g., at least a portion of the image is enclosed by both bounding boxes). The image pixels enclosed by each bounding box can be determined based on the coordinates of the corners / edges of the bounding box. For example, the server program 164 may determine that bounding boxes 1120 and 1122 (in FIG. 11 B) corresponding to detected animals 1124 and 1126 respectively are overlapping bounding boxes. The server program 164 can assess each pair of bounding boxes in an image. For example, in an image containing three bounding boxes Bi, B2 and B3, corresponding to three animals, the server program 164 can evaluate each pair of bounding boxes, e.g., B1B2, B1B3 and B2B3. The server program 164 can perform mounting detection on the image portions corresponding to the bounding boxes only.

[0169] Reference is now made to FIG. 6, which shows a flowchart showing an example embodiment an animal assessment method 600 for determining if a current state of a detected animal in a given image matches a mounting state. Method 600 may be implemented using one or more server programs 164 operating on one or more processors of the processor module 142 in at least one embodiment of the system 100.

[0170] In various embodiments, method 600 may be executed at step 506 of method 500. For example, method 600 maybe be executed for each pair of animals detected and located in a given image at step 504. Method 500 may be executed for each image received at step 502 that pass filtering criteria (e.g., the size of the bounding boxes is larger than a threshold size in order to obtain more accurate results).

[0171] Optionally, the method 600 may be executed as a stand-alone process independent of method 500. For example, the method 600 may be executed automatically (e.g., periodically) to perform animal assessments for captured images / videos. In various embodiments, method 600 may be executed in response to an input identifying one or more animals for assessment. For example, the method 600 may be executed to monitor mounting behavior of animals that have been identified and shown to be receptive to mounting and / or have shown other signs of estrus. The IDs of these identified animals may be stored in a database. The method 500 may be executed for images / videos tagged as including the identified animal(s).

[0172] If method 600 is executed independently of method 500, captured images / videos are received at step 602. The server program 164 uses a machine learning algorithm to perform as an object detector to detect one or more animals in a received image and define a bounding box to indicate the location of the detected animal as described previously. The bounding boxes can be determined as described with reference to method 500. If method 600 is executed at step 508 of method 500, the animal is detected and the bounding box is defined at step 504 of method 500.

[0173] At step 604, the server program 164 determines one or more pairs of overlapping bounding boxes, i.e. , at least a portion of the image is enclosed by both bounding boxes. The server program 164 can determine pairs of overlapping bounding boxes as described with reference to method 500.

[0174] The server program 164 can then perform age classification on the animals located within the overlapping bounding boxes. In various embodiments, the age classification may be performed for all animals included in a received image (e.g., the age classification may be conducted at step 602). Alternatively, the age classification may only be performed for animals corresponding to at least one pair of overlappingbounding boxes. The age classification model may also generate a confidence value for the generated age classification output. In various embodiments, age classification can be performed only on image portions enclosed by bounding boxes. By performing age classification only on a subset of detected animals, the speed and efficiency of the system 100 can be increased and fewer computing resources expended.

[0175] To perform age classification, the server program 164 can use an age classification model, which is an Al model. The age classification model can be a machine learning algorithm that is part of the Al pipeline of Al models that is trained to generate an age classification output such as, for example, the Al pipeline described in United States Patent No. 11 ,910,784. The age classification model may be any suitable image classification model trained to generate classification output for an imaged animal. For example, the age classification model can output a binary classification (i.e., adult, not adult) or a multi-class classification (e.g., an age range of the animal). The age classification model may be any suitable image classification model trained to generate an adult or non-adult classification output for an imaged animal. For example, an age classifier can be implemented using a computer vison model such as a Jigsaw model or DINO / FCN. The extracted features for the age classification may include, for example, a size dimension (e.g., height) of the animal. The age classification model can, for example, classify animals 952 and 954 of image 900 (FIGS. 9A-9B) as adult animals. The age classification model may also generate a confidence value for the generated age classification output. An effective age classification model for animals may be obtained by collecting and preparing a diverse dataset with images labeled by age, followed by preprocessing to standardize and augment the data. The model architecture, which may be a DINO / FCN or Jigsaw model, may be customized forage classification by adjusting the output layer for binary or multi-class classification. Features relevant to age, such as size and morphological characteristics, may be used for training the model, utilizing strategies like transfer learning and regularization to prevent overfitting and enhance performance. Since estrus only occurs in (female) adult animals, determining whether a detected animal is an adult animal can allow the server program 164 to only identify a mounting state in adult animals, and avoid using computing resources for detecting animals being mounted that are not sexually mature.

[0176] The training data used to train the age classification model may include real and / or synthetic images corresponding to adult and non-adult animals as described previously. Each training image may correspond to a single class, for example, an adult class or non-adult class. The model’s effectiveness may be evaluated through k- fold cross-validation and metrics like accuracy and F1 -score, with a focus on generalization. Integration into the Al pipeline and deployment may include generating confidence scores for each prediction to indicate reliability.

[0177] In an alternative embodiment, the identification of the detected animal may be used to determine the age by using information, such as age data, on the identified animal that is stored in a database.

[0178] At step 606, the server program 164 determines if both animals in the bounding boxes are adult animals based on the results of the age classification model for both animals. If both animals are adult animals, method 600 proceeds to step 608. If the server program 164 determines that one or both animals in the bounding boxes are not adult animals, the method 600 proceeds to step 616 and processes the next received image that needs processing.

[0179] At step 608, the server program 164 can perform standing detection using an activity classification model. The activity classification model can be a machine learning algorithm that is part of the Al pipeline of Al models that is trained to generate an activity classification output identifying whether an animal is in a standing position or a non-standing position. In various embodiments, the activity classification may be performed for all detected adulted animals included in a received image (e.g., the age classification may be conducted at step 604). Alternatively, the activity classification may only be performed for animals corresponding to at least one pair of overlapping bounding boxes. In at least one embodiment, the server program 164 only performs standing detection on animals that have been determined to be adult animals at step 606. Since mounting requires an animal to be in a standing position, determining whether an animal is in standing position can allow the server program 164 to only perform mounting detection on adult animals that are in a position that allows for mounting. Though step 608 is shown as being subsequent to step 604, in some cases, step 608 may be performed prior to step 604. In such cases, in at least oneembodiment, the server program 164 only performs age classification on animals that have been determined to be in a standing position. In various embodiments, activity classification can be performed only on image portions enclosed by bounding boxes.

[0180] The activity classification model may be any suitable image classification Al model trained to generate classification output for an imaged animal. The classification output can correspond to an activity of the animal, including standing, having one or more legs not in contact with the ground, and non-standing activities such as, but not limited to, lying down. The activity classification model may use a Vision Transformer (ViT) approach, where preprocessing involves resizing images to a standard dimension (e.g., 224x224 pixels or 256x256 pixels) to ensure uniformity, which is used for the ViT's patch-based processing. Images are also normalized to maintain consistent scale across the dataset. Various parameters that may be set for the ViT include the patch size, such as 16x16 pixels for example, the transformer's depth indicating the number of encoding layers, and the number of attention heads, which aids in capturing global image dependencies. The learning rate is carefully chosen, incorporating a warm-up period and a subsequent decay schedule to enhance training stability and efficacy. Additionally, data augmentation techniques such as random cropping, flipping, and color jittering are applied to improve the model's ability to generalize across diverse data.

[0181] The training data may include real and / or synthetic images corresponding to standing and non-standing animals. The training may generally be performed as described earlier. In supplementing the training of the activity classifier, in at least one embodiment advanced techniques may be incorporated such as He or Xavier initialization for optimal weight settings, batch normalization to enhance stability, and dropout for preventing overfitting. Addressing class imbalances through class weight balancing aids for unbiased learning and employing transfer learning by using pretrained ViT models can expedite convergence and enhance performance. Dynamic learning rate adjustments with schedulers like ReduceLROnPlateau can promote efficient learning, while integrating attention mechanisms within the ViT allows for selective focus on relevant image regions. Comprehensive hyperparameter tuning, alongside diligent monitoring of the training process via validation curves, may be usedfor fine-tuning the model to achieve peak accuracy and generalization in activity classification tasks.

[0182] The extracted features for the activity classification may include, for example, angles between links and / or distance between key points on legs of an animal (e.g., indicating whether the legs are erect). The key points may be determined using various techniques including those described in United States Patent No. 11 ,910,784. The activity classification model can, for example, classify animals 1014 and 1016 of image 1000 (FIGS. 10A-10B) as standing animals and animals 1010 and 1012 as non-standing animals. The activity classification model may also generate a confidence value for the generated activity classification output.

[0183] At step 610, the server program 164 determines if both animals in the overlapping bounding boxes are standing. If both animals are standing, the method 600 proceeds to step 612. If one or both animals are in a non-standing state, the method 600 proceeds to step 616, where it is determined if there are any other overlapping bounding boxes that need to be processed.

[0184] Optionally, at step 612, the server program 164 selects an animal for which to perform mounting detection. In at least one embodiment, the server program 164 only performs mounting detection for the animal in an animal pair that is in a higher position relative to the other animal and / or relative to a reference point (e.g., the ground). In such embodiments, the server program 164 can determine which animal in an animal pair is in a higher position based on the coordinates of the highest / topmost corners / edges of the bounding box of each animal. For example, the server program 164 may determine that bounding box 1120 is higher than bounding box 1122 (in FIG. 11 B). If both animals are at a substantially similar height (e.g., the heights / topmost corners of the bounding boxes are within 1 , 2, 5, or 10 percent of on another), the server program 164 may perform mounting detection for both animals. If mounting is detected, then the animal in the higher bounding box is considered to be the “mounter” and the lower one is considered to be the “mountee”. The cow in the lower bounding box that is the mountee is the one that is in estrus.

[0185] At step 614, the server program 164 performs mounting detection using a mounting state detection model for the animal(s) selected at step 612.

[0186] As described with reference to step 506, the server program 164 uses one or more machine learning algorithms to process the images captured at step 602 to determine if one or more of the animals in a given image matches a mounting state. The server program 164 for example, uses a mounting state detection model. The mounting state detection model can be a machine learning algorithm that is part of the Al pipeline of Al models that is trained to generate a mounting classification output identifying whether an animal is in a mounting state or in a non-mounting state. The mounting detection may only be performed for animals corresponding to at least one pair of overlapping bounding boxes. In at least one embodiment, the server program 164 only performs mounting detection on animals that have been determined to be adult animals at step 606 and / or that have been determined to be in a standing position at step 610. Since mounting requires an animal to be (i) in a standing position, (ii) to be an adult animal, and (iii) to be in close proximity with another animal, performing mounting detection only on animals that have been determined to be (i) adult animals, (ii) in a standing position and (iii) in an overlapping bounding box can allow the mounting detection to only be performed on a subset of animals and / or only on images that include adult animals in a standing position in an overlapping box, reducing computing resources required by the system 100.

[0187] The mounting state detection model may be any suitable image classification model trained to generate classification output for an imaged animal. The classification output can correspond to a mounting state or a non-mounting state. The mounting state detection model may be implemented and trained as was described at step 506 of method 500 in FIG. 5.

[0188] The training data for training the mounting state detection model may include real and / or synthetic images corresponding to mounting and non-mounting animals. FIGS.7A-7C for example, show example synthetic images 710, 720, 730 that may be included in the training data. Synthetic images such as synthetic images 710, 720 and 730 can be generated using real images or using previously generated synthetic images and can correspond to real images or previously generated synthetic images in which patterns (e.g., colors, background, fur patterns, orientation, etc.) have been changed. The training may be performed similarly to the general training that was described previously. For example, the images included in the training data maybe pre-processed before being used for training the mounting state detection model. The processing may include, but not be limited to resizing the images, applying color jitter to the images so that the brightness, contrast, hue, and / or saturation of the images can be varied, and / or normalizing the RGB values of the images, for example.

[0189] The extracted features for mounting detection may include, for example, a position of the shoulder portion and / or back portion of the animal. Referring briefly to FIGS. 12A-12H, which show example images of animals in a mounting state and associated heatmaps, showing features and / or areas of the image that provide highly useful information that contributes to the classification of animals in a mounting state. As shown in heatmaps of FIGS. 12B, 12D, 12F and 12H, features related to the shoulder and back portion of an animal highly contribute to the detection of a mounting state.

[0190] The mounting state detection model can, for example, classify animal 1124 in image 1100c (FIG. 11 C) as an animal in a mounting state. As another example, the server program 164 can classify animal 2010 in image 2000 (FIG. 20A) as an animal in a mounting state. The mounting state detection model may also generate a confidence value for the mounting state classification output.

[0191] In at least one embodiment, the detection and timespan for standing heat may be used along with the detection of a mounting event to improve the accuracy of determining when the mountee animal is experiencing estrus. For example, this may be assessed by determining whether the lower animal (e.g., mountee animal) stays relatively still for a certain period of time.

[0192] As shown in FIG. 11 B, in various embodiments, the server program 164 may only perform mounting detection on an image portion enclosed by a bounding box.

[0193] At step 616, the server program 164 determines if all pairs of overlapping bounding boxes have been processed. If all pairs of overlapping bounding boxes in an image have been processed, the server program 164 proceeds to the next image. If all pairs of overlapping boxes have not been processed, the server program 164 returns to step 604, and repeats steps 604 to 616, until all overlapping bounding boxes have been processed.

[0194] If method 700 is implemented independently, the server program 164 may, at step 614 generate an output indicating that the current state of the detected animal matches a mounting state in response to detecting a mounting state. If method 600 is implemented at step 506, the output may be generated at step 508. The output may be generated as a user notification. The user notification may include information related to the current state of the detected animal, and one or more images or videos of the detected animal. For example, the user notification may include a location information for the animal and an image corresponding to the estrus state or the mounting state. Some examples of user notifications are provided herein.

[0195] Optionally, the server program 164 stores in memory (e.g., in memory module 146) the processed image in association with the detected state of the animal. The stored data may be used for subsequent time-series or historical assessment of one or more animals.

[0196] Returning to FIG. 5, at step 508, the server program 164 generates an output indicating the assessment results. For example, in response to determining that at least one animal in a given image is in a mounting state, the server program 164 can generate a notification. In some cases, the server program 164 may only generate a notification if the confidence value of the mounting classification output meets a predetermined mounting confidence threshold. The notification can indicate that the current state of an animal matches an estrus state or a mounting state. In at least one embodiment, the notification can include identification data associated with the animal.

[0197] The output may be generated as a user notification that is provided to one or more user devices. The user devices can include, for example, personal computers, smartphones, tablet devices, workstations etc. For example, FIG. 22A shows a user notification 2204a provided at a display 2202a of a user smartphone and FIG. 22B shows a user notification 2204b provided at a display 2202b of a user computer. The user notification can be provided in any way that allows a user to be notified that an animal is in an estrus state or a mounting state.

[0198] The user notification 2204 may include information related to the current state of the detected animal, and one or more images or videos of the detected animal. The one or more images or videos can be annotated images or videos, showingbounding boxes and / or labels indicating animals in a mounting state and identification data (e.g., identification number). The user notification may include detected state information 2210, and date and time information 2212, animal location information 2214, image / video 2216 associated with the detected state. In some embodiments, the server program 164 may provide additional data related to the animal assessment. For example, the server program 164 may provide an image 2220 of an animal in an estrus state or a mounting state after providing the related state alert.

[0199] In some embodiments, the output at step 508 may be provided via a custom user interface. For example, FIGS. 23A and 23B show example embodiments of GUI 2300a and GUI 2300b respectively of an administration portal that is displayed on the monitor of a desktop computer. GUI 2300 may display multiple windows including, for example, a menu 2302. The menu 2302 may provide a user with various GUI viewing options and / or animal management / assessment options.

[0200] In the illustrated examples, a “Dashboard” view is shown. In the “Dashboard” view GUI 2300 may include various display portions, for example, an image / video portion 2404, a graph portion 2406, a calendar portion 2408 and a notification portion 2410. The image / video portion 2404 may display maps, captured images / videos, and / or processed images / videos that may include one or more animals. The graph portion 2406 may display one or more animal assessment results.

[0201] In various embodiments, one or more outputs related to the animal assessments may be provided via the notification portion 2310 (for example, a notification 2312 related to estrus detection). In some embodiments, one or more outputs related to the animal assessments may be provided via the calendar portion 2308 (for example, a historical estrus detection indicator 2314).

[0202] In various embodiments, the server program 164 may perform postprocessing related to detected mounting states detected at step 506 (of method 500) and / or act 614 (of method 600). The post-processing may be based, for example, on data stored in memory (e.g., in memory module 146) related to the detected states of the animals. In some embodiments, additional data (e.g., animal identification data, animal location tracking data etc.) may be stored in combination with the data related to the detected states. As part of the post-processing of mounting detection,Confidence Score Assessment, Control on Alert Frequency, and / or Integration with multi-modal large language models (MM-LLMs). Confidence score cutoffs, used in conjunction with bounding boxes sizes, may be used as cut-offs for a mounting indication. Control Alert Frequency checks to ensure that repeated or similar events to not product more than a single notification to a customer. MM-LLMs may be used to take basic JSON outputted data and / or images and convert the notification into a human readable format.

[0203] Referring now to FIG. 13, shown therein is a flowchart showing an example embodiment of an animal assessment method 1300 that may be executed by a server program 164 of the animal assessment system of FIG. 1A, for detecting a standing heat in at least one animal at the site using one or more Artificial Intelligence (Al) models which might be in an Al pipeline. Method 1300 may be used in combination with method 500 and 600. As described, standing heat is a primary sign of estrus and corresponds to a period in the reproductive cycle of a female animal where the animal will allow itself to be mounted. As used herein, a standing heat state is characterized by an animal being continuously receptive to mounting (e.g., continuously being mounted) for a predefined period of time, for example, at least three seconds. Standing heat can be detected using the teachings described herein by evaluating a series of images containing multiple images, or video frames, obtained chronologically. For example, the server program 164 can receive a video showing animals and extract frames chronologically at a predetermined time interval (e.g., one to a few frames per second) or extract frames occurring within a short period of time. Alternatively, the server program 164 can receive timestamped images, captured within a short period of time (e.g., up to within 1 to a few seconds of each other). As used herein, “consecutive frames” or “consecutive images” refers to images obtained within a short period of time and that may not necessarily correspond to images captured immediately after one another (e.g., consecutive frames in a video may not correspond to frames extracted at the frame rate of the video and some frames may be omitted). For example, in the set of images h, l2, , I4, Is obtained at ti, t2, ts, t4 and ts, images h, h and Is may be referred to as consecutive images, if the images are captured within a short period of time. Obtaining and processing several images captured within a short period of time, on the order of about 3 to about 5 seconds, where the timestampsof consecutive images are not more than 1 second apart, can avoid standing heat from being detected in images that are captured within very short periods of time (e.g., within milliseconds, faster than an animal can move).

[0204] At step 1302, the server program 164 receives newly captured images or images stored in a data store. Step 1302 may be substantially similar to step 502 of method 500.

[0205] At step 1304, the server program 164 uses one or more machine learning algorithms to process the images captured at step 1302 and perform as an object detector to detect, locate and optionally identify animals in one or more of the acquired images. Step 1304 may be substantially similar to step 504 of method 500 and the same Al models may be preferably used for object detection, although other may be used in other embodiments. Bounding boxes around the detected animals are also obtained when the animal is detected.

[0206] At step 1306, the server program 164 uses one or more machine learning algorithms to process the images captured at step 1302 to determine if one or more of animals in a given image matches a mounting state. Step 1306 may be substantially similar to step 508 of method 500. Step 1306 can also include steps substantially similar to steps 604-612 of method 600, that is, the server program 164 can identify overlapping bounding boxes, perform age classification and / or activity classification, which may be done as was described for method 600, for example. However, step 1306 may additionally involve identifying the animal being mounted, since a standing heat state is, as described, characterized by an animal being receptive to mounting for a certain period of time. If the server program 164 determines that one or more animals in a given image matches a mounting state, the server program 164 determines the corresponding one or more animal(s) being mounted by the one or more or more animals in a mounting state, by determining the lower bounding box for these animals being mounted and the animals being mounted are identified as being in the estrus state.

[0207] At step 1308, the server program 164 determines if at least one animal in a given image is in a standing heat state. As described above, a standing heat state can be characterized by an animal being continuously receptive to mounting for apredefined time period, for example, at least about three seconds. Accordingly, a standing heat state can be characterized by an animal being mounted in each image in the series of images obtained over the predefined time period (e.g., a mounting time period). For example, for a given video, if the predetermined time period is set to be at least three seconds and frames are obtained at a rate of one frame per second, three consecutive frames containing the same animal being mounted indicate that the animal being mounted is in a standing heat state which also indicates that the animal being mounted is in an estrus state.

[0208] In various embodiments, as described with reference to FIGS.1-2, each animal can be associated with identification data. To determine that the same animal is being mounted in a predetermined number of consecutive images (corresponding to the predefined time period), the server program 164 can use the identification data associated with each animal.

[0209] In at least one embodiment, to determine if an animal being mounted is in a standing heat state, the server program 164 may evaluate the movement of the animal being mounted. Since a standing heat state involves an animal being continuously receptive to mounting for at least a predetermined time period, a standing heat state may be characterized by minimal movement of the animal. In at least one embodiment, to evaluate the movement of the animal being mounted, the server program 164 can determine the overlap between the bounding boxes in each pair of consecutive images. In at least one embodiment, to determine the overlap, the server program 164 may calculate the intersection over union (i.e., the ratio between the intersection and the union) of the bounding box of the animal determined to be mounted of image h obtained at ti and the bounding box of image h obtained at t2. The loll measure is more preferable than the intersection area since that the latter only indicates the size of the overlapping area between two bounding boxes while the loU measure reflects the proportion of this overlap relative to the combined size of both bounding boxes. Accordingly, the loU measure is unaffected by the objects' size, providing a consistent metric regardless of scale. For example, the server program 164 may calculate the intersection between and the union of bounding box 1410 in image 1400a (FIG. 14A), and bounding box 1420 in image 1400b (FIG. 14B), bounding box 1420 in image 1400b (FIG. 14B), and bounding box 1430 in image 1400c (FIG. 14C), bounding box1430 in image 1400c (FIG. 14C), and bounding box 1440 in image 1400d (FIG. 14D) and bounding box 1440 in image 1400d (FIG. 14D), and bounding box 1450 in image 1500e (FIG. 14E). Image 1400f (FIG. 14F) for example shows the bounding box 1460 of the current frame and the bounding box 1450 of frame (image) 1400e.

[0210] The server program 164 can then compare the calculated overlap with a predetermined threshold (which may be referred to as a standing heat threshold). A large overlap between the bounding box at time ti and the bounding box at time t2 (e.g., a ratio between the intersection and union of the bounding boxes tending toward 1) indicates minimal movement of the animal being mounted, between temporally successive / consecutive frames in the sequence of frames being analyzed. Conversely, little overlap between the bounding boxes at time ti and the bounding boxes at time t2 (e.g., a ratio between the intersection and union of the bounding boxes tending toward 0) indicate that the animal being mounted is likely to have moved in between the successive frames, which suggests that standing heat is unlikely to occur, since standing heat is associated with relatively small displacement of the animal being mounted. The threshold can vary according to the sensitivity of the Al model and a desired level of accuracy. For example, the threshold can be 0.45, 0.7, or any other suitable threshold value. The server program 164 may repeat the process for each pair of consecutive images. The server program 164 may determine that an animal is a standing heat state if the calculated intersections for all pairs of consecutive images meet the predetermined threshold.

[0211] In at least one embodiment, upon determining that all pairs of consecutive frames are associated with a calculated intersection meet the predetermined standing heat threshold, the server program 164 can additionally calculate the overlap between the bounding boxof the first image and the bounding box of the last image of the series of consecutive images. For example, for a series of three consecutive images, the server program 164 can calculate the overlap between the bounding box of the first image and the bounding box of the third image. The server program 164 can compare the calculated intersection with a supplemental predetermined standing heat threshold, which may be the same threshold as the standing heat threshold used for comparing bounding boxes in consecutive images or may be a different threshold. If the calculated intersection meets the supplemental predetermined standing heatthreshold, the server program 164 may determine that the animal is in a standing heat state and thus also in an estrus state. The use of these two thresholds aid in reducing the occurrence of false positives. In standing heat, the mountee animal does not move, which is expected to occur after the mounting event. Therefore, when determining whether the mountee animal is experiencing standing heat it is not desirable for the bounding boxes for the mountee animal to be overlapping with other bounding boxes.

[0212] At step 1310, the server program 164 generates an output indicating the assessment results. For example, in response to determining that at least one animal in a series of consecutive images is in a standing heat state and thus the estrus state, the server program 164 can generate a notification. The notification can indicate that the current state of an animal matches a standing heat state. In at least some embodiments, the location and identification of the animal having a standing heat state and thus estrus state may be included and / or saved in a database. Step 1310 may be substantially similar to step 508 of method 500. While FIGS. 22A and 22B only show user notifications for a mounting state, it will be understood that similar user notifications can be generated for a standing heat state.

[0213] Referring now to FIG. 15, shown therein is a flowchart showing an example embodiment of an animal assessment method 1500 that may be implemented using one or more server programs 164 operating on one or more processors of the processor module 142 in at least one embodiment of the system 100 for assessing a state of at least one animal at the site 102 using one or more Al models of an Al pipeline. Method 1500 may be used for assessing if an animal is in a chin-resting state. The other animal that is receiving a chin-resting on its rear end may be identified as being in an estrus state. The animal in the estrus state may be ready for artificial insemination and the animal in the chin-resting state may enter the estrus state within the foreseeable future such as in about 12 hours or so from when the chin-resting detection occurred. Method 1500 may be used in combination with method 500, 600 and / or method 1300 or may be used separately. For example, method 1500 can be used to supplement the detection results obtained using method 500, 600 and / or method 1300 which may all be used as indicators of estrus.

[0214] Similar to method 500, method 1500 may start automatically (e.g., periodically), manually under a user’s command, or when one or more imaging devices 108 send a request to the server program 164 for transmitting captured (and encoded) images thereto.

[0215] At step 1502, the server program 1504 can receive captured images. Similar to step 502 of method 500, the server program 164 commands the imaging devices 108 to capture images or video clips of at least one animal or an entire herd of animals. Step 1502 may be substantially similar to step 502 of method 500.

[0216] At step 1504, the server program 164 uses one or more machine learning algorithms to process the images captured at step 502 and perform as an object detector to detect, locate and optionally identify animals in one or more of the acquired images. Step 1504 may be substantially similar to step 504 of method 500 and the same Al models may be preferably used for object detection, although other may be used in other embodiments. Bounding boxes around the detected animals are also obtained when the animal is detected.

[0217] At step 1506, the server program 164 uses one or more machine learning algorithms to process the images captured at step 1502 to determine if one or more of animals in a given image matches a chin-resting state. For example, the server program 164 can perform chin-resting detection using a chin-resting classification model. The machine learning algorithms can process images to identify all animals in a given image that are in a chin-resting state. The chin-resting classification model can be a machine learning algorithm that is part of the Al pipeline of Al models that is trained to generate a mounting classification output identifying whether an animal is in a chin-resting state or in a non-chin-resting state. As described, chin-resting in cattle can be indicative that a second cow is in an estrus state when it receives a chin on its rear end from a first cow that is in the chin-resting state and accordingly results of assessments determining whether an animal is in an estrus state or chin-resting state can be used for animal management. The chin-resting detection model may be any suitable image classification model trained to generate classification output for an imaged animal. The classification output can correspond to a chin-resting state or a non-chin-resting state. The training data may include real and / or synthetic imagescorresponding to animals that are resting their chins on the backside of other animals and images that shown animals that are not performing chin-resting. The training may generally be performed as described earlier. The implementation of the chin-resting detection model and its training is described with reference to FIG. 16.

[0218] In some embodiments, determining whether one or more animals in a given images match / have a chin-resting state involves assessing the bounding boxes defined at step 1504. In such embodiments, the server program 164 identifies whether at least two animals are present in an image and determines if one or more pairs of bounding boxes overlap (i.e., at least a portion of the image is enclosed by both bounding boxes) similar to step 506 of method 500. The image pixels enclosed by each bounding box can be determined based on the coordinates of the corners / edges of the bounding box. For example, the server program 164 may determine that bounding boxes 1820 and 1822 (in FIG. 18B) correspond to detected animals 1824 and 1826 respectively are overlapping bounding boxes. The server program 164 can assess each pair of bounding boxes in an image. The server program 164 can perform chin-resting detection on the image portions corresponding to the bounding boxes only.

[0219] Reference is now made to FIG. 16, which shows a flowchart showing an example embodiment of an animal assessment method 1600 for determining if a current state of a detected animal in a given image matches a chin-resting state and thus a detected animal receiving the chin on its rear end is in an estrus state. Method 1600 may be implemented using one or more server programs 164 operating on one or more processors of the processor module 142 in at least one embodiment of the system 100.

[0220] In various embodiments, method 1600 may be executed at step 506 of method 500. For example, method 1600 maybe be executed for each animal detected and located in a given image at step 1504. Method 1600 may be executed for each image received at step 1502.

[0221] Optionally, the method 1600 may be executed as a stand-alone process independent of process 1500. For example, the method 1600 may be executed automatically (e.g., periodically) to perform animal assessments for captured / storedimages / videos. In various embodiments, method 1600 may be executed in response to an input identifying one or more animals for assessment. For example, the method 1600 may be executed to monitor chin-resting in animals that have shown signs of estrus. The process 1600 may be executed for images / videos tagged as including the identified animal(s).

[0222] If method 1600 is executed independently of method 1500, captured images / videos are received at step 1602. The server program 164 uses a machine learning algorithm to perform as an object detector to detect one or more animals in a received image and define a bounding box to indicate the location of the detected animal. The bounding boxes can be determined as described with reference to method 500. If method 1600 is executed at step 1510 of method 1500, the animal is detected and the bounding box is defined at step 1504 of method 1500. Otherwise, step 1602 is performed.

[0223] Steps 1602, 1604, 1606, 1608, 1610, 1616 and 1618 may be substantially similar to steps 602, 604, 606, 608, 610, 616 and 618 of method 600, respectively, and implemented in a similar fashion.

[0224] At step 1602, the server program 164 determines one or more pairs of overlapping bounding boxes, i.e. , at least a portion of the image is enclosed by both bounding boxes. The server program 164 can determine pairs of overlapping bounding boxes as described with reference to method 500.

[0225] At step 1604, the server program 164 can perform age classification on detected animals. Age classification can be performed as described with reference to step 604 of method 600 and as shown in FIGS. 9A-9B.

[0226] In various embodiments, the age classification may be performed for all animals included in a received image (e.g., the age classification may be conducted at step 1602). Alternatively, the age classification may only be performed for animals corresponding to at least one pair of overlapping bounding boxes. In various embodiments, age classification can be performed only on image portions enclosed by bounding boxes.

[0227] At step 1606, the server program 164 determines if both animals in a pair of overlapping bounding boxes are adults. If both animals are adults, the method 1600 proceeds to step 1608. If one or both animals are not adults, the method 1600 proceeds to step 1616 to process another pair of overlapping bounding boxes.

[0228] At step 1608, the server program 164 can perform standing detection using an activity classification model. Activity classification can be performed as described with reference to step 608 of method 600 and as shown in FIGS. 10A-10B.

[0229] In various embodiments, the activity classification may be performed for all animals included in a received image (e.g., the age classification may be conducted at step 1602). Alternatively, the activity classification may only be performed for animals corresponding to at least one pair of overlapping bounding boxes. For example, in at least one embodiment, the server program 164 only performs standing detection on animals that have been determined to be adult animals at step 1606. Though step 1608 is shown as being subsequent to step 1604, in some cases, step 1608 may be performed prior to step 1604. In such cases, in at least one embodiment, the server program 164 only performs age classification on animals that have been determined to be in a standing position. In various embodiments, activity classification can be performed only on image portions enclosed by bounding boxes.

[0230] At step 1610, the server program 164 determines if both animals in the pair of overlapping bounding boxes are in standing position. If both animals are in a standing position, the method 1600 proceeds to step 1612. If one or both animals are not in a standing position, the method 1600 proceeds to step 1616 where data related to another pair of bounding boxes is processed.

[0231] At step 1612, the server program 164 combines the overlapping bounding boxes. In at least one embodiment, the server program 164 generates a merged bounding box for the pair of overlapping bounding boxes and performs chin-resting detection on the merged bounding box. The merged bounding box can be generated to include each portion of the image (i.e. , every image pixel) that is enclosed by at least one of the two overlapping bounding boxes. For example, image 1800c (FIG. 18C) shows merged bounding box 1830 corresponding to overlapping bounding boxes1820 and 1822. The server program 164 can perform chin-resting detection only on the merged bounding box 1830 (also corresponding to image 1890d in FIG. 18D).

[0232] At step 614, the server program 164 performs chin-resting detection using a chin-resting detection model for the animals in the merged bounding box generated at step 1612.

[0233] As described with reference to step 1506 of method 1500, the server program 164 uses one or more machine learning algorithms to process the images captured at step 1602 to determine if one or more of animals in a given image matches a chin-resting state. For example, the server program 164 can perform chin-resting detection using a chin-resting classification model. The chin-resting classification model can be a machine learning algorithm that is part of the Al pipeline of Al models that is trained to generate a mounting classification output identifying whether an animal is in a chin-resting state or not in chin-resting state. The chin-resting detection may only be performed for animals corresponding to at least one pair of overlapping bounding boxes. In at least one embodiment, the server program 164 only performs standing detection on animals that have been determined to be adult animals at step 1604 and that have been determined to be in a standing position at step 1608. Since chin-resting indicating estrus requires an animal to be in a standing position, to be an adult animal, and to be in close proximity with another animal, performing chin-resting detection only on animals that have been determined to be adult animals, in a standing position and in an overlapping bounding box can allow the chin-resting detection to only be performed on a subset of animals and / or only on images that include adult animals in a standing position in an overlapping box, reducing computing resources required by the system 100.

[0234] The chin-resting detection model may be any suitable image classification model trained to generate classification output for an imaged animal. For example, the JigSaw model, which is generally described earlier, may be used. The classification output can correspond to a chin-resting state or a non-chin-resting state. The chinresting model uses the union of two animal bounding boxes, that includes the animals and the background, as shown in 1830 in FIG. 18C. Either a DINO / FCN or Jigsaw training methodology may be utilized. For Jigsaw model training, a binary classificationis modeled, by preparing the dataset such that the data is cleanly segmented and appropriately labeled for the two classifications. A layer that is used for binary classification is a dense layer with a single output and a sigmoid activation function, providing a probability score indicating class membership. Using binary cross-entropy as the loss function, with an optimizer such as Adam with an appropriate learning rate (e.g., of 0.001) for efficient training. Model training involves dividing the data into training (80%), validation (10%), and test sets (10%), with input sizes tailored to the input SAM data such as 256x256 pixels. The training data may include real and / or synthetic images corresponding to animals that are resting their chins on the backside of other animals and images that shown animals that are not performing chin-resting. The training may generally be performed as described earlier.

[0235] The training data may include real and / or synthetic images corresponding to chin-resting and non-chin-resting animals. FIGS.17A-17B for example, show example synthetic images 1710, 1720 that may be included in the training data. Synthetic images such as synthetic images 1710, 1720 can be generated using real images or using previously generated synthetic images and can correspond to real images previously generated synthetic images in which patterns (e.g., colors, background, fur patterns, etc.) have been changed. The images included in the training data may also be pre-processed as described with reference to FIGS. 7A-7C.

[0236] During training, the goal may be for a balance between precision and recall, and monitoring metrics such as F1 score and ROC-AUC may be used during validation to adjust hyperparameters like batch size (e.g., 32 or 64) and number of epochs (typically between 10 and 100, with early stopping applied to prevent overfitting). For the optimization process, an SAM algorithm may be used to further refine the model weights. SAM seeks parameters that lie in neighborhoods having uniformly low loss, which can lead to improved generalization by sharpening the minima of the loss landscape. SAM may be used in used in conjunction with a primary optimizer, such as Adam, by periodically adjusting the learning rate (dual-step update) to navigate the loss landscape more effectively. This approach encapsulates the end-to-end process, from data preparation to deployment, focusing on achieving high generalization on unseen data while being open to iterative refinement based on performance feedback.

[0237] The extracted features for chin-resting detection may include, for example, a position and / or angle of an animal’s neck and / or chin, and / or a position of an animal’s chin relative to another animal’s head. Referring briefly to FIGS. 19A-19H, which show example images of animals in a chin-resting state and associated heatmaps, showing features contributing to the classification of animals in a chin-resting state. As shown in heatmaps of FIGS. 19B, 19D, 19F and 19H, features related to the face, chin and neck portion of an animal highly contribute to the detection of a chin-resting state.

[0238] The server program 164 can for example, classify, by performing chinresting detection on the merged bounding box 1830, animal 1824 as an animal in a chin-resting state and animal 1124 in image 1100c (FIG. 11 C) as an animal in a chinresting state. The other animals in each of those images that are receiving the chin on their rear end are identified as animals in an estrus state. As another example, the server program 164 can classify animal 2052 in image 2050 (FIG. 20B) as an animal in a chin-resting state and animal 2054 in image 2050 as an animal not in a chinresting state. The chin-resting detection model may also generate a confidence value for the generated chin-resting classification output.

[0239] As shown in FIG. 18D, in various embodiments, the server program 164 may only perform chin-resting detection on an image portion enclosed by a bounding box.

[0240] Similar to method 600, if method 1600 is implemented independently, the server program 164 may, at step 1614 generate an output indicating that the current state of the detected animal matches a chin-resting state in response to detecting a chin-resting state. If method 1600 is implemented at step 1606, the output may be generated at step 1508. The output may be generated as a user notification. The user notification may include information related to the current state of the detected animal, and one or more images or videos of the detected animal. For example, the user notification may include a location information and an image corresponding to the chinresting state and identifying the detected animal that is in the estrus state.

[0241] Optionally, the server program 164 stores in memory (e.g., in memory module 146) the processed image in associated with the detected state of the animal.The stored data may be used for subsequent time-series or historical assessment of one or more animals.

[0242] Returning back to FIG. 15, at step 1508, the server program 164 generates an output indicating the assessment results. For example, in response to determining that at least one animal in a given image is in a chin-resting state, the server program 164 can generate a notification. The notification can indicate that the current state of an animal matches an estrus state or a chin-resting state and may also include the location and identification of the animal that is in the estrus state or the chin-resting state. Step 1508 may be substantially similar to step 508 of method 500.

[0243] While FIGS. 22A and 22B only show user notifications for a mounting state, it will be understood that similar user notifications can be generated for chin-resting.

[0244] In at least one embodiment, steps 1502-1506 can be repeated over a predetermined period of time to evaluate changes in chin-resting behaviors. For example, the server program 164 can calculate the number of chin-resting events occurring over a predetermined period of time and generate an output at step 1508 to indicate the number of chin-resting events detected. If the number of chin-resting state events exceeds a predetermined threshold over a predetermined time period, which may be referred to as a predetermined chin-resting count threshold and a predetermined chin-resting time period, respectively, then the server program 164 may generate an output indicating a high likelihood of breeding receptivity. As another example, the server program 164 can calculate changes in the number of chin-resting events over a predetermined time period and generate an output at step 1508 based on the calculated changes. An increase for a given time period of chin-resting events can indicate a high likelihood of breeding receptivity. For instance, if it is observed that the number of chin-resting events within a 24-hour period increases by more than 20% compared to the previous 24 hours, then send a notification may be sent to personnel who run the animal facility. The 20% threshold may be adjusted to obtain better test results.

[0245] Referring briefly now to FIG. 21 , shown therein is a flowchart showing another example embodiment of an animal assessment method 2100 that may be executed by a server program of the animal assessment system of FIG. 1A, fordetecting a state indicative of estrus in at least one detected animal at a monitoring site, such as a ranch for example, using one or more Artificial Intelligence (Al) models which may be part of an Al pipeline.

[0246] Method 2100 may combine various steps of methods 500, 600, 1500 and 1600. For example, in at least one embodiment, steps 1922, 2104, 2106 and 2108 may be substantially similar to steps 502, 504, 506, 508 and steps 1502, 1504, 1506 and 1508.

[0247] At step 2110, the server program 164 determines whether to perform mounting state detection and / or chin-resting detection based on several bounding box criteria which may also be referred to as trigger conditions. If the trigger conditions for performing the mounting state detection and the chin-resting state detection are true for animals in overlapping bounding boxes, then both of these will be performed. For example, if a pair of overlapping bounding boxes have two animals that are adults and are in a standing position and one of the bounding boxes is higher than the other bounding box, the server program 164 may proceed to step 2112 and perform mounting detection using data associated with the higher bounding box. Alternatively, the server program 164 may proceed to step 2114 and perform chin-resting detection on the combined bounding box that includes the higher bounding box and the lower bounding box if the trigger condition for the chin-resting detection is met which is there is one adult animal in each of the bounding boxes that are overlapping and the animals are standing. Mounting detection may be performed according to methods 500 and 600. Chin-resting may be performed according to methods 1500 and 1600.

[0248] If the bounding boxes in a pair of overlapping bounding boxes are of the same height, the server program 164 may proceed to step 2114 and perform chinresting detection. Alternatively, the server program 164 may proceed to step 2112 and perform mounting detection on both bounding boxes, separately.

[0249] In at least one embodiment, the server program 164 can be configured to perform both mounting detection and chin-resting detection for the same pair of overlapping bounding boxes. For example, the server program 164 may perform mounting detection for the animal enclosed by the higher bounding box and apply chin-resting detection for the pair of animals in the pair of overlapping bounding boxes.

[0250] At step 2116, the server program 164 may amalgamate the detection results for the detection models that were performed to generate an assessment results based on the determined states of the animals in the bounding boxes such as, but not limited to, if the animal in the higher bounding box is in a mounting state or a chinresting state and if the animal in the lower bounding box is in a standing heat state I estrus state, for example. Confidence scores and confidences score thresholds may also be used to assess the detection results as described below.

[0251] At step 2118, the server program 164 generates an output indicating the assessment results. For example, in response to determining that at least one animal in a given image is in one of a chin-resting state or a mounting state and the other animal being mounted or receiving the chin is in an estrus state, the server program 164 can generate a notification. The notification can indicate that the current state of an animal matches a chin-resting state, a mounting state or an estrus state. Step 2116 may be substantially similar to step 508 of method 500 and step 1508 of method 1500.

[0252] Referring now to FIG. 24A and 24B, shown therein are images of an example user notification 2400 and 2450, respectively, provided by at least one embodiment of an animal assessment system described herein. In such embodiments, the various user notifications may include an image of the detected event as well as text that describes what is in the image and indicates the state of the animals in the image. In FIG. 24A, a portion of an image of user notification 2400 is shown including an image 2410, location and time information 2420, image description text 2430 and a detected event indication 2440. The image 2410 was obtained from a plurality of images based on the analysis performed according to the teachings herein to determine that one of the cows in the image 2410 is experiencing estrus. The image description text 2430 includes text describing what is in the image 2410 and the detected event indication 2440 includes text that describes the detected event. In FIG. 24B, a portion of an image of user notification 2450 is shown including an image 2460, image description text 2480 and a detected event indication 2490. The image 2460 was obtained from a plurality of images based on the analysis performed according to the teachings herein to determine that one of the cows in the image 2460 is experiencing estrus. The image description text 2480 includes text describing what is in the image 2410 and the detected event indication 2490 includes text that describesthe detected event. For example, the image description text 2480 is: “The cow in the image is exhibiting mounting behavior, where it is placing its front legs on the backend of another cow. This posture is indicative of estrus, a period when cows are sexually receptive and ready to mate. The arched back and the positioning are clear signs of mounting behavior.” This notification may then be used by the operator of the animal facility (e.g., ranch) to get the cow that is experiencing estrus ready for artificial insemination.

[0253] In at least one embodiment, user notifications, such as user notifications 2400, for example, may be obtained by processing an input image, such as image 2410, using a multi-modal deep learning Al model that is provided with a first description of estrus detection signs and a second description of estrus detection reporting signs. Both of these descriptions may be written using common language and not necessarily programming language. The descriptions are provided to the multi-modal deep learning Al model to allow the model to learn the estrus detection signs to look for in the input image and also learn the estrus detection reporting signs to use when generating the image description text and the detected event indication in the user notification. The multi-modal Al model is an Al model that can analyze text (e.g., the first and second descriptions) and can analyze images (e.g., the input image) to determine whether an animal in the input image is experiencing estrus and generate text to describe what is shown in the input image.

[0254] For example, the multi-modal deep learning Al model is similar to a large language model but may process both input text and an input image and map them to the same embedding space. The multi-modal deep learning Al model may be an OpenAI model such as, but not limited to, GPT4 Vision, DALL-E, Gemini Pro, or Gemini Ultra, for example. The multi-modal deep learning Al models can process and understand both textual and visual inputs, allowing them to perform tasks that involve understanding and generating content across different media types. The multi-modal deep learning Al models are very computationally intensive and / or time-consuming. Advantageously, in accordance with the teachings herein, the estrus detection models and methods described herein are used, which can be more computationally efficient and orders of magnitude faster, and act as a filter to process a plurality of images to obtain a smaller set of images, that are more likely to include an animal experiencingestrus and can be referred to as candidate estrus images. There may be one or more Al models that are used for prefiltering the plurality of images to obtain the candidate estrus images where these one or more Al models may be a keypoint Al model, an image classification Al model, or an ensemble of Al models where each of these models may be referred to as unimodal Al models that operate on the plurality of images and objects which may be visually related to the images where these objects may be keypoints and / or bounding boxes. The candidate estrus images are then provided to the more computationally intensive multi-modal deep learning Al model which is used to further analyze each candidate estrus image to determine the type of estrus is being experienced based on what is shown in the candidate calving image and provide text to describe the estrus state. In some embodiments, the candidate estrus image may include a bounding box around the two animals that are being assessed for estrus. This may enable the multi-modal model to focus its analysis on the detected animal(s) included in the bounding box and reduce analysis time and / or computational resources consumed. In other embodiments, the candidate estrus image may not include a bounding box. This may enable the multi-modal model to include the environment / context around the detected animal(s) during analysis.

[0255] In at least one embodiment, a multi-modal deep learning Al model may be trained to generate user notifications with more descriptive text, as described above where user notification 2400 is an example, by providing an introduction description, an estrus detection signs description, and a reporting estrus signs description.

[0256] The introduction description is used to train the multi-modal deep learning Al model to perform certain functions. An example of an introduction description may be as follows:"Introduction:","BETSY is a mounting watcher that helps in analyzing early signs of estrus in cows. It acts like the eyes of the rancher.","It visually analyzes the image uploaded and detects if the cows are mounting or not, in turn also indicating",if the cow is in estrus or not.","What is Estrus in Cows?","Estrus in Cows is actually the time-period when cows are attracted towards each other and try to indulge in mounting.","During estrus, cows exhibit increased social behavior such as mounting, which is a sign of submissive behavior and","readiness to mate.","Mounting is explained as follows:","1. [[Mounting]]:","One cow places its front legs on the backend of another cow's body. The cow trying to mount the cow below it has","an arched back. This event is an estrus-related behavior for the cow being mounted, the cow mounting or both of","the cows within the event.","Sometimes, both the cows might not be clearly visible but even if there appears to be one cow trying to stand on its","back legs and rest of the body is in the air, this is a strong sign of mounting.","In short, if the cow's position is mounting position (or half standing up, on its two legs trying to get onto another",'cow from the rear-end, it is a clear sign of Estrus.)","Additionally, even if the cow / cows seems to be having a slightly arched back, this indicated that the cow / cows are","trying to mount, and it is definitely a clear sign of mounting.","Analyzing and Reporting Results:","After utilizing the information provided for the mounting behaviour mentioned above, analyze the image uploaded by the","user and give your decision based on the conditions mentioned below:","-> If the cow / cows in the image are in mounting position, explain the reason for mounting in cows and add the","following Note:","Keep the responses professional, not at all in first person. Also, try to keep the overall analysis to less than 70 words."" To conclude, the cow in the image is mounting."," Classification = Mounting (Please add this at the end, keep the sentence intact without any modifications)","-> If the cows in the image are not in the Mounting position, add the following Note:","Keep the responses professional, not at all in first person. Also, try to keep the overall analysis to less than 70 words."" To conclude, the cows in the image is not mounting."," Classification = Not_Mounting (Please add this at the end, keep the sentence intact without any modifications)"

[0257] In at least one embodiment described herein, the confidence score that is determined by the Al models when classifying whether a detected animal in an image being analyzed is in the mounting, chin-resting and standing heat states may be used as a filter by comparing the confidence score to a confidence score threshold. This may aid with reducing the number of false positives. For example, alerts may be issued only if they meet a first confidence threshold (e.g., at least 80%, 85% or higher) for mounting and standing heat, and a second higher confidence threshold for chinresting.

[0258] In at least one embodiment, another technique which may be used to mitigate the issue of frequent alerts, is to specify that there are a certain number of alerts that are sent per camera, per animal, over a certain period of time, such as about 2 minutes, about 5 minutes, about 10 minutes or about 15 minutes. This reduces the likelihood of redundant alerts. The parameters may be tuned to improve efficacy of the generated alerts.

[0259] In at least one embodiment, mounting may be considered to be more of a typical behavior for cows in estrus compared to chin-resting since chin-resting may occur for cows that are not in estrus. Accordingly, in at least one embodiment, the mounting and standing heat classification models may be used as the primary detection methods, with the chin-resting classification model may be used as a supplementary detector. For example, the number of occurrences of detected chinresting events may be counted and any increase in the frequency of chin-resting may be used to determine whether cows in a herd are in estrus.

[0260] In accordance with the teachings herein, each of the estrus detection methods described herein may be modified to include sex classification as a prefilter / pre-screening step before performing any of the steps for detecting mounting,chin-resting and / or the standing heat by checking that the animals in the pair of bounding boxes are female. This may be done by running a sex classifier, which is an Al model, on the detected animal or using the identification of the detected animal to find information on the animal in a database where the information includes sex data for the detected animal. This will reduce false positives since male animals may exhibit behaviours that may be mistaken as estrus related behaviours. The sex classifier may be implemented by providing data or features from the animal section corresponding to its genitals to an Al model that has been trained to determine the animals sex based on data from this animal section. The section of the image pertaining to the animal’s genitals may be determined using various techniques, some of which are described in United States Patent No. 11 ,910,784.

[0261] Although embodiments have been described above with reference to the accompanying drawings, those of skill in the art will appreciate that variations and modifications may be made.

[0262] For example, while the animal assessments are described primarily in terms of cows, it should be understood that these techniques can be applied to other animals by training the various Al models in a somewhat similar fashion using images and other data from the other species.

[0263] It should also be noted that for any description herein of a process, method or step that may be executed by a server program should be understood as meaning that at least one processor or other computing device / electronics (e.g., an ASIC, dedicated hardware, etc.) is executing one or more programs, which may be stored on a server or another storage device, for performing the functions defined by the program.

[0264] While the applicant's teachings described herein are in conjunction with various embodiments for illustrative purposes, it is not intended that the applicant's teachings be limited to such embodiments as the embodiments described herein are intended to be examples. On the contrary, the applicant's teachings described and illustrated herein encompass various alternatives, modifications, and equivalents, without departing from the embodiments described herein, the general scope of which is defined in the appended claims.

Claims

CLAIMS:1 . A method for animal management comprising: receiving, at a computing device, one or more images captured by one or more imaging devices; processing the one or more images using one or more Al models for: detecting and locating two animals with a pair of overlapping bounding boxes in a given image of the one or more images; determining if a current state of a first animal of the two detected animals in the given image matches a mounting state and when the first animal is in the mounting state identifying a second animal of the two detected animals as being mounted by the first animal which is indicative that the second animal is in an estrus state; and generating a notification when the second animal is in the estrus state or the first animal is in the mounting state.

2. The method of claim 1 , wherein determining if the current state of the first detected animal matches the mounting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model, the age classification model being trained to generate an age classification output, the age classification output being one of: adult or non-adult; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position, using an activity classification model, the activity classification model being trained to generate an activity classification output, the activity classification output being one of: standing or non-standing; and in response to the pair of overlapping bounding boxes corresponding to two adult animals being in a standing position, applying a mounting state detection model to an image portion encapsulated by the bounding box for the first animal, to generate a mounting classification output for the detected animal associated with the image portion, the mounting classification output indicating whether the first animal in the image portion is mounting the second animal which is indicative that the second animal is in the estrus state and the first animal will be in the estrus state in the future.

3. The method of claim 1 , wherein determining if the current state of the first detected animal matches the mounting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model, the age classification model being trained to generate an age classification output, the age classification output being one of: adult or non-adult; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position, using an activity classification model, the activity classification model being trained to generate an activity classification output, the activity classification output being one of: standing or non-standing; in response to the pair of overlapping bounding boxes corresponding to two adult animals being in a standing position, applying a mounting state detection model to an image portion encapsulated by the bounding boxes of the first detected animal, to generate a mounting classification output for the first detected animal associated with the image portion, the mounting classification output indicating whether the first animal in the image portion is mounting the second animal; and if the second animal being mounted stays stationary for a predefined period of time after the mounting has finished, identifying the second animal being mounted as being in a standing heat state and the estrus state.

4. The method of claim 2 or 3, wherein determining if the first detected animal matches the mounting state further comprises: in response to the pair of overlapping bounding boxes corresponding to two adult animals, each adult animal being in a standing position, identifying the bounding box from the pair of bounding boxes corresponding to a higher bounding box relative to a ground in the given image; and applying the mounting state detection model to the image portion corresponding to the higher bounding box to generate the mounting classification output.

5. The method of any one of claims 2 to 4, wherein determining if the current state of the first detected animal matches the mounting state further comprises:in response to the pair of overlapping bounding boxes corresponding to two adult animals, determining a corresponding size of each of the bounding boxes of the pair of overlapping bounding boxes, and in response to the size of each the bounding boxes satisfying a predetermined minimum size, applying the mounting state detection model.

6. The method of any one of claims 2 to 5, wherein the mounting state detection model is trained to: generate a mounting classification output distinguishing the mounting state from other non-mounting states; and generate a confidence value for the mounting classification output.

7. The method of claim 6, further comprising: determining if the confidence value for the mounting classification output satisfies a predetermined threshold; and in response to determining the confidence value for the mounting classification output satisfies the predetermined threshold, generating the notification.

8. The method of any one of claims 3 to 7, wherein the one or more Al models comprise an animal identification model for identifying the one or more detected animals and associating a unique identifier with each of the one or more detected animals in the given image and wherein providing the notification comprises including the unique identifier of the animals in the mounting state, the estrus state and / or the standing heat state.

9. The method of any one of claims 1 to 8, wherein locating a given detected animal in the given image comprises generating a mask identifying each pixel of the given image associated with the given detected animal.

10. The method of any one of claims 1 to 9, wherein the notification comprises information related to the current state of each of the detected animals, and one or more image or videos of each of the detected animals.11 . The method of any one of claims 3 to 10, further comprising: annotating the one or more images to indicate the animal matching the estrus state, the mounting state and / or the standing heat state; and displaying, on a display, the annotated one or more images.

12. The method of claim 11 , wherein the notification comprises an annotated image including the animals matching the estrus state, the mounting state and / or the standing heat state.

13. A method for animal management comprising: receiving, at a computing device, one or more images captured by one or more imaging devices; and processing the one or more images using one or more Al models for: detecting and locating two animals with a pair of overlapping bounding boxes in a given image of the one or more images; and determining if a current state of a first animal of the two detected animals in the given image matches a chin-resting state and when the first animal is in the chin-resting state identifying a second animal of the two detected animals as receiving a chin of an animal nears its rear end which is indicative that the second animal is in an estrus state; and generating a notification when the second animal is in the estrus state or the first animal is in the chin-resting state.

14. The method of claim 13, wherein determining if the current state of the first detected animal matches the chin-resting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model the age classification model being trained to generate an age classification output, the age classification output being one of: adult or non-adult;determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position using an activity classification model, the activity classification model being trained to generate an activity classification output, the activity classification output being one of: standing or non-standing; in response to the pair of overlapping bounding boxes corresponding to two adult animals in a standing position, generating a merged bounding box, the merged bounding box encapsulating an image portion that includes the two standing adult animals; and generating a chin-resting classification output by using a chin-resting state classification model to analyze the merged bounding box, the chin-resting classification output indicating whether the first detected animal in the image portion is in a chin-resting state which is indicative that the second animal is in the estrus state and the first animal will be in the estrus state in the future.

15. The method of claim 13, wherein determining if the current state of the first detected animal matches the chin-resting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model, the age classification model being trained to generate an age classification output; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position using an activity classification model, the activity classification model being trained to generate an activity classification output, the activity classification output being one of: standing or non-standing; in response to the pair of overlapping bounding boxes corresponding to two adult animals in a standing position, generating a merged bounding box, the merged bounding box encapsulating an image portion that includes the two standing adult animals; and generating a chin-resting classification output by using a chin-resting state classification model to analyze the merged bounding box, the chin-resting classification output indicating whether the first detected animal in the image portion is in a chin-resting state which is indicative of the second animal being in the estrus state.

16. The method of any one of claims 14 to 15, wherein determining if the current state of the first detected animal matches the chin-resting state further comprises: in response to the pair of overlapping bounding boxes corresponding to two adult animals, determining a corresponding size of each of the bounding boxes of the pair of overlapping bounding boxes, and in response to the size of each the bounding boxes satisfying a predetermined minimum size, applying the chin-resting state classification model.

17. The method of any one of claims 14 to 16, wherein the chin-resting state classification model is trained to: generate a chin-resting classification output distinguishing the chin-resting state from other non-chin-resting states; and generate a confidence value for the chin-resting classification output.

18. The method of claim 17, further comprising: determining if the confidence value for the chin-resting classification output satisfies a predetermined threshold; and in response to determining the confidence value for the chin-resting classification output satisfies the predetermined threshold, generating the notification.

19. The method of any one of claims 13 to 18, wherein locating a given detected animal in the given image comprises generating a mask identifying each pixel of the given image associated with the given detected animal.

20. The method of any one of claims 13 to 19, wherein the notification comprises information related to the current state of each of the at least one detected animals, and one or more image or videos of each of the detected animal.21 . The method of any one of claims 14 to 20, further comprising: annotating the one or more images to indicate each of the at least one detected animal matching the estrus state, and / or the chin-resting state; and displaying, on a display, the annotated one or more images.

22. The method of claim 21 , wherein the notification comprises an annotated image including the at least one detected animal matching the estrus state, and / or the chinresting state.

23. A method for animal management comprising: receiving, at a computing device, a plurality of images captured by one or more imaging devices, the plurality of images corresponding to images sequentially obtained during a predefined time window; and processing the plurality of images using one or more Al models for: detecting and locating two animals with overlapping bounding boxes in a given image of the one or more images; and determining if a current state of a first animal of the two detected animals in the given image matches a mounting state; when the first animal is in the mounting state, determining if a second animal of the two detected animals is in a standing heat state, wherein the second animal is in the standing heat state if the second animal remains stationary for at least several images sequentially obtained over a predefined period of time after the second animal is no longer being mounted; and providing a notification in response to when the second animal is determined to be in the standing heat state.

24. The method of claim 23, wherein determining if the current state of the first detected animal in the given image matches a mounting state comprises: determining if the pair of overlapping bounding boxes corresponds to two adult animals using an age classification model, the age classification model being trained to generate an age classification output, the age classification output being one of: adult or non-adult; determining if the pair of overlapping bounding boxes corresponds to two animals in a standing position, using an activity classification model the activity classification model being trained to generate an activity classification;in response to the pair of overlapping bounding boxes corresponding to two adult animals being in a standing position, applying a mounting classification image classification model to an image portion encapsulated by the bounding box of the first detected animal to generate a mounting classification output for the first detected animal, the mounting classification output indicating whether the first detected animal in the image portion is mounting the second detected animal.

25. The method of any one of claims 1 to 24, wherein one of the bounding boxes is a rectangular bounding box, a square box or an oriented bounding box.

26. The method of any one of claims of 1 to 25, wherein the method comprises providing a candidate estrus image to a multi-modal deep learning Al model that is trained to analyze the candidate estrus image to generate a user notification including the candidate estrus image, image description text and a detected event indication where the image description text includes text describing what is shown in the candidate estrus image and the detected event indication includes text that describes a detected estrus event.

27. The method of claim 26, wherein the user notification further includes location and time information for indicating where and when the candidate estrus image was obtained.

28. The method of claim 26 or claim 27, wherein determining the candidate estrus image is based on detecting estrus according to the method of any one of claims 1 to 25.

29. The method of any one of claims 26 to 28, wherein the multi-modal deep learning Al model is trained by providing to the multi-modal deep learning Al model: (a) an introduction description that is used to train the multi-modal deep learning Al model to perform certain functions, (b) an estrus detection signs description that is used to train the multi-modal deep learning Al model to learn estrus detection signs to look for in the candidate calving image and (c) a reporting estrus signs description that is used to train the multi-modal deep learning Al model to learn estrus detectionreporting signs to use when generating the image description text and the detected event indication in the user notification.

30. The method of any one of claims 1 to 29, wherein after the at least one animal is detected, a sex of the detected at least one animal is obtained using a sex classifier, and further processing of the given image is not performed when the at least one detected animal is male.31 . The method of any one of claims 3 to 30, wherein the notification indicates that the detected animal in a mounting state or in a chin-resting state will be experiencing estrus within the near future.

32. A computing device comprising a memory storing program instructions and a processor that is coupled to the memory to read and execute the program instructions which configure the processor to perform a method for animal management including detecting at least one animal that is experiencing estrus, wherein the method is defined according to any one of claims 1 to 31 .

33. A non-transitory computer readable medium storing thereon program instructions that, when executed by a processor of a computing device configure the processing for performing a method for animal management including detecting at least one animal that is experiencing estrus, wherein the method is defined according to any one of claims 1 to 31 .

Citation Information

Patent Citations

  • Method for identifying estrus behaviors of ruminant based on artificial intelligence

    CN111685060A

  • Apparatus for detecting mounting behavior of cattle

    KR102527058B1

  • Mounting behavior detection system and detection method

    US20180035648A1

  • Ai-based livestock management system and livestock management method thereof

    US20220022427A1

  • A system and a device for health and fertility management of one or more milch animals

    WO2020031050A1