Human animal conflict mitigation system
The human-animal conflict mitigation system addresses the limitations of existing systems by using image classification and cloud storage to differentiate between dangerous and harmless animals, reducing maintenance needs, and enhancing efficiency.
Patent Information
- Application Number
- PCT/IN2024/052145
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-08
AI Technical Summary
Existing human-animal conflict mitigation systems lack image classification capabilities, unable to differentiate between dangerous and harmless animals, and require periodic physical maintenance due to the lack of cloud storage for captured images.
The system employs a camera to capture images, a reference image module to select reference images, an image comparison module to detect objects, an image cropping module to transmit relevant image portions, and an animal detection and recognition module using a trained animal identification model to classify animals as dangerous or non-dangerous.
The system effectively differentiates between dangerous and harmless animals, reduces bandwidth requirements through image cropping, and eliminates the need for periodic physical maintenance by storing images in the cloud, thereby enhancing the efficiency and reliability of human-animal conflict mitigation.
Smart Images

Figure IN2024052145_08052025_PF_FP_ABST
Abstract
Description
HUMAN ANIMAL CONFLICT MITIGATION SYSTEMCross-Reference To Related Applications:
[0001] This application claims the benefit of and priority to Indian Provisional Application 20231 1074057, filed on October 31, 2023, the entirety of which is incorporated herein by reference.Technical Field:
[0002] The present disclosure relates to human animal conflict mitigation system using artificial intelligence.Background:
[0003] Certain areas (e.g., villages), that are inhabited by humans, are proximate to animal habitats (e.g., forests). Some types of animals are capable of endangering the lives of humans, and as such, by cautioning or alerting the human population about the nearby presence of such dangerous animals, the necessary precautions (e.g., retreating to safety) can be timely taken to thwart the threat posed by the dangerous animals.
[0004] Existing technology to caution the human population or mitigate conflicts between humans and animals, suffer from the following drawbacks. They only utilize motion sensor cameras, and a result, do not have any image classification capabilities. The lack of image classification capability results in the technology being unable to differentiate between a dangerous animal (e.g., a tiger) and a harmless animal (e.g., a rabbit). A harmless animal is less likely to create a conflict with a human, and as such, a human need not be cautioned about the nearby presence of a harmless animal. Further, the images captured by the existing technology are not stored in the cloud. As a result, periodic physical maintenance (e.g., physically replacing the memory card in the camera) is required.Summary:
[0005] These and other problems are generally solved or circumvented, and technical advantages are generally achieved, by advantageous embodiments of the present disclosure.
[0006] A summary' of certain embodiments disclosed herein is set forth below. It should be understood that these aspects are presented merely to provide the reader with a brief summary of these certain embodiments and that these aspects are not intended to limit the scope of this disclosure. Indeed, this disclosure may encompass a variety of aspects that may not be set forth below.
[0007] In a first aspect, a method is provided. The method comprises capturing, by a camera, at least one image of an environment. The method comprises selecting, by a reference image module, a reference image, the reference image depicting the environment. The method comprises comparing, by an image comparison module, the reference image with the at least one captured image. The method comprises determining, by the image comparison module, that the reference image is distinct from the at least one captured image, wherein the at least one captured image includes an object that is absent in the reference image. The method comprises cropping, by an image cropping module, the at least one captured image, the cropping comprising cropping out a first portion of the at least one captured image, wherein the remaining portion of the at least one captured image includes the object. The method comprises transmitting, by an image transmission module, the remaining portion of the at least one captured image along with the coordinates of the first portion of the at least one captured image.
[0008] In a second aspect, a method is provided. The method comprises receiving, by an animal detection and recognition module employing a trained animal identification model, an image comprising an animal. The method comprises detecting, by the trained animal identification model, the animal within the image. The method comprises recognizing, by the trained animal identification model, a species of the animal based on a featureextraction of the animal within the image. The method comprises determining, by an animal classification module, if the recognized species of the animal is accurate. On determining that the recognized species of the animal is accurate, the method comprises classifying, by the animal classification module, the recognized species of the animal as dangerous or non-dangerous.
[0009] In a third aspect, a system is provided. The system comprises at least one memory storing a plurality of instructions, and at least one processor, configured to execute the plurality of instructions, to cause the system to perform the method as disclosed in the first aspect.
[0010] In a fourth aspect, a system is provided. The system comprises at least one memory storing a plurality of instructions, and at least one processor, configured to execute the plurality of instructions, to cause the system to perform the method as disclosed in the second aspect.
[0011] The details of the embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims.Brief Description of Drawings:
[0012] The detailed description is described with reference to the accompanying figures. The same numbers are used throughout the drawings to reference like features and components.FIG. 1 illustrates an environment comprising the human animal conflict-mitigation system (HACMS) and other devices, according to an embodiment of the present disclosure;FIG. 2 illustrates the various components within an image capturing and processing unit of the HACMS, according to an embodiment of the present disclosure;FIG. 3 illustrates the various components within an alert control unit of the HACMS, according to an embodiment of the present disclosure;FIG. 4 illustrates the various types of deterrent devices, according to an embodiment of the present disclosure;FIG. 5A illustrates an example reference image and FIG. SB illustrates an example captured image, according to an embodiment of the present disclosure;FIG. 6 illustrates a cropped version of the example captured image, according to an embodiment of the present disclosure;FIG. 7 illustrates a method performed by an image capturing and processing unit, according to an embodiment of the present disclosure;FIG. 8 illustrates a method performed by an alert control unit for animal identification, according to an embodiment of the present disclosure;FIG. 9 illustrates a method performed by the alert control unit for computing a risk score, according to an embodiment of the present disclosure;FIG. 10 illustrates a logic-based decision tree for computing the risk score, according to an embodiment of the present disclosure;FIG. 11 illustrates a method performed by the alert control unit for training an animal identification model, according to an embodiment disclosed hereinDetailed Description
[0013] Exemplary embodiments now will be described with reference to the accompanying drawings. The invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey its scope to those skilled in the art. The terminology used in the detailed description of the particular exemplary embodiments illustrated in the accompanying drawings is not intended to be limiting. In the drawings, like numbers refer to like elements. The term “exemplary embodiment” is meant to be interpreted as being an example embodiment and is not meant to be interpreted as a preferred embodiment.
[0014] The specification may refer to “an”, “one” or “some” embodiment(s) in several locations. This does not necessarily imply that each such reference is to the same embodiment(s), or that the feature only applies to a single embodiment. Single features of different embodiments may also be combined to provide other embodiments.
[0015] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless expressly stated otherwise. It will be further understood that the terms “includes”, “comprises”, “including” and / or “comprising” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, whenever the phrase “at least one of” or “one or more of’ precedes a list of elements, wherein the elements are joined by “and” or “or”, it means that at least any one of the elements or at least all the elements are present. As used herein, whenever the phrase “one of’ precedes a list of elements, wherein the elements are joined by “and” or “or”, it means that only one of the elements are present at a given instant, unless the context permits a meaning that allows the inclusion of more than one element. The usage of the term “or” is to be understood as “inclusive or” instead of “exclusive or”, unless indicated otherwise by the relevant context. Conditional language, such as among others, “can” or “may”, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments could include,while other embodiments may not include certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required for one or more embodiments. It will be understood that when an element is referred to as being “connected” or “coupled” to another element, it can be directly connected or coupled to the other element or intervening elements maybe present. Furthermore, “connected” or “coupled” as used herein may include wirelessly connected or coupled. As used herein, the term “and / or” includes any and all combinations and arrangements of one or more of the associated listed items.
[0016] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0017] The figures depict a simplified structure only showing some elements and functional entities, ail being logical units whose implementation may differ from what is shown. The connections shown are logical connections; the actual physical connections may be different. In addition, all logical units described and depicted in the figures include the software and / or hardware components required for the unit to function. Further, each unit may comprise within itself one or more components, which are implicitly understood. These components may be operatively coupled to each other and be configured to communicate with each other to perform the function of the said unit. For the sake of illustrative purposes, the embodiments herein recite various units / modules that are associated with a particular functionality. However, this is to be construed as non-limiting as the functionality of the various units / modules may be combinable in any manner. Further, the embodiments may include various units / modules not explicitly mentioned herein, but would be implicitly understood to be present by a person skilled in the art.
[0018] The embodiments herein disclose a Human Machine Conflict Mitigation System (HACMS) that can capture an image, which can include an animal, and analyze said image to determine if the animal present in the images is dangerous. If the animal is classified as being dangerous, the HACMS can further analyze the image to generate a risk score, based on which an alert can be generated to notify individuals proximate to the animal about the potential threat of the animal. In addition to this, the HACMS can also activate a set of measures that would then deter the animal from its path towards an area inhabited by humans. The term “animal” here can be construed broadly to also include insects or any other type of species. It is to be noted that the various units / modules whose functionality have been described herein can also be described as functions performed by the HACMS.
[0019] Referring now to the drawings, and more particularly to FIGS. 1 to 11, where similar reference characters denote corresponding features used consistently throughout the figures relating to the example embodiments disclosed herein.OVERVIEW OF THE HACMS
[0020] FIG. 1 illustrates a block diagram depicting the interaction of the HACMS 102 with a user device 110 and deterrent device(s) 112 for mitigating a human-animal conflict, according to an embodiment of the present disclosure. The HACMS 102 may communicate with the devices 110 and 112 via the network 108.
[0021] The HACMS 102 comprises an image capturing and processing unit (ICPU) 104 and an alert control unit (ACU) 106. Although not shown, the HACMS 102 can be implemented by processing circuitry (“processors”) and one or more storage devices (“memory”). The ACU 106 can be a compute resource on a cloud. By way of example, rather than limitation, the HACMS 102 can operate in the capacity of a server. The server may be a physical or virtual server, and the server may be a web server, an application server, or a cloud server.
[0022] The processing circuitry can include, although is not limited to, a general-purpose central processing unit (CPU), an application specific integrated circuit, multiple processingunits, dedicated circuitry for achieving the functionality of various units / modules that will be described later herein.
[0023] The memory can be any suitable processor-readable storage medium, e.g., random access memory (RAM), read-only memory (ROM), Electrical Erasable Read-only Memory (EEPROM), Flash memory, etc., suitable for storing instructions for execution by the processor, and located separate from the processing circuitry' and / or integrated therewith. The memory' can store various instructions that enable the processing circuitry' to execute the functions of the various modules / units that will be described later herein.
[0024] The network 108 can be the Internet, or a private communication link (e.g., LAN or WAN). The network 108 may use standard communication technologies or protocols.
[0025] The user device 1 10 can include any portable or electronic device (e.g., a smart phone, a laptop, tablet etc.) that can communicate with the HACMS 102 via the network 108. The user device 110 can receive electronic alerts (e.g., text message, application-based notifications, or email) from the HACMS 102.
[0026] The deterrent devices 1 12 can include water cannons, spray jets, hooters, loud speakers, high-intensity lights etc. The deterrent devices 112 are activated, by the HACMS 102, to generate physical alerts that can deter an animal from its present trajectory'. For example, if an animal’s trajectory is such that it is moving towards a human habitat, the physical alerts generated by the deterrent devices 112 can incite a fear in the animal, thereby' causing the animal to retreat or divert from its intended path.Overview of image Capturing & Processing Unit ICPU)
[0027] FIG. 2 illustrates the various components of the ICPU 104. The 1CPIJ 104 comprises, at least, a camera 1041 , a reference image module 1042, an image comparison module 1043, an image cropping module 1044; and an image transmission module 1045.
[0028] The camera 1041 can continuously capture images of an environment associated with the camera / surrounding area of the location where the camera is installed, wherein a subsetof the captured images can include one or more animals. The camera 1041 can be powered via electricity or solar energy'. In an embodiment, where the camera 1041 is powered by solar energy, it can be fixed in a remote location.
[0029] The reference image module 1042 stores a plurality of reference images, with each reference image depicting the coverage area (i.e., field of view) of the camera at different time periods / intervals of the day. For example, one reference image can depict an environment, within the coverage area of the camera, during daytime. Another reference image can depict an environment, within the coverage area of the camera, during nighttime. The reference images pertaining to daytime or nightime, can also be associated with an intensity of light or angle of light. Each reference image may also be object-free, i.e., each reference image may be devoid of an animal. In an embodiment, the reference images can include images captured by the camera 1041 on a particular day. In such an embodiment, the reference image(s) can be captured based on a particular frequency (e.g., every hour of the day, every 15 minutes etc.) and accordingly associated with a particular time instant.
[0030] The image comparison module 1043 compares a plurality of the images captured (hereinafter referred to as “captured images”) by the camera 1041, in real-time and within a specific time period, with a reference image that corresponds to the specific time period. For example, if the images are captured in a time range of, for example, 3:00 PM to 4:00 PM, the reference image may also be an image that has been captured within that time range (e.g., on a previous day or same day). In another embodiment, the lighting conditions of the reference image and the at least one captured image may be similar, i.e., the reference image and the at least one captured image depict the environment surrounding the camera under daytime and / or nighttime. The phrase “similar lighting conditions” can also refer to the reference image and the at least one captured image being captured within a particular time window, or being within a range of brightness levels. The objective of the comparison (performed by the image comparison module 1043) is to detect the presence of an animal in the vicinity of the camera, based on the captured image. In other words, the image comparison module 1043 compares the object-less reference image with the captured images, to see if there is any distinction between thereference image and the captured images. The distinction can be atributed to the presence of objects (e.g., an animal or a human) in the captured images. The comparison between the reference image and the captured images can be based on the difference in the pixel value of the images. The comparison between the captured images and the reference image can be performed using metrics such as Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Mean Squared Error (MSE). Through this comparison, it can be determined if there is any distinguishing feature / distinction between the captured images and the reference image, with the distinguishing feature possibly being the animal, human, or any other object. Hence, if therd is a distinguishing feature between the captured image and the reference image (devoid of any animal or human), then it can mean that the captured image includes an animal (or human), and such there is a confirmation that an animal (or human) has been spotted in the vicinity. FIG. 5A illustrates an example reference image, which has no objects (e.g., animals or humans) in the image. FIG. 5B illustrates an example captured image with a tiger (example of an animal). The image comparison module 1043 can compare FIG. 5B with FIG. 5A to determine that FIG. 5B is distinct from FIG. 5A because a tiger (the distinguishing feature in this case) is present in FIG. 5B, and not present in FIG, 5A.
[0031] On a determination, based on the image comparison, that one of the captured images includes, for example, an animal, the image cropping module 1044 may crop the captured image (with the animal) to a portion depicting the animal (i.e., the distinguishing feature). FIG. 6 illustrates a cropped version of FIG. 5B, according to an example embodiment.
[0032] The captured image (e.g., FIG. 5B) can be considered as a matrix, with each pixel in the captured image having a set of coordinates. The cropped image (cropped by the image cropping module 1044), along with the coordinates of the cropped-out portions of the captured image, are transmitted to the ACU 106, by the image transmission module 1045 (e.g., implemented by a transceiver). In one example embodiment, the ICPU 104 and the ACU 106 can be physically located at separate entities / de vices and at different locations, wherein the ICPU 104 may be communicatively coupled to the ACU 106.Overview of the Alert Control Unit
[0033] FIG, 3 illustrates the various components of the alert control unit (ACU) 106, according to an embodiment of the present disclosure. The ACU 106 comprises an image reconstruction module 1061 , an animal detection & recognition module 1062, an animal classification module 1063, a score computation module 1064, and a device triggering module 1066.
[0034] As previously stated, the cropped image, with the coordinates of the cropped-out portions, is transmitted to the ACU 106. By transmitting the coordinates of the cropped-out portions to the ACU 106, the image reconstruction module 1061 can reconstruct the entirety of the captured image from just the cropped image. In the event that an alert needs to be triggered to the user device 110, the alert can, in an example embodiment, include the reconstructed captured image with the animal so that a user can, for example, discern details such as the species of the animal that is in the vicinity' and its last known location.
[0035] The reconstructed image or the cropped image can be fed as input to the animal detection & recognition module 1062. The animal detection & recognition module 1062, may at least be implemented by one or more machine learning models. There may be a first machine learning model that can detect the animal in the transmitted image. The first machine learning model may utilize any computer vision technique for detection of the animal, using, for example, bounding boxes. There may be a second machine learning model which can be used for recognizing the species of the animal in the image. In an example embodiment, the functionalities of the first machine learning model and the second machine learning model may be combined into one model or collectively referred to as “animal identification model”.
[0036] The training of the animal identification model can be based on a dataset of labelled animal images, wherein the images capture the animals in a daytime setting, dim setting or low-light setting (e.g., dawn or nightime). In some embodiments, the training dataset may undergo a form of preprocessing, for example, augmentation techniques like hue and saturation adjustments, to improve the animal identification model’s performance (e.g., detecting and recognizing animals in different light settings). In one or moreembodiments, the preprocessing of the training dataset may also include bounding box brightness adjustments. The animal identification model may, without limitation, include a convolutional neural network (CNN), that can be trained to extract various features from the captured image, and accordingly identify (i.e., detect and recognize) the animal in the captured Image.
[0037] The animal identification process can be a two step-process; The first step can involve detecting an animal. The second step can involve recognizing the species of the animal. By way of example, rather than limitation, the species of the animal can include tigers, leopards, lions, fly etc. The first step of animal detection can involve generating bounding boxes around the animal. The bounding boxes highlight the portion of the input image that corresponds to the detected animal.
[0038] After detecting the animal, the second step of the identification process, i.e., recognizing the species of the animal, can be executed. The animal identification model can be trained on a dataset comprising images of animals belonging to a plurality of species, wherein each image can have a label corresponding to the species of the animal in the image. Based on the training of the animal identification model, it can generate a probabilityscore for various labels (i.e., multi-label classification) that could represent the species of animal in the picture. The recognized animal (i.e., species of the animal) can be the animal whose label has the highest probability score. In another embodiment, the animal identification model can also be trained to output an individual of the recognized animal. For instance, tigers can be differentiated based on, for example, the stripes on the body of the tiger. Similarly, leopards can be differentiated based on, for example, the spots on their body. Accordingly, the animal identification model can also determine the individual tiger or leopard, as part of the recognition process. In another embodiment, the animal identification model can include, for example, a plurality of binary' decision trees, with each free corresponding to a specific species and / or species individual label of an animal.
[0039] The training of the animal identification model, for example, the CNN, can comprise selecting hyperparameters such as epoch number, batch size, dropout, image size, classification loss weight etc. An activation function that can be used is the leaky ReLU(rectified linear unit) activation function, with skip connections in the cross stage partial (CSP) layers. The training can also comprise freezing of certain layers (to prevent further learning on the training dataset) and unfreezing of certain layers (to continue with the learning of the training dataset). In case the input image to the animal identification model includes an object, other than an animal, the animal identification model can output that object in the input image is not an animal, and accordingly, further processing is not required.
[0040] The recognized animal (e.g., tiger, leopard etc.) can be predetermined to be either dangerous or harmless. Therefore, based on the recognized animal, the animal classification module 1063 classifies the animal as dangerous or harmless, wherein if the animal is classified as dangerous, then further processing / analysis can occur to assess the level of danger that the recognized animal poses to a human.
[0041] In an embodiment, before the classification of the recognized animal can occur, there may be a confirmation step (i.e., a confidence threshold step) as to whether the recognized animal (i.e., output of the animal identification model) is a true positive (i.e., accurate). Essentially, there may be situations where the detected animal in the image can end up being incorrectly recognized as a tiger, a leopard, a lion, or any other species of dangerous animals. Consequently, this incorrect recognition could also result in an incorrect classification of the detected animal, wherein the detected animal would be incorrectly classified as being dangerous. An example situation of where such an incorrect classification could occur is when a fly (i.e., detected animal) that is close to the camera, can have a large bounding box size, and it could be recognized as a tiger or a leopard (examples of dangerous animals).
[0042] Therefore, to eliminate false positives, a filtering step can be performed by the animal detection and recognition module 1062, wherein the bounding box size (e.g., height and width of the box) of the detected animal can be compared with a threshold size of the bounding box corresponding to the recognized animal. Continuing w'ith the previous example, assume that the image fed to the animal identification model comprises a fly, wherein the animal identification model detects the fly using the bounding boxes, butincorrectly recognizes the fly as a leopard (an example of a dangerous animal). Then, to confinn whether this recognition is a true positive or a false positive, the bounding box size of the fly (the detected animal) can be compared with the threshold size of the bounding box for the leopard (the recognized animal). If the bounding box size of the fly exceeds the threshold bounding box size of the tiger or leopard, then this can be considered as a false positive.[004.3] In the case of a false positive, the animal identification model filters out these detections, However, if it is determined based on the bounding box size comparison that the recognized animal is a true positive, then the animal classification module 1063 classifies the animal as being dangerous or harmless. In the case of the animal classification being “harmless”, further processing may not be required. But, in the case of the animal classification being “dangerous”, further processing may occur by the score computation module 1064 to assess the level of danger that the recognized animal would pose to a human. In other words, the score computation module 1064 generates a risk score that is representative of the recognized animal’s inclination to attack a human or create a conflict.
[0044] The score computation module 1064 may use a logic-based decision tree (as shown in FIG. 10) to compute / generate the risk score, wherein based on the risk score, it is determined if electronic and / or physical alerts need to be generated for alerting humans about the nearby presence of the dangerous animal. The device triggering module 1066 can determine if the risk score exceeds a threshold, and accordingly trigger the deterrent devices 112.
[0045] FIG. 10 illustrates a logic-based decision tree where the branches represent a non- exhaustive criteria (physical characteristics of an animal and / or geographical information) for gauging an animal’s inclination to attack or create a conflict. The branches of the logic-based decision tree can consider at least one of the following for its output:● Age of the animal● Physical damage to the animal● Chipped tooth● Tiger movement● Herbivore density
[0046] Age of the animal: The age of the animal can provide insight on the agility of the animal for attacking its prey, where the younger the animal, the more agile it is. On the other hand, an older animal may not be fast enough to catch its prey, thereby making it less inclined to atack a human. As previously stated, the animal identification model can be trained to classify animals based on a labelled dataset. The labelled dataset may also include the age of the animals, thereby enabling the animal identification model to estimate the age of the detected animal. This way, the age of the recognized animal can be determined and fed to the logic-based decision tree. In an embodiment, a branch of the logic-based decision tree can have a threshold age, wherein animals whose age is below the threshold age can be determined to be agile, and as such likelier to be threatening to a human. Similarly, those animals whose age is higher than the threshold age can be determined to be slow or sluggish, thereby making them less likely to attack a human,
[0047] Physical damage: The physical damage to the animal, like a laceration in the body due to a fight or damaged limbs, can cause a problem for the animal due to which it is prone to creating a conflict. In an embodiment, a branch of the logic-based decision free can represent the presence of physical damage to the animal in binary values (0 or 1), with ‘0’ indicating that there is no physical damage and ‘1’ indicating that there is physical damage. The physical damage to the animal can be determined by the animal identification model through more labelling in the training dataset.
[0048] Chipped Tooth: The animal’s teeth are very essential for hunting its prey, where without canines it cannot hunt. Sometimes, the animal’s tooth may be chipped while engaging with a prey, due to which depending on extent of the chipping of the tooth, it can be determined if the animal is susceptible to conflict with humans. In an embodiment, a branch of the logic-based decision tree can represent the presence of chipped tooth in binary values (0 or 1), with ‘0’ indicating that there is no chipped tooth and ‘ 1 ’ indicatingthat there is chipped tooth. The animal identification model can be trained to determine whether any tooth of the animal is chipped.
[0049] Animal movement: If an animal (e.g., a tiger) is on a continuous movement, then it is likely that it has lost its territory. As such, the animal may be more inclined to attract conflict. In an embodiment where multiple cameras may be installed at different locations, then by comparing the distance between the two locations whose cameras have captured the same animal and the time instants at which each camera captured the animal, it can be determined the amount of distance travelled by the animal and accordingly whether the animal has been on a continuous movement. In an example embodiment, a branch of the logic-based decision tree can represent a threshold distance travelled by an animal, where if the animal has travelled more than the threshold distance, the output of the branch is ‘ 1 ’, else it is ‘O’. In one example embodiment, where the animal is a tiger, the threshold distance can be 100 km. This value can be manually input to the score computation module 1064.
[0050] Herbivore density: If the area where the tiger has been detected has a herbivore density below a threshold, then it can mean that due to the inadequate supply of food (based on the herbivore density), the tiger may be prone to attacking humans for its food. The herbivore density may be manually input to the score computation module 1064.
[0051] Accordingly, based on the logic-based decision tree, a risk score is calculated / computed. If the risk score exceeds a certain threshold, then the device triggering module 1066 transmits a signal to the user device 1 10 and / or the deterrent device 112 to generate the corresponding alerts. Although the disclosure herein explains how in one embodiment the animal’s inclination to atack can be determined using a logic-based decision tree, it is to be noted that other embodiments can utilize machine learning models (including deep learning models) for the same purpose.
[0052] The alerts sent to the user device 110 can be used for alerting various individuals of the nearby presence of a dangerous animal, and to take the necessary precautions. The alerts sent to the deterrent device 112, can result in the activation of, for example, spray jets,water cannons, hooters, and / or high-intensity Sights, for deterring a dangerous animal from engaging in a conflict. As previously stated, the alerts can be transmited to the user device 110 in the form of a text message, an email notification, or a notification from an application installed in the user device 110. Additionally, as the ACU 106 receives a cropped version of the captured image, the image reconstruction module 1061 can reconstruct the captured image in its entirety using the coordinates of the portions of captured image that were cropped out. As part of the transmitted alert, the reconstructed captured image can be sent to the user device 1 10 to indicate to the user of the user device 110 about the exact whereabouts of the recognized animal and also the species of the recognized animal.
[0053] FIG. 7 illustrates a method 700 performed by the ICPU 104, according to an embodiment of the present disclosure. At step 702, the ICPU 104 can select a reference image, wherein the reference image is to be used as a basis for comparison against a set of captured images. The reference image can be an image captured by a camera of the ICPU 104. The reference image can depict an environment in which the camera is located and can be devoid of any objects. At step 704, the ICPU 104 can capture at least one image of the surrounding environment (i.e., an environment within the field of view of the camera of the ICPU 104), which is to be compared with the reference image. At step 706, the ICPU 104 can compare the reference image with the at least one captured image, At step 708, the comparison can lead to the ICPU 104 determining that there is a distinguishing feature ' between the reference image and the at least one captured image. In other words, the ICPU 104 determines that an object is present in the at least one captured image, as the presence of the object can be the distinguishing feature between the reference image and the at least one captured image. The object can, for example, be a human or a person. At step 710, the ICPU 104 can crop the at least captured image with the object, wherein the cropped image includes the object. The cropping can involve cropping out a first portion, that partially or wholly does not include the object, and the remainder of the picture (i.e., non-cropped portion) can include the object. The at least one captured image can be represented as a matrix, with each pixel in the at least one captured image having a set of coordinates. At step 712, the cropped image and the coordinates of the cropped-outportion (i.e., first portion) of the at least one captured image can be transmitted to, for example, the ACU 106.
[0054] FIG. 8 illustrates a method 800 performed by an alert control unit (ACU) 106, according to an embodiment disclosed herein. More particularly, the method 800 relates to the performing of animal identification by the ACU 106. At step 802, the ACU 106 can receive an image comprising an animal. In an embodiment, the received image can be the cropped image transmitted by the ICPU 104. At step 804, the ACU 106 can detect the presence of the animal in the received image using, for example, bounding boxes. The ACU 106 may comprise an animal identification model that is trained to perform the object (e.g., animal) detection using bounding boxes.
[0055] Upon detecting the animal, at step 806, the ACU 106 can recognize the animal in the received image. In one embodiment, the recognition process can involve the animal identification model performing multilabel classification, with each label having a certain probability value and the recognized animal corresponding to the label having the highest probability value. The recognition process can also comprise recognizing an individual of the recognized species of the animal.
[0056] At step 808, the ACU 106 can check whether the recognized animal (i.e., the recognized species of the animal or an individual of the recognized species of the animal) is a true positive or a false positive. In one context, the phrase “true positive” can be understood as whether the animal identification model correctly recognizes the animal in the received image, irrespective of whether the animal is dangerous or harmless (non-dangerous). In another context, the phrase “true positive” can be understood as whether the recognized animal, which may be deemed to be dangerous, has been correctly identified by the animal identification model. For example, a tiger may be deemed to be a dangerous animal, and if the animal identification model recognizes the animal in the received image to be a tiger, then a true positive is when the animal in the received image is actually a tiger. Similarly, a false positive (i.e., not “true positive”) is when the animal in the received image is actually a harmless animal (e.g., a fly). The determination of whether the recognized animal is a true positive can be based on a comparison of the bounding boxsize of the detected animal in the received image and a threshold bounding box size of the recognized animal.
[0057] If the recognized animal is a true positive, then at step 812, the animal can be classified as dangerous or non-dangerous / harmless. In one or more embodiments, the classification of the animal may be pre-determined, i.e., some animals (e.g., rabbits) are always considered as harmless and some animals are always considered as dangerous. Otherwise, if the recognized animal is a false positive, then the ACU 106 can discard the analysis of the model or refrains from engaging in further processing.
[0058] FIG. 11 illustrates a method 1100 performed by the ACU 106 for training the animal identification model, according to an embodiment of the present disclosure. The training dataset can comprise a diverse set of images (i.e., not including similar images), where class balance is maintained to prevent bias towards any specific category. In other words, if there are 4 classes / labels (e.g., tiger, fly, leopard, bear), then the training dataset can comprise an equal number of images for each of those classes / labels so that the animal identification model is uniformly trained to recognize each of those classes / labels. At step 1102, the ACU 106 can preprocess the training dataset using data augmentation techniques, such as hue and saturation, so that the animal identification model is trained to detect and recognize animals in different lighting conditions, while emphasizing their distinct features. At step 1104, the training dataset can also be annotated with bounding boxes for animal detection and may include labels associated with the animals present in the training dataset. At steps 1106 and 1198, after feeding the training dataset to the animal identification model, it is trained to detect an animal and recognize the animal, respectively.
[0059] FIG. 9 illustrates a method performed by the ACU 106 for computing a risk score associated with the recognized animal, according to an embodiment disclosed herein. At step 902, the ACU 106 can classify the recognized animal as “dangerous” or “harmless.” There may be a table or a database that maps each animal (its species and / or individual of species) or output label (from the animal identification model) to a class (e.g.,dangerous or harmless). In other words, the classification of an animal can be pre-defined. At step 904, the ACU 106 can classify the recognized animal as being dangerous, based on the mapping. At step 906, the ACU 106 can classify the recognized animal as being harmless. At step 908, which follows from step 906, the ACU 106 can compute a risk score associated with the recognized animal. The risk score can indicate the animal’s inclination to attack or a threat level. In one embodiment, the risk score can be computed based on the logic-based decision tree as shown in FIG. 10. If the risk score is above a threshold, then at step 912, the ACU 106 can enable the HACMS 102 to generate alerts that users need to be vigilant about the nearby presence of a dangerous animal. However, if the risk score is below a threshold or if the animal was classified as being harmless, then at step 914, no alert is generated and the ACU 106 may refrain from performing further processing.Technical Advantages Effects
[0060] The following is a non-exhaustive list of technical advantages / effects achieved by the one or more embodiments disclosed herein. At the ICPU 104, by comparing the captured image(s) with the reference image, and as a result, only transmitting those captured images which have an object in it, there can be a saving of resources that would unnecessarily be consumed by the alert control unit (ACU) 106 for analyzing a captured image that is devoid of an object. Further, the captured image (with the animal) that is transmitted to the ACU 106 can be a cropped image, wherein the captured image is cropped to just the portion of the image that comprises the distinguishing feature between the captured image and the reference image. This way, the technical advantage here is the reduction in the bandwidth requirements, as the quantity of images that are transmitted is reduced, and the size of the transmited images (i.e., cropped images) are smaller than the size of the captured images. The reduction in the bandwidth requirements is further desirable, because one or more embodiments of the HACMS 102 may be deployed in remote areas, where the network speed might be limited. The dataset for training the animal identification model can include a diverse set of images while maintaining classbalance, thereby reducing the bias of the animal identification model and improving its decision-making ability. At the ACU 106 level, by checking if the output of the animal identification model is a false positive, then the unnecessary processing power utilized by the score computation module 1064 (for computing the score) can be eliminated as the computation of the score is not required. For example, if the animal recognition model classifies the detected animal as a leopard (i.e., a false positive), then by comparing the bounding box of the detected animal to see if it exceeds a threshold relative to the size of the detected animal, it can be determined that this is a false positive (i.e., the detected animal could be a non-dangerous animal, like a fly), at which point no further processing needs to be performed. In some embodiments, the training of the animal recognition model can involve freezing a set of layers while leaving the remaining layers unfrozen. One purpose for this is to limit further learning by the frozen layers. Hence, by freezing a set of the layers (i.e., not training these layers), resources that would have otherwise been consumed by the frozen layers are also avoided.
[0061] In the drawings and specification, there have been disclosed exemplary' embodiments of the invention. Although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation. It will be apparent to those having ordinary skill in this art that various modifications and variations may be made to the embodiments disclosed herein, consistent with the present invention, without departing from the spirit and scope of the present invention. Other embodiments consistent with the present invention will become apparent from consideration of the specification and the practice of the description disclosed herein.
Claims
We Claim:
1. A method (700), comprising: capturing (702), by a camera (1041), at least one image of an environment; selecting (704), by a reference image module (1042), a reference image, the reference image depicting the environment; comparing (706), by an image comparison module (1043), the reference image with the at least one captured image determining (708), by the image comparison module (1043), that the reference image is distinct from the at least one captured image, wherein the at least one captured image includes an object that is absent in the reference image; cropping (710), by an image cropping module (1044), the at least one captured image, the cropping comprising: cropping out a first portion of the at least one captured image, wherein the remaining portion of the at least one captured image includes the object: and transmitting (712), by an image transmission module (1045), the remaining portion of the at least one captured image along with the coordinates of the first portion of the at least one captured image.
2. The method as claimed in claim 1, wherein the image comparison module determines that the reference image is distinct from the at least one captured image using at least one of: structural similarity index (SS1M); peak signal-to-noise ratio (PSNR); and means squared error (MSE).
3. The method as claimed in claim 1, wherein the reference image and the at least one captured image depict the environment under similar lighting conditions.
4. The method as claimed in claim 1, wherein the coordinates of the first portion can be used for reconstructing the entirety of the at least one captured image from the remaining portion of the at least one captured image.
5. The method as claimed in claim 1, wherein the object is an animal or a human.
6. A method (800 / 900), comprising: receiving (802), by an animal detection and recognition module employing a trained animal identification model, an image comprising an animal; detecting (804), by the trained animal identification model, the animal within the image; recognizing (806), by the trained animal identification model, a species of the animal based on a feature extraction of the animal within the image; determining (808), by an animal classification module, if the recognized species of the animal is accurate; and on determining that the recognized species of the animal is accurate, classifying (812 / 902), by the animal classification module, the recognized species of the animal as dangerous or non- dangerous.
7. The method as claimed in claim 6, wherein on classifying the recognized species of the animal as dangerous, the method comprises: computing (908), by a score computation module employing a machine learning model, a risk score, wherein the computation of the risk score is based on at least one of: an age of the animal in the image; physical damage to the animal; presence of chipped tooth of the animal; a trajectory of the animal; and a herbivore density of a location of the animal; and determining (910), by a device triggering module, if the computed risk score exceeds a threshold.
8. The method as claimed in claim 7, wherein on determining that the computed risk score exceeds a threshold, the method comprises: triggering (912), by the device triggering module, at least one of: electronic-based alerts; or physical alerts.
9. The method as claimed in claim 6, wherein the trained animal identification model is trained to detect and recognize one or more animals in different lighting conditions based on a preprocessed training dataset depicting one or more animals in the different lighting conditions.
10. The method as claimed in claim 8, wherein the trained animal identification model includes a convolutional neural network, and wherein the machine learning model employed by the score computation module is a logic -based decision tree.1 1. The method as claimed in claim 6, wherein: the detecting of the animal within the image is done using bounding boxes; and the determining if the recognized species is accurate is based on a comparison of the bounding box size of the animal in the image and a bounding box size associated with the recognized species of the animal.
12. The method as claimed in claim 6, wherein: recognizing the species of the animal based on the feature extraction, comprises recognizing an individual of the species of the animal; determining if the recognized species of the animal is accurate, comprises determining if the individual of the recognized species of the animal is accurate; and classifying the recognized species of the animal as dangerous or non-dangerous, comprises classifying the individual of the recognized species of the animal as dangerous or non-dangerous.
13. A system (102), comprising: at least one memory storing a plurality of instructions; and at least one processor, configured to execute the plurality' of instructions, to cause the system to perform the methods as claimed in any of methods 1 to 5.
14. A system (102), comprising:at least one memory storing a plurality of instructions; and at least one processor, configured to execute the plurality of instructions, to cause the system to perform the methods as claimed in any of methods 6 to 12.
Citation Information
Patent Citations
Distance Metric for Image Comparison
US20140169684A1
Image cropping and synthesizing method, and imaging apparatus
US7209149B2
Animal visual identification, tracking, monitoring and assessment systems and methods thereof
WO2022077113A1