System and method for animal detection

By using animal detection systems and methods in farms to generate probability maps and affinity field maps of key points, the problem of animal counting difficulties in farms has been solved, achieving high-precision, real-time animal detection and reducing manual labor costs.

CN116964646BActive Publication Date: 2026-02-10PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280011995.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-28
Filing Date
2022-04-29
Publication Date
2026-02-10
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

In farms, existing technologies struggle to quickly and accurately count large numbers of crowded animals, leading to increased manual labor costs.

Method used

An animal detection system and method are employed to detect animals in regions of interest by receiving images and generating probability maps and affinity field maps of key points using an animal detection model, determining the connectivity graph.

Benefits of technology

It achieves high-precision, real-time or near-real-time animal detection, reducing the possibility of omissions and false detections, and reducing the need for manual counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116964646B_ABST
    Figure CN116964646B_ABST
Patent Text Reader

Abstract

A system (101) and method for detecting animals in a region of interest receives an image (202) capturing a scene in a region of interest. The image (202) is fed to an animal detection model (250) to produce a set of probability maps (208) of a set of keypoints and a set of affinity field maps (206) of a set of keypoints. One or more connection graphs (216) are determined based on the set of probability maps (208) and the set of affinity field maps (206). Each connection graph (216) outlines a presence of an animal in the image (202). One or more animals present in the region of interest are detected based on the one or more connection graphs (216).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Patent Application No. 17 / 361,258, filed June 28, 2021, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This invention relates to animal detection technology, and more specifically, to a method, electronic device, and storage medium for detecting animals in a region of interest. Background Technology

[0004] The growing demand for meat has fueled the continued growth of the livestock industry. In feedlots, hundreds or even thousands of animals may be raised simultaneously in different pens. For example, a pig farm may have dozens or hundreds of pens, each housing dozens of pigs living together. Manually counting the pigs in a pig farm would require a significant amount of manual labor. Furthermore, if some pigs occasionally bump into the wrong pens, it can make manual counting even more difficult. Therefore, the operating costs of feedlots can increase due to the rising labor costs in these facilities. Summary of the Invention

[0005] In one aspect, this invention discloses a method for detecting animals in a region of interest. The method involves receiving an image capturing a scene within the region of interest. The image is fed into an animal detection model to generate a set of probability maps for a set of keypoints and a set of affinity field maps for a set of keypoints. One or more connectivity graphs are determined based on the set of probability maps and the set of affinity field maps. Each connectivity graph outlines the presence of an animal in the image. One or more animals present in the region of interest are detected based on the one or more connectivity graphs.

[0006] In another aspect, the present invention discloses a system for detecting animals in a region of interest. The system includes a memory and a processor. The memory is configured to store instructions. The processor is coupled to the memory and configured to execute the instructions to perform a process including: receiving an image capturing a scene in the region of interest; feeding the image to an animal detection model to generate a set of probability maps of a set of keypoints and a set of affinity field maps of a set of keypoints; determining one or more connection graphs based on the set of probability maps and the set of affinity field maps, wherein each connection graph outlines the presence of an animal in the image; and detecting one or more animals present in the region of interest based on the one or more connection graphs.

[0007] In another aspect, the present invention discloses a non-transient computer-readable storage medium. The computer-readable storage medium is configured to store instructions that, in response to execution by a processor, cause the processor to perform a process comprising: receiving an image of a scene in a region of interest; feeding the image to an animal detection model to generate a set of probability maps of a set of keypoints and a set of affinity field maps of a set of keypoints; determining one or more connection graphs based on the set of probability maps and the set of affinity field maps, wherein each connection graph outlines the presence of an animal in the image; and detecting one or more animals present in the region of interest based on the one or more connection graphs.

[0008] It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory, and do not limit the claimed invention. Attached Figure Description

[0009] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate implementations of the invention and, together with the specification, further serve to explain the invention and enable those skilled in the art to make and use it.

[0010] Figure 1A A block diagram of an exemplary operating environment of a system disclosed in this invention configured to detect an animal in a region of interest is shown.

[0011] Figure 1B A block diagram of another exemplary operating environment of the system disclosed in this invention, configured to detect an animal in a region of interest, is shown.

[0012] Figure 2A An exemplary procedure for detecting an animal in a region of interest, as disclosed in this invention, is shown.

[0013] Figure 2B Another exemplary procedure for detecting an animal in a region of interest, as disclosed in this invention, is shown.

[0014] Figure 2C A schematic diagram of an exemplary structure of the animal detection model disclosed in this invention is shown.

[0015] Figure 3A A schematic diagram of an exemplary structure of a feature extraction model according to an embodiment of the present invention is shown.

[0016] Figure 3B A schematic diagram of an exemplary structure of the depth-separable convolutional block disclosed in this invention is shown.

[0017] Figure 3C A schematic diagram of another exemplary structure of a depthwise separable convolutional block in the feature extraction model disclosed in this invention is shown.

[0018] Figure 3D An exemplary table is shown, which includes a list of parameters for a list of convolutional sequences as disclosed in this invention.

[0019] Figure 4A A schematic diagram of an exemplary structure of the two-level detection model disclosed in this invention is shown.

[0020] Figure 4B A schematic diagram of an exemplary structure of the neural network in the two-level detection model disclosed in this invention is shown.

[0021] Figure 4C A schematic diagram of an exemplary structure of a convolutional block in a neural network disclosed in this invention is shown.

[0022] Figure 5 This is a flowchart of an exemplary method for detecting animals in a region of interest, as disclosed in this invention.

[0023] Figure 6 This is a flowchart of an exemplary method for determining one or more connection graphs disclosed in this invention.

[0024] Figure 7 This is a flowchart of an exemplary method for forming a cluster of fragments for a set of key points, as disclosed in this invention.

[0025] Figure 8A -8B is a graphical representation of a set of key points marked in an animal, as disclosed by an exemplary embodiment of the present invention.

[0026] Figure 9A This is a graphic representation of an exemplary image of an animal enclosure disclosed in this invention.

[0027] Figure 9B This is a graphical representation of an exemplary location diagram of key shoulder points disclosed in this invention.

[0028] Figure 9C This is a graphical representation of an exemplary location diagram of key points on the buttocks disclosed in this invention.

[0029] Figure 10A This is a graphical representation of an exemplary affinity field diagram of the key point set disclosed in this invention.

[0030] Figure 10B This invention discloses a method based on... Figure 10A A graphical representation of an exemplary process for generating fragment clusters from affinity field diagrams.

[0031] Figure 11A This is a graphic representation of an exemplary image of an animal enclosure disclosed in this invention.

[0032] Figure 11B This is the basis for the disclosure of this invention. Figure 11A A graphical representation of an exemplary affinity field map of the keypoint set generated from an image.

[0033] Figure 12 This is a graphical representation of an exemplary connection chart depicting an animal in an image, as disclosed in this invention.

[0034] The implementation of the present invention will be described with reference to the accompanying drawings. Detailed Implementation

[0035] Reference will now be made in detail to exemplary embodiments, examples of which are shown in the accompanying drawings. Wherever possible, the same reference numerals are used in all the drawings to denote the same or similar parts.

[0036] Computer vision technology can be applied to animal husbandry to monitor animals in farms. For example, cameras can be installed in farms to monitor animals kept in different enclosures. Images or videos captured by the cameras can be processed to detect animals in different enclosures and automatically calculate the total number of animals in the farm. As a result, labor costs in farms can be reduced.

[0037] Animals raised in farms often exhibit a variety of habits, such as standing together in groups, lying back-to-back on the floor, resting together in corners, overlapping each other, or crowding together along feeding troughs. As a result, one or more parts of an animal can easily be obscured by another animal, making it difficult to detect from images (or video frames).

[0038] In some examples, single-level network models (e.g., YOLO, RetinaNet) can be used to detect objects in images using non-maximum suppression algorithms. This single-level network model can be used to detect animals in a farm enclosure. However, because animals in enclosures (e.g., pigs) may have a habit of huddling together to rest or feeding, images capturing the scene of an enclosure can depict animals stacked together or overlapping each other. When non-maximum suppression is applied to detect animals from the image, multiple overlapping animals can be identified as a single animal, resulting in missed detections of animals in the enclosure. As a result, the single-level network model may fail to detect animals in the enclosure, especially when the animals are crowded together.

[0039] In some examples, two-level network models (e.g., Faster-RCNN, Mask-RCNN) can be used to detect objects in images. However, due to the structural complexity of two-level network models, their detection speed is slow. Two-level network models cannot be used for real-time or near-real-time object detection. Therefore, applying a two-level network model to detect animals in a fenced area results in significant detection latency.

[0040] In some examples, pose detection models based on keypoint determination can be used to detect the pose of objects in an image. When using a pose detection model to detect animals in an enclosure, a regression method based on coordinate difference prediction can be used to establish relationships between keypoints, allowing keypoints to be matched with different animals. However, when animals are crowded together in the enclosure, the regression method may not match keypoints, causing pose detection to fail. Consequently, the pose detection model may also fail to detect animals in the enclosure, especially when animals are crowded together.

[0041] This invention provides an animal detection system and method capable of detecting animals in a region of interest (ROI), even when the animals are crowded together within the ROI. Specifically, a camera module is used to acquire images capturing the scene within the ROI. By applying an animal detection model, the animal detection system and method described herein can determine one or more connection graphs of one or more animals appearing in an image. Each connection graph can outline the presence of an animal in the image. The animal detection system and method can detect one or more animals present in the ROI based on one or more connection graphs.

[0042] For example, an image can be fed into an animal detection model to generate a set of probability maps for a set of keypoints and a set of affinity field maps for a set of keypoints. One or more connection graphs can then be determined based on this set of probability maps and the set of affinity field maps. The detection results can then be determined using one or more connection graphs, including but not limited to the total number of animals present in the region of interest, one or more geographic locations of one or more animals in the region of interest, and one or more poses of one or more animals.

[0043] The animal detection systems and methods described herein can provide detection results of regions of interest in real-time or near real-time with high accuracy. The possibility that one or more animals may be missed in the detection results (e.g., the possibility of missed detection) and the possibility that one or more animals may be falsely detected (e.g., the possibility of one or more animals not matching) can both be reduced.

[0044] For example, the animal detection systems and methods described herein can redefine a set of keypoints for each animal. By using a set of keypoints, the likelihood of missed detections can be reduced even when animals are closely clustered together. Furthermore, the possibility of mismatches between keypoints of different animals can be reduced.

[0045] In another example, the animal detection system and method described in this paper apply the affinity field of a keypoint set to measure the different degrees of association between different locations of keypoints in the keypoint set, and match different locations of keypoints with different animals based on the degree of association. This matching method is more stable than regression methods based on coordinate difference prediction and can significantly reduce keypoint mismatches.

[0046] In yet another example, the animal detection model described herein may include a feature extraction model. Feature extraction models can reduce the number of parameters in a neural network model while maintaining high accuracy in the detection results. For example, when the animal detection model is only 11MB in size, it can achieve an accuracy of 95.4% in the detection results.

[0047] In another example, the animal detection model described herein can be trained using multiple training images, which capture animals with multiple body shapes and poses in multiple living environments. Multiple training images can be captured at multiple times with different lighting conditions. Therefore, the diversity and robustness of the animal detection model can be improved by training with multiple training images.

[0048] According to the present invention, the term "near real-time" can refer to data processing that provides a rapid response to events with slight delays. The delay can be in milliseconds (ms), seconds, minutes, etc., depending on various factors such as computing power, available memory space, and signal sampling rate. For example, the animal detection system and method described herein can be implemented almost in real-time with a millisecond delay.

[0049] According to the present invention, the animals described herein can be, for example, livestock raised in farms (e.g., pigs, sheep, cattle, horses, chickens, geese, or ducks), animals kept in zoos (e.g., tigers, lions), animals living in national parks, etc. As an example, the following description will refer to animals raised in farms. It should be understood that this specification can also be applied to any other type of animal.

[0050] According to the present invention, key points may represent a part or joint of an animal, such as the head, shoulder, abdomen, hip, elbow, foot, etc. In some embodiments, a set of key points of an animal may include one or more body key points located within the animal's body, one or more limb key points located in one or more limbs of the animal, or combinations thereof. For example, one or more body key points may include head key points, shoulder key points, abdominal key points, hip key points (or tail key points), or combinations thereof. The one or more limb key points may include the elbow joint key point of the left foreleg, the left forefoot key point, the elbow joint key point of the right foreleg, the right forefoot key point, the elbow joint key point of the left hind leg, the left hind foot key point, the elbow joint key point of the right hind leg, and the right hind foot key point. See below. Figures 8A-8BExplain key points about pigs.

[0051] According to the present invention, a set of keypoints may include two or more keypoints forming an animal segment. The animal segment may be an animal fragment associated with two or more keypoints, such as a body segment or limb of an animal. For example, a set of keypoints may include head keypoints and shoulder keypoints, and the connection between the head keypoints and shoulder keypoints may represent the neck portion of the animal. In another example, the set of keypoints may include head keypoints, shoulder keypoints, and hip keypoints. A first connection from the head keypoint to the shoulder keypoint and a second connection from the shoulder keypoint to the hip keypoint may be combined to represent an animal body segment (e.g., the torso of the animal) from the head to the hip.

[0052] Figure 1A A block diagram of an exemplary operating environment 100 for a system 101 disclosed in this invention, configured to detect animals in a region of interest, is shown. The region of interest can be, for example, an animal enclosure, a cage, or any other area of ​​interest within a farm. Operating environment 100 may include system 101, computing device 112, user device 114, camera module 116, and any other suitable components. Components of operating environment 100 may be coupled to each other via network 110.

[0053] In some embodiments, system 101 may be contained within a cloud computing device. Alternatively, system 101 may be contained within a local computing device. The computing device may be, for example, a server, desktop computer, laptop computer, tablet computer, or any other suitable electronic device including a processor and memory. In some embodiments, system 101 may include a processor 102, memory 103, and storage device 104. It should be understood that system 101 may also include any other suitable components for performing the functions described herein.

[0054] For example, system 101 may have different components in a single device, such as an integrated circuit (IC) chip, or a separate device with a dedicated function. The IC may be implemented as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). In another example, one or more components of system 101 may be located in a cloud computing environment, or alternatively in a single location or distributed locations, but communicate with each other via network 110.

[0055] It should be understood that when system 101 is implemented in a cloud computing environment, the farm may need to provide network connectivity with specific requirements (e.g., specific bandwidth) so that camera module 116 and computing device 112 in the farm can communicate with system 101 via network 110. However, system 101 can also be implemented in a local computing environment, such as... Figure 1BAs shown. In this case, no network connection is required in the farm, which is suitable for applications in remote areas where network connectivity is unavailable.

[0056] Processor 102 may include any suitable type of general-purpose or special-purpose microprocessor, digital signal processor, microcontroller, or graphics processing unit (GPU). Processor 102 may include one or more hardware units (e.g., one or more portions of an integrated circuit) designed to be used in conjunction with other components or to execute part of a program. This program may be stored on a computer-readable medium and, when executed by processor 102, may perform one or more functions. Processor 102 may be configured as a separate processor module dedicated to animal detection. Alternatively, processor 102 may be configured as a shared processor module for performing other functions.

[0057] The processor 102 may include several modules, such as a detection module 105, an analysis module 106, and a training module 107. Although in Figure 1A The diagram shows a detection module 105, an analysis module 106, and a training module 107 within a single processor 102. These modules may also be implemented on different processors located close to or far from each other. For example, the training module 107 could be implemented by a processor (e.g., a GPU) dedicated to offline training of the animal detection model, and the detection module 105 could be implemented by another processor for detecting animals in regions of interest from an image.

[0058] The detection module 105, analysis module 106, and training module 107 (and any corresponding submodules or subunits) may be hardware units (e.g., portions of an integrated circuit) of the processor 102, designed to be used in conjunction with other components or software units implemented by the processor 102 by executing at least a portion of a program. The program may be stored on a computer-readable medium such as memory 103 or storage device 104, and when executed by the processor 102, it may perform one or more functions.

[0059] Memory 103 and storage device 104 may include any suitable type of mass storage device provided for storing any type of information that processor 102 may need to operate. For example, memory 103 and storage device 104 may be volatile or non-volatile, magnetic, semiconductor-based, magnetic tape-based, optical, removable, non-removable, or other types of storage devices or tangible (i.e., non-transient) computer-readable media, including but not limited to ROM, flash memory, dynamic RAM, and static RAM. Memory 103 and / or storage device 104 may be configured to store one or more computer programs that can be executed by processor 102 to perform the functions disclosed herein. For example, memory 103 and / or storage device 104 may be configured to store a program that can be executed by processor 102 to perform animal detection. Memory 103 and / or storage device 104 may also be configured to store information and data used by processor 102.

[0060] Camera module 116 can be configured to acquire images or videos of a region of interest. For example, camera module 116 can be placed on top of (or to the side of) an animal enclosure to generate images or videos capturing a scene of the animal enclosure. In another instance, camera module 116 can be mounted on a drone (UAV) capable of flying over an animal enclosure. In some embodiments, camera module 116 can preprocess images or videos. For example, camera module 116 can perform deblurring, super-resolution, or any other suitable operation on the images or videos.

[0061] For example, camera module 116 may include an infrared camera with an ultra-wide-angle lens. The infrared camera's shooting area can cover the entire animal enclosure, and the infrared camera can acquire images or videos during the day or at night. Camera module 116 can be mounted above the animal enclosure, with the lens facing downwards towards the animal enclosure. Camera module 116 can output images or videos capturing the entire animal enclosure. The animal enclosure may have a size of 6m x 5m, and camera module 116 can be mounted 4m above the animal enclosure.

[0062] In some embodiments, camera module 116 can forward images or videos to system 101, allowing system 101 to detect animals in an animal enclosure based on the images or videos. In some embodiments, camera module 116 can forward images or videos to computing device 112, allowing computing device 112 to display the images or videos on a screen. Although in Figure 1AOnly one camera module 116 is shown in the illustration. The operating environment 100 may include multiple camera modules 116, one of which is used for an animal enclosure in a farm. In other embodiments, more than one camera module 116 may be provided for an animal enclosure in a farm. The number of camera modules 116 used for an animal enclosure may be the same or different, depending on the different scenario or application (e.g., the size of the animal enclosure, the number of animals typically present in the animal enclosure, etc.).

[0063] The computing device 112 may be located in the farm. For example, the computing device 112 may be a server, desktop computer, laptop computer, tablet computer, or any other suitable electronic device with a processor and memory located in the farm. The computing device 112 may display images or videos acquired by the camera module 116 on a display device. In some embodiments, the computing device 112 may also display messages received from the system 101 on a display device.

[0064] User device 114 may be a computing device including a processor and memory. For example, user device 114 may be a desktop computer, laptop computer, tablet computer, smartphone, game controller, television (TV), music player, wearable electronic device such as a smartwatch, internet device, smart vehicle, or any other suitable electronic device with a processor and memory. User device 114 may be operated by a user associated with the farm (e.g., owner, manager, worker, or any other person). In some embodiments, user device 114 may receive messages from system 101 and display those messages on the screen of user device 114. In some embodiments, user device 114 may receive images or videos from camera module 116 and display those images or videos on the screen of user device 114.

[0065] Figure 1B A block diagram of another exemplary operating environment 150 of the system 101 disclosed in this invention, configured to detect animals in a region of interest, is shown. Operating environment 150 may include system 101, computing device 112, user device 114, camera module 116, and any other suitable components.

[0066] System 101, computing device 112, and camera module 116 may be located in a breeding farm. Camera module 116 may be communicatively coupled to system 101 via a wired connection (e.g., cable connection, Universal Serial Bus (USB) connection) or a wireless connection (e.g., Bluetooth connection). System 101 may be communicatively coupled to computing device 112 via a wired or wireless connection. User equipment 114 may be communicatively coupled to computing device 112 via a wired or wireless connection.

[0067] The computing device 112 may include a processor 155, a memory 156, and a storage device 158. The processor 155 may include an analysis module 106 and a training module 107. The processor 155, memory 156, and storage device 158 may have structures similar to those of processor 102, memory 103, and storage device 104, respectively, and similar descriptions will not be repeated here.

[0068] System 101 may include processor 102, memory 103, and storage device 104. Processor 102 may include detection module 105. In some embodiments, system 101 may be embodied on a system-on-a-chip (SoC). Processor 102, memory 103, and storage device 104 may be implemented in an embedded integrated circuit (IC) of the SoC. In some embodiments, the SoC may be placed in the same location as camera module 116. For example, the SoC and camera module 116 may be placed together above an animal enclosure.

[0069] In some embodiments, the embedded IC of the SoC may include a neural network module that can implement the animal detection model described herein or any other neural network model with low power consumption. The embedded IC of the SoC may also include input / output interfaces that support image formats, video formats, or combinations thereof. For example, camera module 116 can input images of the animal fence to the SoC via a wired or wireless connection. Detection module 105, installed in the SoC, can process the images to generate detection results and forward them to computing device 112, allowing a user to monitor the animal fence in real-time or near real-time on the screen of computing device 112.

[0070] Figure 2A An exemplary flow diagram of the operation for detecting animals in a region of interest disclosed in this invention is shown. The region of interest may be, for example, an animal enclosure in a farm. In some embodiments, a camera module 116 located above or to the side of the animal enclosure can capture an image 202 of the scene in the region of interest. The camera module 116 can then forward the image 202 to the detection module 105.

[0071] The detection module 105 can be configured to feed the image 202 to the animal detection model 250 to generate a set of affinity field maps 206 for a set of keypoints and a set of probability maps 208 for a set of keypoints. In some embodiments, a set of keypoints may include one or more keypoints, and a set of probability maps 208 may include one or more probability maps 208 corresponding to one or more keypoints respectively. A set of keypoints may include one or more keypoint sets, and a set of affinity field maps 206 may include one or more affinity field maps 206 corresponding to one or more keypoint sets respectively.

[0072] For example, detection module 105 can input image 202 into animal detection model 250, causing animal detection model 250 to generate a set of feature maps 204 from image 202 using a series of depthwise separable convolutional blocks. Animal detection model 250 can be configured to generate a set of probability maps 208 and a set of affinity field maps 206 from the set of feature maps 204. For example, animal detection model 250 is configured to generate a probability map 208 for each keypoint and an affinity field map 206 for each set of keypoints. Each probability map 208 and each affinity field map 206 can be output from different channels (or layers) of animal detection model 250. See Appendix below. Figure 3A -4C provides a more detailed description of the animal detection model 250.

[0073] The affinity field diagram 206 of the keypoint set can depict an affinity field describing the connection trends between keypoints in the keypoint set. For example, the keypoint set may include a first keypoint and a second keypoint, and form an animal fragment (e.g., a limb or body fragment) between the first keypoint and the second keypoint. Image 202 may include one or more instances of the animal fragment, where each instance of the animal fragment belongs to a different animal. For example, image 202 may include one or more instances of a left foreleg, where each instance of the left foreleg belongs to a different animal. The affinity field diagram 206 of the keypoint set can then include one or more vector fields (e.g., two-dimensional (2D) vector fields). Each vector field may correspond to an instance of an animal fragment in image 202 and encode the position and orientation of the instance of the animal fragment in image 202.

[0074] In some embodiments, for each pixel belonging to a region of an animal fragment instance, the vector field of the animal fragment instance may include a vector encoding a direction from a first keypoint to a second keypoint. For example, if the point is located within a region belonging to the animal fragment instance, the vector field may include a unit vector for that point, where the unit vector points from the first keypoint to the second keypoint. If the point is located outside a region of the animal fragment instance, the vector field may include a zero-value vector for that point.

[0075] Figure 10A An exemplary affinity field diagram of the keypoint set is shown below. Go to... Figure 10A The keypoint set may include shoulder keypoints and hip keypoints. Figure 10A Image 1001 depicts two pigs and therefore includes two instances of body segments from shoulder keypoints to hip keypoints for each pig (e.g., each pig has an instance of a body segment). The affinity field diagram of the keypoint set includes a vector field 1002 for a first instance of the body segment for the first pig and a vector field 1004 for a second instance of the body segment for the second pig.

[0076] For each point located within the region belonging to the first instance of the body segment, the vector field 1002 may include a unit vector for that point, where the unit vector points from the shoulder keypoint of the first pig to the hip keypoint. If the point is located outside the region of the first instance of the body segment, the vector field 1002 may include a zero-value vector for that point. Here, the region belonging to the first instance of the body segment may be, for example, a rectangular region on the body of the first pig from the shoulder to the hip.

[0077] Similarly, for each point located within the region belonging to the second instance of the body segment, the vector field 1004 may include a unit vector for that point, where the unit vector points from the shoulder keypoint of the second pig to the hip keypoint. If the point is located outside the region of the second instance of the body segment, the vector field 1004 may include a zero-value vector for that point. The region belonging to the second instance of the body segment may also be a rectangular region on the body of the second pig extending from the shoulder to the hip.

[0078] In some embodiments, the x and y components can be used to represent a vector field (e.g., vector field 1002 or 1004), where the x and y components represent offsets in the x and y directions, respectively. The orientation and position of each point in the vector field can be determined based on a combination of the x and y components. For example, referencing... Figure 10A The x-component can be used to represent the offset from the shoulder keypoint to the hip keypoint in the x-direction, and the y-component can be used to represent the offset from the shoulder keypoint to the hip keypoint in the y-direction.

[0079] Return to Figure 2A In some embodiments, for a keypoint set including more than two keypoints (e.g., first, second, and third keypoints), a first affinity field map 206 can be generated for the first and second keypoints, and a second affinity field map 206 can be generated for the second and third keypoints. The affinity field map 206 for the keypoint set can then be generated as a combination of the first and second affinity field maps 206.

[0080] Refer again Figure 2AThe probability map 208 for keypoints may include an array of probability values ​​for keypoints, where each probability value corresponds to a location (e.g., a pixel location) in image 202 and describes the probability that a keypoint appears at that location in image 202. For example, for a first location in image 202, probability map 208 may include a probability value "0.8", indicating that the probability of a keypoint appearing at the first location in image 202 is 0.8; for a second location in image 202, probability map 208 may include a probability value "0.5", indicating that the probability of a keypoint appearing at the second location in image 202 is 0.5. In some embodiments, each probability map 208 may be normalized by a sigmoid function such that each probability value included in the probability map 208 is within the range of 0 and 1.

[0081] Next, for each keypoint, the detection module 105 can process the probability map 208 corresponding to the keypoint and generate a location map 210 for the keypoint. For example, the detection module 105 can use a local maximum algorithm to determine one or more locations of keypoints in image 202, such that each location of a keypoint corresponds to a pixel location in the probability map 208 that has a local maximum probability value. The detection module 105 generates the location map 210 for the keypoints, making it possible to identify one or more locations of keypoints in the location map 210. As a result, the detection module 105 can generate a set of location maps 210 for the set of keypoints from the set of probability maps 208.

[0082] The detection module 105 can combine the set of location maps 210 to generate a combined location map 212. The combined location map 212 can identify one or more locations of each key point, such that at least some or all of the locations of a set of key points are identified in the combined location map 212.

[0083] The detection module 105 can also be configured to determine a set of fragment clusters 214 for a set of keypoints based on the combination of the location map 212 and the affinity field map 206, as described in more detail below. Each fragment cluster 214 corresponding to a set of keypoints includes one or more instances of an animal fragment associated with the set of keypoints. For example, suppose the set of keypoints includes head keypoints and shoulder keypoints that form the neck portion of an animal. The fragment clusters 214 of the keypoint set may include one or more instances of a neck portion, each instance of the neck portion representing the neck of a different animal in image 202.

[0084] Specifically, for each keypoint set including a first keypoint and a second keypoint, the detection module 105 can determine one or more first positions of the first keypoint and one or more second positions of the second keypoint from the combined location map 212. The detection module 105 can match one or more first positions of the first keypoint with one or more second positions of the second keypoint to form a fragment cluster 214 of the keypoint set based on the affinity field map 206 of the keypoint set.

[0085] For example, for each first location of a first keypoint, detection module 105 can measure one or more correlations between the first location of the first keypoint and one or more second locations of a second keypoint based on the affinity field map 206 of the keypoint set. Detection module 105 can then determine the maximum correlation from the one or more correlations and determine whether the maximum correlation satisfies a correlation threshold. The correlation threshold can have a value of 0.5 or another suitable value. In response to the maximum correlation satisfying the correlation threshold (e.g., the maximum correlation is greater than or equal to the correlation threshold), detection module 105 can form instances of animal fragments in the fragment cluster 214 of the keypoint set by associating the first location of the first keypoint with the second location of the second keypoint corresponding to the maximum correlation. Instances of animal fragments appear between the first and second locations in image 202.

[0086] refer to Figure 10B This illustrates an exemplary process for generating fragment clusters for keypoint sets. Go to... Figure 10B In image 1001, two locations s1 and s2 are identified as shoulder keypoints for two pigs, and two locations t1 and t2 are identified as rump keypoints for two pigs. To match the shoulder keypoint locations s1 and s2 with the rump keypoint locations t1 and t2, the correlation coefficient (E(S)) is measured between each shoulder keypoint location and each rump keypoint location. i , t j ), and 1≤i≤2 and 1≤j≤2).

[0087] In some embodiments, the positions s of the shoulder keypoint and the hip keypoint i and t j N sampling points can be uniformly identified between two locations, and the correlation degree E(s) can be calculated using the following equation. i ,t j ):

[0088]

[0089] In the equation above, F(g) n ) represents sampling point g n The vector field at point 1 ≤ n ≤ N. This can be derived from... Figure 10AAffinity field diagram determines F(g) n ). F(g) represents n )and The dot product between them. When each sampling point g n Vector field F(g) at point n The direction of the shoulder key point and the position of the shoulder key point i The location of the key point of the buttocks t j When the directions are the same, the correlation degree E(s) i ,t j The maximum value can be 1, representing the position s of the key point on the shoulder. i and the location of key points on the buttocks t j There is the strongest correlation between them. When each sampling point g... n Vector field F(g) at point n ) has a zero-value vector (e.g., F(g) n When )=0), the correlation degree E(s) i ,t j The value can reach 0, indicating the position of the shoulder key point s. i and the location of key points on the buttocks t j There is no connection between them.

[0090] For example, refer to Figure 10B To determine the correlation E(s1,t2) between the shoulder keypoint s1 and the hip keypoint t2, three sampling points p1, p2, and p3 are uniformly identified along the line connecting s1 and t2. The correlation E(s1,t2) can then be calculated as follows: Similarly, to calculate the correlation E(s1,t1) between the shoulder keypoint position s1 and the hip keypoint position t1, three sampling points q1, q2, and q3 are uniformly identified along the line connecting s1 and t1. Then, E(s1,t1) can be calculated as follows:

[0091] exist Figure 10B In, at each sampling point q n The vector field F(q) at (1≤n≤3) n ) has a zero-valued vector (e.g., F(q) n Since )=0), and therefore E(s1,t1) can have a value of 0, indicating that there is no correlation between the position s1 of the shoulder keypoint and the position t1 of the hip keypoint. On the other hand, since each sampling point p n Vector field F(q) at point nThe direction of E(s1,t2) is essentially the same as the direction from the shoulder keypoint position s1 to the hip keypoint position t2, therefore the direction of E(s1,t2) can have a non-zero value, and thus be greater than E(s1,t1). The value E(s1,t2) can also be greater than the association threshold (e.g., threshold 0.5). Therefore, the shoulder keypoint position s1 and the hip keypoint position t2 are associated to form the first instance of the body segment. That is, both the shoulder keypoint position s1 and the hip keypoint position t2 belong to the first instance of the body segment (e.g., s1 and t2 belong to the first pig). The first instance of the body segment is in Figure 10B It is shown in the figure by connection 1006 and appears between position s1 and position t2 in image 1001.

[0092] Similarly, the position s2 of the shoulder keypoint is associated with the position t1 of the hip keypoint to form a second instance of the body segment (e.g., s2 and t1 belong to the second pig). The second instance of the body segment is... Figure 10B It is shown in the figure as 1008, and appears between position s2 and position t1 in image 1001.

[0093] As a result, a fragment cluster is generated for the keypoint set. The fragment cluster includes a first instance of the body fragment of the first pig (shown as connection 1006) and a second instance of the body fragment of the second pig (shown as connection 1008).

[0094] Return to Figure 2A By performing operations similar to those described above, the detection module 105 can generate a fragment cluster for each set of keypoints. Therefore, a fragment cluster 214 is generated for each set of keypoints.

[0095] The detection module 105 can classify each instance of each animal fragment in the cluster 214 of fragments into one or more connection graphs 216, such that one or more instances of one or more animal fragments belonging to the same animal are grouped into the same connection graph 216. Each connection graph 216 can summarize the presence of animals in image 202.

[0096] For example, suppose a first set of keypoints may include head keypoints and shoulder keypoints forming the neck, and a second set of keypoints may include shoulder keypoints and a third keypoint (e.g., the left forearm keypoint) forming a limb (e.g., the left forearm keypoint). A first fragment cluster 214 of the first set of keypoints may include one or more instances of the neck portion. For example, a first instance of the neck portion appears between the position L11 of the first keypoint and the position L21 of the second keypoint, while a second instance of the neck portion appears between the position L12 of the first keypoint and the position L22 of the second keypoint. A second fragment cluster 214 of the second set of keypoints may include one or more instances of the left forearm keypoint. For example, a first instance of the left forearm keypoint appears between the position L21 of the second keypoint and the position L31 of the third keypoint, while a second instance of the left forearm keypoint appears between the position L22 of the second keypoint and the position L32 of the third keypoint.

[0097] Then, detection module 105 can determine that the first instance of the neck portion and the first instance of the left forearm belong to the first animal appearing in image 202 because they share the location L21 of a common second keypoint. A first connection graph 216 can be generated for the first animal to include connections representing the first instance of the neck portion and connections representing the first instance of the left forearm. Similarly, detection module 105 can determine that the second instance of the neck portion and the second instance of the left forearm belong to the second animal appearing in image 202 because they share the location L22 of the second keypoint. A second connection graph 216 can be generated for the second animal to include connections representing the second instance of the neck portion and connections representing the second instance of the left forearm.

[0098] The detection module 105 can also be configured to detect one or more animals present in the region of interest based on one or more connection graphs 216. For example, the detection module 105 can determine that the total number 218 of animals present in the region of interest is equal to the total number of connection graphs 216 in the image 202.

[0099] In another example, detection module 105 can determine the geographic location 220 of each animal present in the region of interest based on the location of the corresponding connection graph 216 in image 202. The location of the corresponding connection graph 216 in image 202 can be the location of a point (e.g., the center point) of the corresponding connection graph 216. Specifically, detection module 105 can convert the location of the corresponding connection graph 216 in image 202 into a geographic location 220 in the region of interest. The geographic location 220 can be, for example, geographic coordinates within the region of interest. For example, if the location of connection graph 216 is at the center of image 202, then the geographic location 220 of the animal corresponding to connection graph 216 is at the center point of the region of interest.

[0100] The detection module 105 can also be configured to determine one or more postures of one or more animals based on one or more connection graphs 216. Specifically, the detection module 105 can determine the posture 222 of an animal based on the corresponding connection graph 216 of each animal present in the region of interest. For example, the corresponding connection graph 216 can indicate that the animal's posture 222 can be a standing posture, a lying posture, or any other suitable posture.

[0101] In some embodiments, each connection diagram 216 may include one or more body connections (or trunk connections), one or more limb connections, or a combination thereof. Body connections may be connections formed by body keypoints. Limb connections may be connections formed by limb keypoints or connections formed by a combination of limb keypoints and body keypoints. Exemplary body connections and limb connections are shown in... Figure 8A middle.

[0102] In some embodiments, the detection module 105 can determine the total number 218 of animals and the geographic location 220 of each animal based on one or more body connections in each connection graph 216. For example, the total number 218 of animals can be equal to the total number of connection graphs 216 with at least one body connection. In other words, if a connection graph 216 only has limb connections, it may not be considered a single animal in the process of calculating the total number 218. Furthermore, the geographic location 220 of each animal can be determined based on the position of the body connections in the corresponding connection graph 216.

[0103] In some embodiments, the detection module 105 can determine the animal's posture 222 based on one or more body connections and one or more limb connections in the corresponding connection diagram 216. For example, if image 202 is a top view, and if the corresponding connection diagram 216 includes one or more body connections and one or more limb connections (e.g., the animal's body and one or more limbs are visible in image 202), the animal's posture 222 can be determined as a down posture. On the other hand, if the corresponding connection diagram 216 only includes one or more body connections (e.g., only the animal's body is visible in image 202), the animal's posture can be determined as a standing posture.

[0104] Analysis module 106 can be configured to perform behavioral analysis 224 on the one or more animals detected in image 202 based on one or more postures of the one or more animals and generate analysis results. Analysis module 106 can perform diagnosis on the one or more animals based on the analysis results to generate a diagnostic report 226. Analysis module 106 can also provide messages describing the analysis results, diagnostic reports, or a combination thereof.

[0105] For example, the analysis module 106 can determine whether any animals are missing from the animal enclosure based on the total number 218 animals detected in the enclosure and the assumed number of animals in the enclosure. If at least one animal is missing from the enclosure, an alert message can be generated to warn the farm's users about the missing animal.

[0106] In another example, analysis module 106 can perform behavioral analysis on one or more animals based on one or more postures to identify animals exhibiting abnormal behavior. Analysis module 106 can diagnose animals exhibiting abnormal behavior to generate a diagnostic report and can provide a warning message describing the animal's abnormal behavior, the diagnostic report, or a combination thereof. For example, an animal exhibiting abnormal behavior might be one that remains in a lying position for a predetermined period of time while other animals in the same enclosure gather together to eat along a feeding trough. Analysis module 106 can determine that the animal exhibiting abnormal behavior may be sick. Analysis module 106 can provide a warning message to user device 114, thereby notifying the farm's users of the sick animal.

[0107] In some embodiments, before applying animal detection model 250 to detect animals in regions of interest, training module 107 can be configured to train animal detection model 250 using multiple training images. Multiple training images can capture animals with multiple body shapes and poses in multiple living environments. Multiple training images can be captured at a set of times with different illumination levels. Therefore, the diversity and robustness of animal detection model 250 can be improved by training with multiple training images.

[0108] In some embodiments, for each animal captured in the training images, the training module 107 may mark one or more keypoints of the animal at one or more locations in the training images. The training module 107 may assign a visibility attribute to each keypoint marked at the corresponding location in the training images. The visibility attribute may indicate whether the keypoint marked at the corresponding location in the training images is visible in the training images. For example, if the keypoint marked at the corresponding location in the training images is visible in the training images, the visibility attribute may be identified as "visible".

[0109] In another example, if a keypoint marked at a corresponding location in the training image is invisible in the training image, but its position in the training image is predictable based on the positions of other keypoints in the animal, the visibility attribute can be labeled "invisible but predictable." For example, even if a keypoint is invisible in the training image, its position and the positions of other keypoints in the animal can conform to the physical characteristics of the animal's body pattern. As a result, the position of invisible keypoints in the training image can be predicted based on the positions of other keypoints. Figure 8BExemplary keypoints that are not visible in the image but whose locations are predictable in the image are shown, which will be described in more detail below.

[0110] Therefore, the training images of the animal detection model 250 can be processed to identify (1) keypoints visible in the images, and (2) keypoints invisible in the images but whose positions are predictable. After training the animal detection model 250 using these training images, not only can the animal detection model 250 process the keypoint information visible in the images, but also, if the keypoint information occluded conforms to the physical characteristics of the animal's body pattern, the animal detection model 250 can process the keypoint information occluded (e.g., invisible) in the images.

[0111] In some embodiments, each keypoint labeled in the training image is configured to have a two-dimensional Gaussian distribution. The covariance of the Gaussian distribution may be proportional to the minimum distance between the keypoint and one or more adjacent keypoints, with a scaling factor of 0.15 or any other suitable value.

[0112] For example, a set of keypoints includes head keypoints, shoulder keypoints, and abdomen keypoints. A connection diagram for the keypoints can include connections from head keypoints to shoulder keypoints and connections from shoulder keypoints to abdomen keypoints. The covariance of the distribution of head keypoints can be proportional to the distance between head keypoints and shoulder keypoints. A larger distance indicates a larger covariance of the distribution. Since shoulder keypoints connect to head keypoints and abdomen keypoints, the covariance of the distribution of shoulder keypoints can be proportional to the minimum of a first distance between head keypoints and shoulder keypoints and a second distance between shoulder keypoints and abdomen keypoints.

[0113] In some embodiments, a training database may be established in system 101 or computing device 112 for training animal detection model 250. The training database may include multiple training images. For example, training images (e.g., 2,500 images) may be extracted from multiple videos (e.g., 100 videos) taken at different farms, where each video has a duration of several minutes (e.g., 5 minutes).

[0114] Figure 2B Another exemplary procedure for detecting an animal in a region of interest, as disclosed in this invention, is shown. (Refer to...) Figure 1B The operating environment 150 shown is used to describe Figure 2B The operation flow is as follows. In some embodiments, the camera module 116 can perform image acquisition and preprocessing operations 230 to generate an image. The camera module 116 can forward the image to system 101, which is contained on an embedded IC of the SoC.

[0115] The embedded IC may include a neural network module that can be configured to implement the operation of the animal detection model 250. For example, the detection module 105 in the embedded IC may use the neural network module to implement the operation of the animal detection model 250, thereby generating a set of affinity field maps of a set of keypoints and a set of probability maps of a set of keypoints from an image.

[0116] The detection module 105 in the embedded IC can implement the key point matching and connection graph generation process 232. For example, by performing a process similar to the one described above... Figure 2A The operations described herein allow the detection module 105 to generate one or more connection graphs based on a set of affinity field graphs and a set of probability graphs.

[0117] The detection module 105 in the embedded IC can perform animal detection operation 234 to detect one or more animals present in a region of interest based on one or more connection graphs. For example, by performing something similar to the above reference... Figure 2A The operations described herein allow the detection module 105 to generate detection results describing the total number of animals present in the region of interest, the geographical location of each animal present in the region of interest, one or more poses of one or more animals, or a combination thereof.

[0118] The detection module 105 in the embedded IC can forward the detection results to the analysis module 106 implemented in the computing device 112 via a wired or wireless connection (e.g., Wi-Fi or Bluetooth). The analysis module 106 can perform animal behavior analysis and diagnosis 236 based on the detection results. For example, by performing something similar to the above reference... Figure 2A The described operations allow the analysis module 106 to identify animals exhibiting abnormal behavior and generate a diagnostic report for those animals. The analysis module 106 can also generate a warning message describing the animal's abnormal behavior and the diagnostic report, and can forward the warning message to the user device 114. The user device 114 can execute warning operation 238 to present the warning message to the farm's users.

[0119] Figure 2C A schematic diagram of an exemplary structure of the animal detection model 250 disclosed in this invention is shown. The animal detection model 250 may include a feature extraction model 254 applied in a serial manner and a two-stage detection model 256. The feature extraction model 254 may be configured to receive an image capturing a scene in a region of interest and generate a set of feature maps from the image using a series of depthwise separable convolutional blocks. Reference is made below to the appendix. Figure 3A - 3D description feature extraction model 254. The two-level detection model 256 can be configured to generate a set of probability maps and a set of affinity field maps based on a set of feature maps. See the appendix below. Figure 4A -4C describes a two-level detection model 256.

[0120] Figure 3A A schematic diagram illustrating an exemplary structure of the feature extraction model 254 disclosed in this invention is shown. The feature extraction model 254 may include convolutions 302, one or more convolutional sequences 303A, 303B, ..., 303N (also individually or collectively referred to as convolutional sequences 303), a merging layer 306, and a fully connected layer 308. In some embodiments, the feature extraction model 254 may have a structure similar to MobileNet.

[0121] Convolution 302 can be a standard convolution with a kernel size of 3×3, 32 filters, and a span of 2. The input to convolution 302 can have, for example, 2^24... 2 The size is ×3. The output of a 302 convolution can have 112. 2 The size is ×32.

[0122] Each convolutional sequence 303 may include one or more depthwise separable convolutional blocks 304A, ..., 304N (also individually or collectively referred to as depthwise separable convolutional blocks 304). The depthwise separable convolutional blocks 304 may include depthwise separable convolutions, which are in the form of factorized convolutions. The depthwise separable convolutional blocks 304 can factorize a standard convolution into depthwise convolutions and 1×1 convolutions called pointwise convolutions. This factorization has the effect of greatly reducing the model size (e.g., the parameters in the model can be greatly reduced).

[0123] Each depthwise separable convolutional block 304 may include sequentially applied extension layers, depthwise convolutional layers (e.g., depthwise convolutions), and pointwise convolutional layers (e.g., pointwise convolutions). Each of the extension layers, depthwise convolutional layers, and pointwise convolutional layers is followed by group normalization.

[0124] Group normalization can be a simple alternative to batch normalization. Group normalization divides the channels into groups and calculates the normalized mean and variance within each group. For example, each group can have 4, 6, or other suitable numbers of channels. The calculation of group normalization is independent of the batch size, and its accuracy is stable over a wide range of batch sizes.

[0125] Figure 3BA schematic diagram of an exemplary structure of the depthwise separable convolutional block 304 disclosed in this invention is shown. In such an embodiment, block 304 has a span of 2. The depthwise separable convolutional block 304 may include one or more of the following: a serially applied extension layer 320, a group normalization 322, an activation function 324 (e.g., ReLU6), a depthwise convolutional layer 326, a group normalization 328, an activation function 330 (e.g., ReLU6), a pointwise convolutional layer 332, a group normalization 334, and an activation function 336 (e.g., a linear activation function). The output of the depthwise separable convolutional block 304 is generated as the output from activation function 336.

[0126] Figure 3C A schematic diagram of another exemplary structure of the depth-separable convolution block 304 disclosed in this invention is shown. In such an embodiment, block 304 has a span of 1. Figure 3C The depthwise separable convolutional block 304 in the middle can include and Figure 3B The components in the depth-separable convolutional block 304 are similar to those in other components. However, Figure 3C The output of the depthwise separable convolutional block 304 is generated by adding the output from the activation function 336 to the input of the depthwise separable convolutional block 304.

[0127] Figure 3D An exemplary table (e.g., Table 1) of a parameter list including a list of convolutional sequences 303 disclosed in this invention is shown. The parameter “n” represents the number of repetitions of the depthwise separable convolutional blocks 304 in each convolutional sequence 303 (e.g., the number of depthwise separable convolutional blocks 304 in each convolutional sequence 303). The parameter “s” represents the number of spans of the first depthwise separable convolutional block 304 in each convolutional sequence 303, while the number of spans of any remaining depthwise separable convolutional blocks 304 in each convolutional sequence 303 is 1. The parameters “t” and “c” represent the number of spread factors and filters (e.g., output channels) in each depthwise separable convolutional block 304 of the convolutional sequence 303, respectively. All spatial convolutions use a 3×3 kernel.

[0128] Each row in Table 1 describes the parameter values ​​for the corresponding convolutional sequence 303. For example, the first row of Table 1 may specify the parameter values ​​for the first convolutional sequence 303, the second row of Table 1 may specify the parameter values ​​for the second convolutional sequence 303, and so on. For example, based on the second row of Table 1 (e.g., row 390), the second convolutional sequence 303 may include two depthwise separable convolutional blocks 304 (e.g., n = 2), the first depthwise separable convolutional block 304 having a span of 2 (e.g., s = 2), and the second depthwise separable convolutional block 304 having a span of 1. Each depthwise separable convolutional block 304 in the second convolutional sequence 303 may have a spread factor of 6 (e.g., t = 6) and 32 filters (e.g., c = 32).

[0129] Figure 4A A schematic diagram of an exemplary structure of the two-level detection model 256 disclosed in this invention is shown. The two-level detection model 256 may include a first-level neural network 402 and a second-level neural network 406. The first-level neural network 402 may be configured to generate a set of affinity field maps based on a set of feature maps from the feature extraction model 204.

[0130] The second-level neural network 406 can be configured to generate a set of probability maps based on the set of affinity field maps and the set of feature maps. For example, a set of affinity field maps and a set of feature maps can be concatenated via a concatenation operation 404 and input into the second-level neural network 406. Then, the second-level neural network 406 can generate a set of probability maps based on the concatenation of the set of affinity field maps and the set of feature maps.

[0131] Figure 4B A schematic diagram of an exemplary structure of a neural network 415 (e.g., a first-level neural network 402 or a second-level neural network 404) in a two-level detection model 256 disclosed in this invention is shown. For example, the neural network 415 may include one or more convolutional blocks 420A, 420B, 420C, and 420D (also individually or collectively referred to as convolutional block 420).

[0132] The input to neural network 415 can be processed by convolutional block 420A to generate a first output. The first output can be input to convolutional block 420B to generate a second output. The second output can be input to convolutional block 420C to generate a third output. The first, second, and third outputs are concatenated by concatenation operation 424 and input to convolutional block 420D, causing convolutional block 420D to generate the output of neural network 415.

[0133] Figure 4C A schematic diagram of an exemplary structure of a convolutional block 420 in a neural network 415 disclosed in this invention is shown. In some embodiments, the convolutional block 420 may include a convolutional layer 430, followed by a group normalization 432 and an activation function 434 (e.g., a parameter-corrected linear unit (PReLU) activation function).

[0134] Figure 5 This is a flowchart of an exemplary method 500 for detecting an animal in a region of interest, as disclosed in this invention. Method 500 can be implemented by system 101, particularly detection module 105, and may include steps 502-508 as described below. Some steps may be optional to perform the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in conjunction with... Figure 5 The different sequences shown will be executed.

[0135] In step 502, the detection module 105 can receive an image of the scene captured in the region of interest.

[0136] In step 504, the detection module 105 can feed the image to the animal detection model to generate a set of probability maps for a set of keypoints and a set of affinity field maps for a set of keypoints. For example, the detection module 105 can perform the same operation as described above. Figure 2A or Figure 2B The operations described are similar to those described above, to generate a set of probability maps and a set of affinity field maps.

[0137] In step 506, the detection module 105 can determine one or more connectivity graphs based on the set of probability graphs and the set of affinity field graphs. For example, the detection module 105 can perform the same steps as described above. Figure 2A or Figure 2B The described operations are similar to those used to determine one or more connection graphs. In some embodiments, each connection graph may outline the presence of an animal in an image.

[0138] In step 508, the detection module 105 can detect one or more animals present in the region of interest based on one or more connection graphs. For example, based on one or more connection graphs, the detection module 105 can determine the total number of animals present in the region of interest, the geographical location of each animal present in the region of interest, and one or more poses of one or more animals.

[0139] Figure 6 This is a flowchart of an exemplary method 600 for determining one or more connection graphs disclosed in this invention. Method 600 can be implemented by system 101, particularly detection module 105, and can include steps 602-608 as described below. In some embodiments, method 600 can be performed to implement... Figure 5 Step 506. Some steps may be optional to perform the disclosure provided herein. Furthermore, some steps may be performed concurrently or in conjunction with... Figure 6 The different sequences shown will be executed.

[0140] In step 602, for each keypoint, the detection module 105 can process the probability map corresponding to that keypoint to generate a location map of that keypoint. As a result, a set of location maps are generated from the set of probability maps.

[0141] In step 604, the detection module 105 can combine the set of location maps to generate a combined location map. The combined location map identifies one or more locations of each keypoint appearing in the image.

[0142] In step 606, for each keypoint set including a first keypoint and a second keypoint, the detection module 105 can match one or more first positions of the first keypoint with one or more second positions of the second keypoint to form a fragment cluster of the keypoint set based on the affinity field map of the keypoint set. As a result, a set of fragment clusters is generated for each set of keypoints. Each fragment cluster corresponding to a keypoint set may include one or more instances of animal fragments associated with the corresponding keypoint set.

[0143] In step 608, the detection module 105 can classify the instances of animal fragments in the fragment cluster into one or more connection graphs, so that one or more instances of one or more animal fragments belonging to the same animal are grouped into the same connection graph.

[0144] Figure 7 This is a flowchart of an exemplary method 700 for forming a cluster of fragments containing a set of key points, as disclosed in this invention. Method 700 can be implemented by system 101, particularly detection module 105, and may include steps 702-714 as described below. In some embodiments, method 700 can be performed to implement this method. Figure 6 Step 606. Some steps may be optional to perform the disclosure provided herein. Furthermore, some steps may be performed concurrently or in conjunction with... Figure 7 The different sequences shown will be executed.

[0145] In some embodiments, the keypoint set may include a first keypoint and a second keypoint, and the first and second keypoints form an animal fragment, such as a body segment or limb segment. The first keypoint may be identified at one or more first locations in the combined location map. The second keypoint may be identified at one or more second locations in the combined location map.

[0146] In step 702, the detection module 105 can select the first position to be processed from one or more first positions of the first key point.

[0147] In step 704, the detection module 105 can measure one or more correlations between the first position of the first key point and one or more second positions of the second key point based on the affinity field map of the key point set.

[0148] In step 706, the detection module 105 can determine the maximum correlation degree from one or more correlation degrees.

[0149] In step 708, the detection module 105 can determine whether the maximum correlation degree meets the correlation threshold. If the maximum correlation degree meets the correlation threshold, method 700 proceeds to step 710. Otherwise, method 700 proceeds to step 712.

[0150] In step 710, the detection module 105 can form instances of animal fragments in the fragment cluster by associating the first position of the first keypoint with the second position of the second keypoint corresponding to the maximum correlation.

[0151] In step 712, the detection module 105 can determine whether there are any remaining first positions of the first keypoint to be processed. In response to the existence of at least one remaining first position of the first keypoint to be processed, method 700 returns to step 702. Otherwise, method 700 proceeds to step 714.

[0152] In step 714, the detection module 105 may output a cluster of fragments associated with the keypoint set. The cluster of fragments may include one or more instances of animal fragments associated with the keypoint set.

[0153] Figure 8A -8B is a graphical representation of a set of key points for marking an animal (e.g., a pig) as disclosed by an exemplary embodiment of the present invention. Figure 8A A side perspective view of a pig is shown, while Figure 8B A top-down perspective view of the pig is shown.

[0154] In some examples, the marking method can be used to identify key points of a pig, where key points may include the pig's two ears, a point on the upper surface of the pig's shoulder ("shoulder point"), and a point on the upper surface of the pig's rump ("rump point"). Because this marking method only marks key points on the surfaces of different pigs, mismatches between key points of different pigs can easily occur. For example, in the case of two pigs placed back-to-back on the floor, the shoulder and rump points of the two pigs are close to each other, and mismatches between the shoulder and rump points of the two pigs can easily occur.

[0155] Furthermore, because the key points identified by this labeling method are unevenly distributed on the pig's surface, some pigs may be missed in the detection results. For example, if the pigs are huddled together or eating along the edge of the container, some pigs' ear and shoulder points may be obscured and not visible, with only the pig's rump point being exposed. Using this labeling method, these pigs may not be identified as valid individuals because only their rump points are exposed. Therefore, these pigs may be missed in the detection results.

[0156] Unlike the marking methods described above, the key points described in this invention can be marked at the geometric center of the animal's torso or limbs (rather than on the animal's surface). For example, a set of key points can be marked as follows: Figure 8AThe pig is marked as shown. The set of key points may include one or more of the following: head key point (marked as "0"), shoulder key point ("1"), abdomen key point ("2"), rump key point ("3"), left foreleg elbow key point ("4"), left foreleg key point ("5"), right foreleg elbow key point ("6"), right foreleg key point ("7"), left hind leg elbow key point ("8"), left hind leg key point ("9"), right hind leg elbow key point ("10"), and right hind leg key point ("11").

[0157] Furthermore, each keypoint described herein may be assigned a visibility attribute. In some embodiments, the visibility attribute may indicate that the keypoint is visible in the image. Alternatively, the visibility attribute may indicate that the keypoint is not visible in the image, but its location is predictable in the image. Or, the visibility attribute may indicate that the keypoint is not visible in the image and its location is unpredictable in the image.

[0158] For example, refer to Figure 8A Keypoints 0-5, 7-9, and 11 are visible in image 802 and are depicted using circles. Each of keypoints 0-5, 7-9, and 11 is assigned a visibility attribute, indicating that the keypoint is visible in image 802. On the other hand, each of keypoints 6 and 10 is not visible in image 802, but their positions can be predicted based on the pig's body pattern. Each of keypoints 6 and 10 is depicted using rectangles and assigned a visibility attribute, indicating that the keypoint is not visible in image 802 but its position is predictable in image 802.

[0159] In another example, refer to Figure 8B Keypoints 0-3 are visible in image 804 and are depicted using circles. Each of keypoints 0-3 is assigned a visibility attribute, indicating that the keypoint is visible in image 804. On the other hand, each of keypoints 4, 6, 8, and 10 is invisible in image 802, but their positions can be predicted based on the pig's body pattern. Each of keypoints 4, 6, 8, and 10 is depicted using rectangles and assigned a visibility attribute, indicating that the keypoint is not visible in image 804 but its position is predictable in image 804. Furthermore, keypoints 5, 7, 9, and 11 are invisible in image 804, and their positions are unpredictable in image 804. Therefore, each of keypoints 5, 7, 9, and 11 is not identified in image 804 and is assigned a visibility attribute, indicating that the keypoint is not visible in image 804 and its position is unpredictable in image 804.

[0160] Refer again Figure 8AThe diagram also shows a connection diagram of a pig. The connection diagram includes one or more body connections (or trunk connections) and one or more limb connections. The one or more body connections include one or more of the following: connections from keypoint 0 to keypoint 1, connections from keypoint 1 to keypoint 2, and connections from keypoint 2 to keypoint 3. The one or more limb connections include one or more of the following: connections from keypoint 1 to keypoint 4, connections from keypoint 4 to keypoint 5, connections from keypoint 1 to keypoint 6, connections from keypoint 6 to keypoint 7, connections from keypoint 3 to keypoint 10, connections from keypoint 10 to keypoint 11, connections from keypoint 3 to keypoint 8, and connections from keypoint 8 to keypoint 9.

[0161] Figure 9A This is a graphic representation of an exemplary image 902 of an animal enclosure disclosed in this invention. The animal enclosure may be a pig enclosure. Image 902 is taken from a top view. Figure 9B This is a graphical representation of an exemplary location diagram 904 of the shoulder key points disclosed in this invention. Each location highlighted in location diagram 904 represents a corresponding shoulder key point of a pig in image 902. Figure 9C This is a graphical representation of an exemplary location diagram 906 of the rump key points disclosed in this invention. Each location highlighted in location diagram 906 represents a corresponding pig rump key point in image 902.

[0162] Figure 10A This is a graphical representation of an exemplary affinity field map of the keypoint set disclosed in this invention. The keypoint set may include shoulder keypoints and hip keypoints. The affinity field map of the keypoint set (including vector fields 1002 and 1004) is generated from image 1001 and overlaid on image 1001, as shown below. Figure 10A As shown above. Figure 10A And similar descriptions will not be repeated here.

[0163] Figure 10B This invention discloses a method based on... Figure 10A A graphical representation of an exemplary process for generating fragment clusters from affinity field maps. Based on Figure 10A The affinity field diagram identifies two instances of animal fragments associated with the keypoint set. The two instances of animal fragments are illustrated using connections 1006 and 1008. This has been described above. Figure 10B And similar descriptions will not be repeated here.

[0164] Figure 11A This is a graphical representation of an exemplary image 1102 showing an animal enclosure, and Figure 11B This illustrates the basis of an embodiment of the invention. Figure 11AA graphical representation of an exemplary affinity field map 1104 of the keypoint set generated from image 1102. The animal enclosure could be a pig enclosure. The keypoint set could include shoulder keypoints and hip keypoints. The affinity field map 1104 of the keypoint set can be generated based on image 1102. For example, the vector field F(p) at point p in the affinity field map 1104 can be identified using the x and y components, e.g., F(p) = (x(p), y(p)). In some embodiments, a field value v(p) representing the angle of the vector field F(p) at point p can be calculated. The field value v(p) can be in the range between -π and π. For example, the field value v(p) can be calculated as:

[0165]

[0166] Figure 12 This is a graphical representation of an exemplary connection diagram of the animal depicted in image 1201 of this invention. The animal may be a pig. Each connection diagram may include connections from the head keypoint to the shoulder keypoint and connections from the shoulder keypoint to the rump keypoint. For each pig in image 1201, in Figure 12 The diagram shows manually marked connection diagrams and connection diagrams generated by detection module 105. For example, for pig 1206, manually marked connection diagram 1204 and connection diagram 1202 generated by detection module 105 are provided.

[0167] exist Figure 12 In the image 1201, the manually labeled connection graph and the connection graph generated by the detection module 105 are substantially consistent. In some cases, the locations of some key points identified by the detection module 105 are more accurate than the locations of the manually labeled key points. Therefore, the animal detection system and method described herein can accurately identify animals from image 1201.

[0168] Another aspect of the invention relates to a non-transient computer-readable medium storing instructions that, when executed, cause one or more processors to perform the methods described above. The computer-readable medium may include volatile or non-volatile, magnetic, semiconductor, magnetic tape, optical, removable, non-removable, or other types of computer-readable media or computer-readable storage devices. For example, as disclosed, the computer-readable medium may be a storage device or memory module storing computer instructions. In some embodiments, the computer-readable medium may be a disk or flash drive on which computer instructions are stored.

[0169] According to one aspect of the present invention, a method for detecting animals in a region of interest is disclosed. An image capturing a scene within the region of interest is received. The image is fed into an animal detection model to generate a set of probability maps of a set of keypoints and a set of affinity field maps of a set of keypoints. One or more connectivity graphs are determined based on the set of probability maps and the set of affinity field maps. Each connectivity graph summarizes the presence of an animal in the image. One or more animals present in the region of interest are detected based on the one or more connectivity graphs.

[0170] In some embodiments, detecting one or more animals present in a region of interest includes determining that the total number of animals present in the region of interest is equal to the total number of connection graphs in one or more connection graphs.

[0171] In some embodiments, detecting one or more animals present in a region of interest includes determining the geographic location of each animal present in the region of interest based on the location of its corresponding connection graph in the image.

[0172] In some embodiments, one or more postures of one or more animals are determined based on one or more connection diagrams.

[0173] In some embodiments, behavioral analysis is performed on one or more animals based on one or more postures to generate analysis results. Diagnoses are then made on the one or more animals based on the analysis results to generate diagnostic reports. A message describing the analysis results, diagnostic reports, or a combination thereof is provided.

[0174] In some embodiments, the set of key points includes one or more of the following: head key points, shoulder key points, abdominal key points, hip key points, left foreleg elbow key points, left forefoot key points, right foreleg elbow key points, right forefoot key points, left hind leg elbow key points, left hind foot key points, right hind leg elbow key points, and right hind foot key points.

[0175] In some embodiments, the animal detection model is configured to generate a set of feature maps from an image using a series of depthwise separable convolutional blocks. From this set of feature maps, a set of probability maps and a set of affinity field maps are generated.

[0176] In some embodiments, the animal detection model includes a feature extraction model configured to generate a set of feature maps from an image using a sequence of depthwise separable convolutional blocks. Each depthwise separable convolutional block includes a concatenated extended layer, a depthwise convolutional layer, and a pointwise convolutional layer. Each of the extended layer, the depthwise convolutional layer, and the pointwise convolutional layer is followed by a group normalization.

[0177] In some embodiments, the animal detection model includes a two-level detection model comprising a first-level neural network and a second-level neural network. The first-level neural network is configured to generate the set of affinity field maps based on the set of feature maps. The second-level neural network is configured to generate a set of probability maps based on the set of affinity field maps and the set of feature maps. Each of the first-level and second-level neural networks includes one or more convolutional blocks, each convolutional block including a convolutional layer followed by a group normalization and PReLU activation function.

[0178] In some embodiments, a plurality of training images are used to train an animal detection model, the plurality of training images depicting animals with multiple body shapes and poses in multiple living environments. The plurality of training images are captured at multiple times with different illumination levels.

[0179] In some embodiments, for each animal captured in the training image, one or more keypoints of the animal are marked at one or more locations in the training image. A visibility attribute is assigned to each keypoint marked at the corresponding location in the training image.

[0180] In some embodiments, each keypoint labeled in the training image is configured to have a two-dimensional Gaussian distribution, the covariance of which is proportional to the minimum distance between the keypoint and one or more adjacent keypoints.

[0181] In some embodiments, determining the one or more connection graphs includes: generating a combined location map based on the set of probability maps; determining a set of fragment clusters of the set of keypoints based on the combined location map and the set of affinity field maps, wherein each fragment cluster corresponding to the set of keypoints includes one or more instances of animal fragments associated with the set of keypoints; and classifying each instance of each animal fragment in the set of fragment clusters into the one or more connection graphs such that one or more instances of one or more animal fragments belonging to the same animal are clustered into the same connection graph.

[0182] In some embodiments, generating the combined location map based on a set of probability maps includes: for each key point, processing the probability map corresponding to the key point, generating a location map of the key point, generating a set of location maps for the set of probability maps; and combining the set of location maps to generate the combined location map.

[0183] In some embodiments, determining a set of fragment clusters for a set of keypoints includes: for each set of keypoints including a first keypoint and a second keypoint, matching one or more first positions of the first keypoint with one or more second positions of the second keypoint to form a fragment cluster of the keypoint set, thereby generating the set of fragment clusters for the set of keypoints.

[0184] In some embodiments, matching the one or more first positions of the first keypoint with the one or more second positions of the second keypoint to form the fragment cluster includes: for each first position of the first keypoint, measuring one or more correlation degrees between the first position of the first keypoint and the one or more second positions of the second keypoint based on the affinity field map of the keypoint set; determining a maximum correlation degree from the one or more correlation degrees; determining whether the maximum correlation degree satisfies an association threshold; and in response to the maximum correlation degree satisfying the association threshold, forming an instance of the animal fragment in the fragment cluster by associating the first position of the first keypoint with a second position of the second keypoint corresponding to the maximum correlation degree. Instances of animal fragments appear between the first and second positions in the image.

[0185] According to another aspect of the present invention, a system for detecting animals in a region of interest is disclosed. The system includes a memory and a processor. The memory is configured to store instructions. The processor is coupled to the memory and configured to execute the instructions to perform a process including: receiving an image capturing a scene in the region of interest; feeding the image to an animal detection model to generate a set of probability maps of a set of keypoints and a set of affinity field maps of a set of keypoints; determining one or more connection graphs based on the set of probability maps and the set of affinity field maps, wherein each connection graph outlines the presence of an animal in the image; and detecting one or more animals present in the region of interest based on the one or more connection graphs.

[0186] In some embodiments, the processor and memory are implemented in the embedded IC of the SoC.

[0187] In some embodiments, the processor and memory are implemented in a cloud computing device.

[0188] In some embodiments, the system further includes a camera module configured to acquire images of a region of interest.

[0189] In some embodiments, in order to detect one or more animals present in a region of interest, the processor is configured to execute instructions to perform a process that further includes determining that the total number of animals present in the region of interest is equal to the total number of connection graphs in one or more connection graphs.

[0190] In some embodiments, in order to detect one or more animals present in a region of interest, the processor is configured to execute instructions to perform a process that further includes determining the geographic location of each animal present in the region of interest based on the location of a corresponding connection graph in the image.

[0191] In some embodiments, the processor is configured to execute instructions to perform a process that further includes determining one or more postures of one or more animals based on one or more connection graphs.

[0192] In some embodiments, the processor is configured to execute the instructions to perform the process, further comprising: performing behavioral analysis on the one or more animals based on the one or more postures to generate analysis results; diagnosing the one or more animals based on the analysis results to generate a diagnostic report; and providing a message describing the analysis results, the diagnostic report, or a combination thereof.

[0193] In some embodiments, a set of key points includes one or more of the following: head key points, shoulder key points, abdominal key points, hip key points, left foreleg elbow key points, left forefoot key points, right foreleg elbow key points, right forefoot key points, left hind leg elbow key points, left hind foot key points, right hind leg elbow key points, and right hind foot key points.

[0194] In some embodiments, the animal detection model is configured to generate a set of feature maps from an image using a series of depthwise separable convolutional blocks. From this set of feature maps, a set of probability maps and a set of affinity field maps are generated.

[0195] In some embodiments, the animal detection model includes a feature extraction model configured to generate a set of feature maps from an image using a sequence of depthwise separable convolutional blocks. Each depthwise separable convolutional block includes a concatenated extended layer, a depthwise convolutional layer, and a pointwise convolutional layer. Each of the extended layer, the depthwise convolutional layer, and the pointwise convolutional layer is followed by a group normalization.

[0196] In some embodiments, the animal detection model includes a two-level detection model comprising a first-level neural network and a second-level neural network. The first-level neural network is configured to generate the set of affinity field maps based on the set of feature maps. The second-level neural network is configured to generate a set of probability maps based on the set of affinity field maps and the set of feature maps. Each of the first-level and second-level neural networks includes one or more convolutional blocks, each convolutional block including a convolutional layer followed by a group normalization and PReLU activation function.

[0197] In some embodiments, the processor is configured to execute instructions to perform a process that further includes training an animal detection model using multiple training images depicting animals with multiple body shapes and poses in multiple living environments. The multiple training images are captured at multiple times with different illumination levels.

[0198] In some embodiments, the processor is configured to execute the instructions to perform the process further includes: for each animal captured in the training image, marking one or more keypoints of the animal at one or more locations in the training image; and assigning a visibility attribute to each keypoint marked at the corresponding location in the training image.

[0199] In some embodiments, each keypoint labeled in the training image is configured to have a two-dimensional Gaussian distribution, the covariance of which is proportional to the minimum distance between the keypoint and one or more adjacent keypoints.

[0200] In some embodiments, in order to determine the one or more connection graphs, the processor is configured to execute the instructions to perform the process further includes: generating a combined location map based on the set of probability maps, wherein the combined location map identifies one or more locations of each keypoint in the image; determining a set of fragment clusters of the set of keypoints based on the combined location map and the set of affinity field maps, wherein each fragment cluster corresponding to the set of keypoints includes one or more instances of animal fragments associated with the set of keypoints; and classifying each instance of each animal fragment in the set of fragment clusters into the one or more connection graphs such that one or more instances of one or more animal fragments belonging to the same animal are grouped into the same connection graph.

[0201] According to another aspect of the present invention, a non-transient computer-readable storage medium is disclosed. The computer-readable storage medium is configured to store instructions that, in response to execution by a processor, cause the processor to perform a process comprising: receiving an image capturing a scene in a region of interest; feeding the image to an animal detection model to generate a set of probability maps of a set of keypoints and a set of affinity field maps of a set of keypoints; determining one or more connection graphs based on the set of probability maps and the set of affinity field maps, wherein each connection graph outlines the presence of an animal in the image; and detecting one or more animals present in the region of interest based on the one or more connection graphs.

[0202] The above description of a particular implementation can be readily modified and / or adapted to various applications. Therefore, based on the teachings and instructions presented herein, such adaptations and modifications are intended to be within the meaning and scope of equivalents of the disclosed implementation.

[0203] The breadth and scope of this invention should not be limited by any of the above exemplary implementations, but should be defined solely by the appended claims and their equivalents.

Claims

1. A method for detecting animals in a region of interest, characterized in that, The method includes: Receive images of the scene captured in the region of interest; The image is fed into an animal detection model to generate a set of probability maps of a set of keypoints and a set of affinity field maps of a set of keypoints. The animal detection model is configured to generate a set of feature maps from the image using a series of depthwise separable convolutional blocks. The animal detection model further includes a first-level neural network and a second-level neural network. The first-level neural network is configured to generate the set of affinity field maps based on the set of feature maps, and the second-level neural network is configured to generate the set of probability maps based on the set of affinity field maps and the set of feature maps. Based on the set of probability maps and the set of affinity field maps, one or more connection graphs are determined, wherein each connection graph summarizes the presence of animals in the images; and Detect one or more animals present in the region of interest based on the one or more connection graphs.

2. The method according to claim 1, characterized in that, The detection of one or more animals present in the region of interest includes: The total number of animals present in the region of interest is determined to be equal to the total number of connection graphs in the one or more connection graphs.

3. The method according to claim 1, characterized in that, The detection of one or more animals present in the region of interest includes: The geographic location of each animal existing in the region of interest is determined based on the location of the corresponding connection chart in the image.

4. The method according to claim 1, characterized in that, The method also includes: One or more postures of the one or more animals are determined based on the one or more connection diagrams.

5. The method according to claim 4, characterized in that, The method also includes: Behavioral analysis is performed on the one or more animals based on the one or more postures to generate analysis results; Based on the analysis results, a diagnosis is performed on the one or more animals to generate a diagnostic report; and Provide a message describing the analysis results, the diagnostic report, or a combination thereof.

6. The method according to claim 1, characterized in that, The set of key points includes one or more of the following: head key points, shoulder key points, abdominal key points, hip key points, left foreleg elbow joint key points, left forefoot key points, right foreleg elbow joint key points, right forefoot key points, left hind leg elbow joint key points, left hind foot key points, right hind leg elbow joint key points, and right hind foot key points.

7. The method according to claim 1, characterized in that, The animal detection model also includes: A feature extraction model configured to generate the set of feature maps from the image using the series of depthwise separable convolutional blocks. Each depthwise separable convolutional block includes concatenated extended layers, depthwise convolutional layers, and pointwise convolutional layers, as well as... Each of the extended layer, the depthwise convolutional layer, and the pointwise convolutional layer is followed by group normalization.

8. The method according to claim 1, characterized in that, Each of the first-level neural network and the second-level neural network includes one or more convolutional blocks, each convolutional block including a convolutional layer followed by a group normalized and parameter-corrected linear unit (PReLU) activation function.

9. The method according to claim 1, characterized in that, The method also includes: The animal detection model is trained using multiple training images depicting animals with multiple body shapes and poses in multiple living environments, wherein the multiple training images are captured at multiple times with different illumination levels.

10. The method according to claim 9, characterized in that, The method also includes: For each animal captured in the training images, Mark one or more key points of the animal at one or more locations in the training image; and The visibility attribute is assigned to each keypoint marked at the corresponding location in the training image.

11. The method according to claim 10, characterized in that, Each keypoint labeled in the training image is configured to have a two-dimensional Gaussian distribution, the covariance of which is proportional to the minimum distance between the keypoint and one or more adjacent keypoints.

12. The method according to claim 1, characterized in that, The process of determining one or more connection graphs includes: A combined location map is generated based on the set of probability maps, wherein the combined location map identifies one or more locations of each key point in the image; Based on the combined location map and the set of affinity field maps, a set of fragment clusters of the set of keypoints is determined, wherein each fragment cluster corresponding to the keypoint set includes one or more instances of animal fragments associated with the keypoint set; and Each instance of each animal fragment in the set of fragment clusters is classified into one or more connection graphs, such that one or more instances of one or more animal fragments belonging to the same animal are grouped into the same connection graph.

13. The method according to claim 12, characterized in that, The generation of the combined location map based on the set of probability maps includes: For each keypoint, process the probability map corresponding to the keypoint, generate a location map of the keypoint, and generate a set of location maps for the set of probability maps; and The set of location maps are combined to generate the combined location map.

14. The method according to claim 12, characterized in that, The process of determining a cluster of fragments from the set of key points includes: For each set of keypoints including a first keypoint and a second keypoint, based on the affinity field map of the keypoint set, one or more first positions of the first keypoint are matched with one or more second positions of the second keypoint to form a fragment cluster of the keypoint set, thereby generating the set of fragment clusters for the set of keypoints.

15. The method according to claim 14, characterized in that, The step of matching one or more first positions of the first key point with one or more second positions of the second key point to form a fragment cluster of the key point set includes: For each first position of the first key point The affinity field map of the keypoint set is used to measure one or more correlations between the first position of the first keypoint and one or more second positions of the second keypoint; Determine the maximum correlation from the one or more correlation degrees; Determine whether the maximum correlation degree meets the correlation threshold; and In response to the maximum correlation satisfying the correlation threshold, an instance of the animal fragment in the fragment cluster is formed by associating the first position of the first keypoint with the second position of the second keypoint corresponding to the maximum correlation, wherein the instance of the animal fragment appears between the first position and the second position in the image.

16. An electronic device for detecting an animal in a region of interest, characterized in that, The electronic device includes: Memory configured to store program instructions; and A processor coupled to the memory and configured to execute the program instructions to implement the method as claimed in any one of claims 1 to 15.

17. The electronic device according to claim 16, characterized in that, The processor and the memory are implemented in an embedded integrated circuit of a system-on-a-chip, or the processor and the memory are implemented in a cloud computing device.

18. A non-transitory computer-readable storage medium storing program instructions, characterized in that, When the program instructions are executed by the processor, the method described in any one of claims 1 to 15 is implemented.

Citation Information

Patent Citations

  • An attention-based important object detection method

    CN109711463A

  • Ranging method and device, storage medium and electronic equipment

    CN112906691A