An intelligent stray cat capture device and method based on deep learning
Through the intelligent capture device based on deep learning, accurate identification and capture of stray cats is achieved, frightened and repeated capture is reduced, suitable environment and food is provided, and the efficiency and benefits of stray cat sterilization are improved.
Patent Information
- Application Number
- CN202410955677.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-07-17
AI Technical Summary
The existing stray cat capture devices have problems such as indiscriminate capture, frightening stray cats, relying on manual operation, inconvenient movement and installation difficulties, and repeated capture, which affects the efficiency and benefits of stray cats.
It adopts intelligent capture devices based on deep learning, including intelligent identification module, drive module, automatic feeding module, automatic sealing module and real-time monitoring module. It uses the camera and main control board to identify the breed and ear tag of cats, drive away non-target animals through ultrasonic alarms, automatically feed cat food, and capture target stray cats by revolving doors, and provide a constant temperature environment.
It improves the intelligence of the capture device, reduces repeated capture, provides suitable environment and food, and improves the efficiency and benefits of sterilizing stray cats.
Smart Images

Figure CN118947677B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of catchers, in particular to an intelligent stray cat catching device and method based on deep learning. Background Art
[0002] Stray cats generally refer to those feline animals that do not have a fixed residence and live outdoors for a long time. They may have had owners before but were abandoned for various reasons, or their offspring reproduced on their own without human care. Stray cats are different from wild animals because they usually have a certain dependence on humans and may look for food and shelter in residential areas, parks or other places where human activities are frequent. These cats may face health problems, overpopulation, and the risk of conflicts with other animals or humans due to lack of proper care and management.
[0003] However, there are more or less some defects in the existing stray cat capture or shelter systems. That is, these devices catch stray cats indiscriminately and may even capture the wrong targets; these devices have a high probability of scaring stray cats and causing stress, and there are few that can provide appropriate food and water for stray cats; traditional capture devices also rely heavily on human operation, which is undoubtedly disadvantageous for catching those stray cats that are cautious and timid; in addition, some devices are not convenient to move, difficult to install or difficult to remove after installation.
[0004] Nowadays, with the increasing attention of society to the welfare issues of stray cats, some stray cats will find new homes after being rescued and sterilized, but some will just have marks on their ears after sterilization and continue their stray life. This also leads to frequent repeated captures when using traditional capture devices to catch stray cats, seriously affecting the efficiency of promoting the universal sterilization of stray cats. Summary of the Invention
[0005] The present invention overcomes the deficiencies of the prior art and provides an intelligent stray cat catching device and method based on deep learning.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] The first aspect of the present invention discloses an intelligent stray cat catching device based on deep learning, and the catching device includes:
[0008] An intelligent recognition module, the recognition module includes a camera and a main control board, the camera is connected to the main control board, and a recognition model is obtained by learning and training the characteristic image data, and the recognition model is input into the main control board, so that the recognition module can identify the breed of the cat entering the device and whether it has an ear tag during operation; wherein the characteristic image data includes the appearance of cats of different types and breeds and ear tag image data;
[0009] An intelligent driving module, wherein the driving module includes an ultrasonic alarm. If the recognition module identifies that the animal entering the device is a non-target animal, the ultrasonic alarm is controlled to emit an alarm to drive the non-target animal away;
[0010] Automatic feeding module, which can deliver cat food into the device at regular intervals to prevent non-target animals from eating up all the food in the device and making it impossible for the device to continue to lure target stray cats;
[0011] The automatic sealing module includes a revolving door and a steering gear, and the steering gear is connected to the revolving door; when the recognition module identifies that the animal entering the device is a target stray cat, the steering gear controls the revolving door to close, completing the capture; and sends a signal to the cloud device to inform that the capture is completed, prompting timely processing of the captured object;
[0012] Real-time monitoring and image transmission module, through which the images captured by the camera can be transmitted to the cloud device, and the camera can be remotely accessed through the real-time monitoring and image transmission module to observe the real-time image situation inside the device;
[0013] The temperature control module includes a temperature sensor and a temperature control device. The temperature control module is used in conjunction with an intelligent central control unit to achieve automatic temperature regulation to submit the internal temperature environment of the device.
[0014] Furthermore, in a preferred embodiment of the present invention, the feeding module is arranged at the top of the device, and the feeding module includes a feeding turntable and a motor, and the feeding turntable is driven by the motor to rotate to deliver cat food to the bottom of the device.
[0015] Furthermore, in a preferred embodiment of the present invention, a power supply is also provided inside the device, and the power supply provides power to each module inside the device.
[0016] Further, in a preferred embodiment of the present invention, the trained recognition model is deployed to the main control board. The main control board makes a judgment based on the data captured by the camera. If the animal entering the device is a non-target animal, it controls the ultrasonic alarm to emit an alarm sound to drive away the non-target animal. If it is recognized that the animal entering the device is a target stray cat, the rotating door is controlled to close by the servo motor to complete the capture.
[0017] The second aspect of the present invention discloses a training method for an identification module of a stray cat intelligent capture device. The recognition model includes a cat face recognition model and a cat ear defect recognition model.
[0018] The cat face recognition model fuses local and global features, uses ResNet50 as the backbone network to obtain feature data, and sets three branches according to the characteristics of the extracted features: the Middle branch, the Global branch, and the Part branch. The Middle branch extracts the intermediate-dimensional global features of the backbone network, the Global branch extracts high-dimensional global features, and the Part branch extracts local features of uniformly sized blocks. Finally, the extracted features are fused as the feature representation for cat face re-identification.
[0019] The cat ear defect recognition model includes a backbone network, a neck network, a detection head network, and a separable self-attention module. In the input processing stage, the image data first undergoes preprocessing steps of normalization and data augmentation. Subsequently, the backbone network performs deep feature extraction on the preprocessed image. The neck network is to fuse the extracted feature maps to enhance the model's ability to recognize multi-scale targets. The detection head is responsible for class judgment and position localization based on the fused features. The separable self-attention module enhances the context information expression and at the same time enables the feature extractor to focus on the defect key areas.
[0020] It further includes the following steps:
[0021] Global average pooling and global max pooling are performed on the feature map generated by Layer4 of the backbone network ResNet50 through the Global branch to obtain 2048-dimensional feature vectors G_Avg and G_Max. Then, these two feature vectors are added to obtain a new feature vector G. Subsequently, the vector G is sent to the classifier to obtain a classification prediction score. The classification prediction score is first smoothed by label smoothing, and the formula is as follows:
[0022]
[0023] In the formula, ε is the smoothing parameter; K represents the total number of labels in the classification task; y is the true label; q k is the smoothed label; k is the original label;
[0024] Subsequently, the logits are converted into a probability distribution p through Softmax k :
[0025]
[0026] In the formula, K represents the total number of labels in the classification task; logits k is the original prediction score for the original label k; logits i is the original prediction score for the i-th label in the classification task; e is the natural constant;
[0027] When p k and q k are obtained, cross-entropy loss is used for loss calculation to obtain a loss value L that measures the difference between the model prediction and the smoothed label:
[0028]
[0029] Meanwhile, G_Avg and G_Max use triplet loss for loss calculation, and its calculation formula is as follows:
[0030]
[0031] In the formula, N represents the feature vector of the negative sample; D ia,ip represents the distance between the i-th anchor and the i-th positive sample; D ia,in represents the distance between the i-th anchor and the i-th negative sample; is a hyperparameter representing the minimum distance difference expected between the anchor-positive sample pair and the anchor-negative sample pair; L t is the loss value used to optimize the feature representation.
[0032] It also includes the following steps:
[0033] The features extracted by the Middle branch are the feature maps generated by Layer3 of ResNet50. Then, global average pooling GAP and global max pooling GMP are performed to generate 1024-dimensional feature vectors M_Avg and M_Max respectively. They are added to obtain a new 1024-dimensional vector M. M is fed into the classifier to obtain classification prediction scores, and the classification prediction scores are calculated using the Softmax cross-entropy loss after label smoothing; similar to the Global branch, M_Avg and M_Max use triplet loss for loss calculation;
[0034] The Part branch first uses 4 spatial attention network STNs to extract significant information regions, forming 4 new feature maps. Then, each feature map generates two 2048-dimensional feature vectors via global average pooling (GAP) and global max pooling (GMP) respectively, and the two vectors are superimposed to obtain a new 2048-dimensional vector. Finally, the new vector is fed into the classifier, and the classification prediction score is obtained through the classifier; similar to the Global branch, the Softmax cross-entropy loss after label smoothing is used for loss calculation.
[0035] It also includes the following steps:
[0036] For the backbone network in the cat ear defect recognition model, the input image first passes through a root convolutional layer to extract shallow features, and then passes through four feature extraction layers composed of C2f modules and convolutional modules. Among them, an SPPF module is also added to the last feature extraction layer to improve the model calculation efficiency;
[0037] For the neck network in the cat ear defect recognition model, the FPN network is used to increase the scale of the feature map through upsampling, and the feature map of the corresponding scale of the backbone network is concatenated to capture information of targets at different scales. And a data path from the bottom up is set in the PAN, so that upsampling is first performed on the low-resolution feature map, and then downsampling is performed on the high-resolution feature map, and the two are linked to construct a new path method to reduce information loss and retain richer details;
[0038] For the detection head network in the cat ear defect recognition model, the detection head network includes a classification branch and a regression branch, which are responsible for predicting the class label and the target box label in the feature map respectively. Each branch contains three basic convolutional blocks;
[0039] For the separable self-attention module in the cat ear defect recognition model, the separable self-attention module can separate the feature map into two parts. One part uses the self-attention mechanism for feature processing, and the other part retains the original information. Finally, the two parts are concatenated to obtain a complete feature map, which can extract information of the context defect region while retaining the original information.
[0040] The present invention solves the technical defects existing in the background technology, and the present invention has the following beneficial effects:
[0041] This device is highly intelligent. The entire device is controlled by a core board and has a deep learning model deployed inside, which can identify cats and whether they are sterilized. There is a light source inside the device, which can adjust the internal brightness independently. There is an automatic feeding device inside the device, which can feed food at regular intervals. The device also carries an ultrasonic repellent device. If a target object that does not meet the capture requirements stays inside the device for a long time, the ultrasonic wave will be activated for channeling. There is a constant temperature device inside the device. If extreme weather occurs, the internal equipment can be adjusted to provide stray cats with a suitable environment to survive the extreme weather. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, drawings of other embodiments can be obtained based on these drawings without paying creative work.
[0043] Figure 1 It is a structural schematic diagram of the device;
[0044] Figure 2 This is a simplified control flow chart of the device;
[0045] In the figure: 1. power supply; 2. motor; 3. feeding turntable; 4. servo; 5. revolving door; 6. camera; 7. main control board; 8. light source; 9. ultrasonic alarm; 10. temperature sensor; 11. temperature control equipment. DETAILED DESCRIPTION
[0046] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0047] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0048] like Figure 1 , 2 As shown, the first aspect of the present invention discloses a stray cat intelligent capture device based on deep learning, the capture device comprising:
[0049] An intelligent recognition module, the recognition module includes a camera 6 and a main control board 7, the camera 6 is connected to the main control board 7, and a recognition model is obtained by learning and training the characteristic image data, and the recognition model is input into the main control board 7, so that the recognition module can identify the breed of the cat entering the device and whether it has an ear tag during operation; wherein the characteristic image data includes the appearance of cats of different types and breeds and ear tag image data;
[0050] It should be noted that the capture device is provided with a camera 6, which is connected to the main control board 7. By inputting image data of the appearance of cats of various types and breeds and the important feature of ear tags into the core board, the core board learns and learns to identify the breed of the cat entering the device and whether it has an ear tag.
[0051] An intelligent driving module, wherein the driving module includes an ultrasonic alarm 9. If the recognition module identifies that the animal entering the device is a non-target animal, the ultrasonic alarm 9 is controlled to emit an alarm to drive away the non-target animal;
[0052] It should be noted that in order to prevent the capture of wrong objects or the device being occupied by wrong objects and affecting work, a driving device is provided inside the device. Once the camera 6 recognizes that the animal in the device is not the target animal to be captured, an alarm will be sounded to drive it away.
[0053] Automatic feeding module, which can deliver cat food into the device at regular intervals to prevent non-target animals from eating up all the food in the device and making it impossible for the device to continue to lure target stray cats;
[0054] It should be noted that the storage device for automatically delivering cat food in the device can deliver cat food into the device at a regular interval to prevent non-target animals from eating up all the food in the device and causing the device to be unable to continue to lure target stray cats.
[0055] The automatic sealing module includes a revolving door 5 and a steering engine 4, and the steering engine 4 is connected to the revolving door 5; when the recognition module identifies that the animal entering the device is a target stray cat, the steering engine 4 controls the revolving door 5 to close, completing the capture; and sends a signal to the cloud device to inform that the capture is completed, prompting the captured object to be processed in time;
[0056] It should be noted that after the camera 6 identifies the stray cat to be captured, the door of the capture device will be controlled to close automatically to complete the capture. A signal prompt is sent to the cloud device to inform that the capture is completed and to deal with the captured object in time.
[0057] The real-time monitoring and image transmission module can transmit the image captured by the camera 6 to the cloud device, and at the same time, through the real-time monitoring and image transmission module, the camera 6 can be remotely accessed to observe the real-time internal situation of the device;
[0058] It should be noted that for the convenience of observing the status of stray cats in real time and dealing with them in a timely manner, the camera 6 will regularly take pictures of stray cats and transmit them to a computer or other connected devices. At the same time, the movements of stray cats can also be observed in real time by remotely accessing the camera 6.
[0059] The temperature control module, the temperature control module includes a temperature sensor 10 and a temperature control device 11. The temperature control module is used in cooperation with the intelligent central control unit to achieve automatic temperature adjustment to improve the internal temperature environment of the device.
[0060] It should be noted that the temperature control module is used in cooperation with the intelligent central control unit to achieve automatic temperature adjustment and reduce manual intervention. Provide a suitable temperature for stray cats, reduce the stress and discomfort of stray cats, and prevent stray cats from suffering from hypothermia or heatstroke under extreme weather conditions.
[0061] Further, in a preferred embodiment of the present invention, the feeding module is arranged at the top of the device. The feeding module includes a feeding turntable 3 and a motor 2. The motor 2 drives the feeding turntable 3 to rotate to deliver cat food to the bottom of the device.
[0062] Further, in a preferred embodiment of the present invention, a power supply 1 is also arranged inside the device, and the power supply 1 provides power for each module inside the device.
[0063] It should be noted that the power required by the main control board 7 is provided by the power supply 1. The main control board 7 judges whether there is a target animal inside the device through the camera 6, obtains the internal temperature environment of the device through the temperature sensor 10, changes the internal temperature of the device through the temperature control device 11, drives away non-target animals through ultrasonic waves, controls the feeding turntable 3 to select and feed through the motor 2, and controls the rotation door 5 to open and close through the servo motor 4.
[0064] Further, in a preferred embodiment of the present invention, the trained recognition model is deployed in the main control board 7. The main control board 7 makes a judgment based on the data captured by the camera 6. If the animal entering the device is a non-target animal, it controls the ultrasonic alarm 9 to emit an alarm sound to drive away the non-target animal; if it recognizes that the animal entering the device is a target stray cat, it controls the rotation door 5 to close through the servo motor 4 to complete the capture.
[0065] It should be noted that the trained recognition model is deployed in the main control board 7. The main control board 7 makes a judgment based on the data captured by the camera 6. If it is an inappropriate target animal, ultrasonic waves are used to drive it away. There is a feeding device inside the device that feeds at regular intervals to attract stray cats into the device. During the execution process, there is a constant temperature device inside that can provide a suitable temperature.
[0066] The second aspect of the present invention discloses a training method for an identification module of a stray cat intelligent capture device. The recognition model includes a cat face recognition model and a cat ear defect recognition model;
[0067] The cat face recognition model fuses local and global features, uses ResNet50 as the backbone network to obtain feature data, and sets three branches according to the characteristics of the extracted features: the Middle branch, the Global branch, and the Part branch. The Middle branch extracts the intermediate-dimensional global features of the backbone network, the Global branch extracts high-dimensional global features, and the Part branch extracts local features of uniformly sized blocks. Finally, the extracted features are fused as the feature representation for cat face re-identification;
[0068] The cat ear defect recognition model includes a backbone network, a neck network, a detection head network, and a separable self-attention module; in the input processing stage, the image data first undergoes preprocessing steps of normalization and data augmentation; subsequently, the backbone network performs deep feature extraction on the preprocessed image; the neck network is to fuse the extracted feature maps to enhance the model's recognition ability for multi-scale targets; the detection head is responsible for class judgment and position localization based on the fused features; the separable self-attention module enhances the information expression of the context and at the same time enables the feature extractor to focus on the key defect areas.
[0069] It should be noted that the cat face recognition model fuses local and global features, uses ResNet50 as the backbone network to obtain feature data, and designs three branches according to the characteristics of the extracted features. The Middle branch extracts the intermediate-dimensional global features of the backbone network, the Global branch extracts high-dimensional global features, and the Part branch extracts local features of uniformly sized blocks. Finally, they are fused as the feature representation for cat face re-identification.
[0070] It also includes the following steps:
[0071] The feature maps generated by Layer4 of the backbone network ResNet50 are subjected to global average pooling (GAP, Global Average Pooling) and global max pooling (GMP, Global Max Pooling) through the Global branch to obtain the feature vectors G_Avg and G_Max with a dimension of 2048. Then, these two feature vectors are added together to obtain a new feature vector G. Subsequently, the vector G is fed into the classifier to obtain the classification prediction score. The classification prediction score is first smoothed by the label smoothing, and the formula is as follows:
[0072]
[0073] In the formula, ε is the smoothing parameter; K represents the total number of labels in the classification task; y is the true label; q k is the smoothed label; k is the original label;
[0074] Subsequently, the logits are converted into a probability distribution p through Softmax k :
[0075]
[0076] In the formula, K represents the total number of labels in the classification task; logits k is the original prediction score of the original label k; logits i is the original prediction score of the i-th label in the classification task; e is the natural constant;
[0077] When p k and q k are obtained, the cross-entropy loss is used for loss calculation to obtain a loss value L that measures the difference between the model prediction and the smoothed label:
[0078]
[0079] At the same time, G_Avg and G_Max are used for loss calculation using the triplet loss, and its calculation formula is as follows:
[0080]
[0081] In the formula, N represents the feature vector of the negative sample; D ia,ip represents the distance between the i-th anchor point and the i-th positive sample; D ia,in represents the distance between the i-th anchor point and the i-th negative sample; is a hyperparameter representing the desired minimum distance difference between the anchor-positive sample pair and the anchor-negative sample pair; L t is the loss value used to optimize the feature representation.
[0082] It also includes the following steps:
[0083] The features map generated by Layer3 of ResNet50 is extracted by the Middle branch, and then global average pooling GAP and global max pooling GMP are performed to generate 1024-dimensional feature vectors M_Avg and M_Max respectively. The two vectors are added to obtain a new 1024-dimensional vector M. M is fed into the classifier to obtain a classification prediction score, and the loss of the classification prediction score is calculated by the Softmax cross-entropy loss after label smoothing; Similar to the Global branch, the triple loss is used to calculate the loss of M_Avg and M_Max.
[0084] It should be noted that the process of the Middle branch is similar to that of the Global branch. However, the features map generated by Layer3 of ResNet50 is extracted by the Middle branch, and then global average pooling GAP and global max pooling GMP are performed to generate 1024-dimensional feature vectors M_Avg and M_Max respectively. The two vectors are added to obtain a new 1024-dimensional vector M. M is fed into the classifier to obtain a classification prediction score, and the loss of the classification prediction score is calculated by the Softmax cross-entropy loss after label smoothing. Similarly, the triple loss is used to calculate the loss of M_Avg and M_Max.
[0085] The Part branch first uses 4 spatial attention networks STN to extract significant information regions, forming 4 new features maps. Then each features map is respectively passed through global average pooling GAP and global max pooling GMP to generate two 2048-dimensional feature vectors, and the two vectors are stacked to obtain a new 2048-dimensional vector. Finally, the new vector is fed into the classifier, and the classification prediction score is obtained through the classifier; Similar to the Global branch, the Softmax cross-entropy loss after label smoothing is used to calculate the loss.
[0086] It should be noted that the Part branch first uses 4 spatial attention networks STN to extract significant information regions, forming 4 new features maps. Then each features map is respectively passed through global average pooling GAP and global max pooling GMP to generate two 2048-dimensional feature vectors, and the two vectors are stacked to obtain a new 2048-dimensional vector. Finally, the new vector is fed into the classifier, and the classification prediction score is obtained through the classifier. Similar to the loss calculation method of the Global layer, the Softmax cross-entropy loss after label smoothing is used to calculate the loss.
[0087] It also includes the following steps:
[0088] For the backbone network in the cat ear defect recognition model, the input image first passes through a root convolutional layer to extract shallow features, and then through four feature extraction layers composed of C2f modules and convolutional modules. Among them, an SPPF module is also added to the last feature extraction layer to improve the model's calculation efficiency;
[0089] For the neck network in the cat ear defect recognition model, the FPN network is used to increase the scale of the feature map through upsampling, and the feature maps of the corresponding scales of the backbone network are concatenated to capture information of targets at different scales. A data path from the bottom up is set in the PAN to first upsample from the low-resolution feature map, then downsample from the high-resolution feature map, and link the two to construct a new path method to reduce information loss and retain richer details;
[0090] For the detection head network in the cat ear defect recognition model, the detection head network includes a classification branch and a regression branch, which are responsible for predicting the class label and the target box label in the feature map respectively. Each branch contains three basic convolutional blocks;
[0091] For the separable self-attention module in the cat ear defect recognition model, the separable self-attention module can separate the feature map into two parts. One part uses the self-attention mechanism for feature processing, and the other part retains the original information. Finally, these two parts are concatenated to obtain a complete feature map, which can increase the information extraction of the context defect area while retaining the original information.
[0092] It should be noted that the cat ear defect recognition model is improved based on the YOLOv8 object detector. The entire algorithm can be divided into four parts, namely the backbone network, the neck network, the detection head network, and the separable self-attention module. In the input processing stage, the image data first undergoes preprocessing steps such as normalization and data augmentation; subsequently, the backbone network performs deep feature extraction on these preprocessed images; the neck network mainly fuses the extracted feature maps to enhance the model's ability to recognize multi-scale targets; the detection head is responsible for accurate class judgment and position localization based on these fused features; the separable self-attention module enhances the information expression of the context and at the same time enables the feature extractor to focus on the key defect areas.
[0093] The backbone network of YOLOv8 references the CSPDarkNet-53 network of YOLOv5 and is improved on this basis, using the C2f module instead of the C3 module. The C2f module mainly separates the input feature channels into two parts. One part retains the original feature information, and the other part mines deep information through the bottleneck layer. This not only reduces the computational dimension and training time but also improves the robustness of the model. Specifically, the input image first passes through a root convolutional layer to extract shallow features, and then through four feature extraction layers composed of C2f modules and convolutional modules. Among them, the SPPF module is also added to the last feature extraction layer to improve the model's computational efficiency.
[0094] YOLOv8, which adopts the PAN-FPN idea in the neck network, combines FPN and PAN for multi-scale feature fusion. First, the large stride (such as 16 times) of the final feature map in the backbone network may lead to a relatively low best recall rate. Second, the overlapping of objects in the ground truth boxes may lead to intractable ambiguities, that is, how to determine the position of the bounding box in the overlapping regression. Therefore, using the FPN network to increase the scale of the feature map through upsampling and stitching the corresponding scale feature maps of the backbone network is beneficial to capturing information of objects at different scales, while PAN further improves the feature fusion effect on the basis of FPN. Compared with FPN, PAN adds a new data path from the bottom up. The original design intention of PAN is to first upsample from the low-resolution feature map, then downsample from the high-resolution feature map, and link the two to build a new path way. Different from the simple addition method of FPN, PAN concatenates and combines feature maps of different levels, aiming to reduce information loss and retain richer details, thereby significantly improving the accuracy of object detection.
[0095] To improve the expression clarity and logic, YOLOv8 abandons the previous design and instead adopts a decoupled detection head and an anchor box-based sample matching method to improve detection accuracy. The new detection head network includes a classification branch and a regression branch, which are responsible for predicting the class label and the target box label in the feature map respectively. Each branch contains three basic convolutional blocks. Through this design adjustment, YOLOv8 achieves more accurate object detection and has a significant improvement in performance.
[0096] To enable the model to pay more attention to the semantic information extraction of the defect area, a separate self-attention module is added to the feature maps of different scales in the backbone network. Since the computational cost of the self-attention mechanism is relatively high, this module refers to the idea of the C2f module, separating the feature map into two parts. One part uses the self-attention mechanism for feature processing, and the other part retains the original information. Finally, these two parts are concatenated to obtain a complete feature map, which can increase the information extraction of the context defect area while retaining the original information, enabling the model to pay more attention to the defect area without consuming a large amount of training time.
[0097] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings between the various components shown or discussed, or direct couplings, or communication connections can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical or other forms.
[0098] The units described as separate components above may or may not be physically separated. The components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0099] In addition, in each embodiment of the present invention, the functional units can all be integrated in one processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0100] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: mobile storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks or optical disks and other various media that can store program codes.
[0101] Alternatively, if the above-integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention essentially or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as a removable storage device, ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0102] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An intelligent stray cat capture device based on deep learning, characterized in that, The capturing device comprises: The recognition module includes a camera and a main control board, the camera is connected to the main control board, and a recognition model is obtained by learning and training the characteristic image data, and the recognition model is input into the main control board, so that the recognition module can identify the breed of the cat entering the device and whether it has an ear tag during operation; wherein the characteristic image data includes the appearance of cats of different types and breeds and the ear tag image data; A driving module, wherein the driving module includes an ultrasonic alarm. If the recognition module identifies that the animal entering the device is a non-target animal, the ultrasonic alarm is controlled to emit an alarm to drive the non-target animal away. Automatic feeding module, which can deliver cat food into the device at regular intervals to prevent non-target animals from eating up all the food in the device and making it impossible for the device to continue to lure target stray cats; The automatic sealing module includes a revolving door and a steering gear, and the steering gear is connected to the revolving door; when the recognition module identifies that the animal entering the device is a target stray cat, the steering gear controls the revolving door to close, completing the capture; and sends a signal to the cloud device to inform that the capture is completed, prompting timely processing of the captured object; Real-time monitoring and image transmission module, through which the images captured by the camera can be transmitted to the cloud device, and the camera can be remotely accessed through the real-time monitoring and image transmission module to observe the real-time image situation inside the device; A temperature control module, which includes a temperature sensor and a temperature control device. The temperature control module is used in conjunction with an intelligent central control unit to achieve automatic temperature regulation to submit the internal temperature environment of the device; It also includes a training method for a recognition module of the stray cat intelligent capture device, wherein the recognition model includes a cat face recognition model and a cat ear defect recognition model; The cat face recognition model combines local and global features, uses ResNet50 as the backbone network to obtain feature data, and sets three branches according to the characteristics of the extracted features: Middle branch, Global branch and Part branch. The Middle branch extracts the global features of the intermediate dimension of the backbone network, the Global branch extracts the global features of high dimensions, and the Part branch extracts the local features of uniformly divided blocks of size. Finally, the extracted features are fused as the feature representation for cat face re-recognition. The cat ear defect recognition model includes a backbone network, a neck network, a detection head network and a separable self-attention module; in the input processing stage, the image data first undergoes normalization and data enhancement preprocessing steps; then, the backbone network performs deep feature extraction on the preprocessed image; the neck network performs feature fusion on the extracted feature map to enhance the model's recognition ability for multi-scale targets; the detection head is responsible for category judgment and position positioning based on the fused features; the separable self-attention module enhances the expression of contextual information and enables the feature extractor to focus on key defect areas.
2. The intelligent stray cat capture device based on deep learning according to claim 1, wherein: The automatic feeding module is arranged at the top of the device. The automatic feeding module includes a feeding turntable and a motor. The motor drives the feeding turntable to rotate to deliver cat food to the bottom of the device.
3. The intelligent stray cat capturing device based on deep learning according to claim 1, wherein: A power supply is also arranged inside the device, and the power supply provides power for each module inside the device.
4. The intelligent stray cat capturing device based on deep learning according to claim 1, wherein: The trained recognition model is deployed into the main control board. The main control board makes a judgment based on the data captured by the camera. If the animal entering the device is a non-target animal, it controls the ultrasonic alarm to emit an alarm sound to drive away the non-target animal; if it is recognized that the animal entering the device is a target stray cat, the rotating door is controlled by the servo to close to complete the capture.
5. The intelligent stray cat capture device based on deep learning according to claim 1, wherein: The feature map generated by Layer4 of the backbone network ResNet50 is subjected to global average pooling and global max pooling through the Global branch to obtain 2048-dimensional feature vectors G Avg and G Max. Then, these two feature vectors are added to obtain a new feature vector G. Subsequently, the vector G is sent to the classifier to obtain a classification prediction score. The classification prediction score is first smoothed by labels, and the formula is as follows: , Where ε is the smoothing parameter; K represents the total number of labels in the classification task; y is the true label; q k is the smoothed label; k is the original label; Subsequently, the logits are converted into a probability distribution p through Softmax k : , where K represents the total number of labels in the classification task; logits k is the original prediction score for the original label k; logits i is the original prediction score for the i-th label in the classification task; e is the natural constant; When p is obtained k and q k are obtained, cross-entropy loss is used for loss calculation to obtain a loss value L that measures the difference between the model prediction and the smoothed label: , At the same time, G Avg and G Max use triplet loss for loss calculation, and its calculation formula is as follows: , Where N represents the feature vector of negative samples; D ia,ip represents the distance between the i-th anchor point and the i-th positive sample; D ia,in represents the distance between the i-th anchor point and the i-th negative sample; ∂ is a hyperparameter representing the minimum distance difference expected between the anchor-positive sample pair and the anchor-negative sample pair; L t is the loss value used to optimize the feature representation.
6. The intelligent stray cat capture device based on deep learning according to claim 5, wherein: The feature map generated by Layer3 of ResNet50 is extracted by the Middle branch, and then global average pooling GAP and global max pooling GMP are performed to generate 1024-dimensional feature vectors M Avg and M Max respectively. The two are added to obtain a new 1024-dimensional vector M. The M is sent to the classifier to obtain a classification prediction score, and the classification prediction score is calculated by the Softmax cross-entropy loss after label smoothing; the same as the Global branch, M Avg and M Max use triplet loss for loss calculation; The Part branch first uses 4 spatial attention networks STN to extract significant information regions to form 4 new feature maps. Then, each feature map is respectively subjected to global average pooling GAP and global max pooling GMP to generate two 2048-dimensional feature vectors, and the two vectors are superimposed to obtain a 2048-dimensional new vector. Finally, the new vector is sent into the classifier, and the classification prediction score is obtained through the classifier; the same as the Global branch, the Softmax cross-entropy loss after label smoothing is used for loss calculation.
7. The intelligent stray cat capture device based on deep learning according to claim 1, wherein: For the backbone network in the cat ear defect recognition model, the input image first passes through a root convolutional layer to extract shallow features, and then passes through four feature extraction layers composed of C2f modules and convolutional modules. Among them, an SPPF module is also added to the last feature extraction layer to improve the model calculation efficiency; For the neck network in the cat ear defect recognition model, the FPN network is used to increase the scale of the feature map through upsampling, and the feature maps of the corresponding scales of the backbone network are concatenated to capture the information of targets at different scales. A data path from the bottom up is set in the PAN so that upsampling is first performed on the low-resolution feature map, then downsampling is performed on the high-resolution feature map, and the two are linked to construct a new path method to reduce information loss and retain richer details; For the detection head network in the cat ear defect recognition model, the detection head network includes a classification branch and a regression branch, which are responsible for predicting the class label and the target box label in the feature map respectively. Each branch contains three basic convolutional blocks; For the separable self-attention module in the cat ear defect recognition model, the separable self-attention module can separate the feature map into two parts. One part uses the self-attention mechanism for feature processing, and the other part retains the original information. Finally, the two parts are concatenated to obtain a complete feature map, which can increase the information extraction of the context defect area while retaining the original information.
Citation Information
Patent Citations
Deep learning-based cow face re-recognition method, system and device, and medium
CN113989836A
Method for feeding stray animals and method for interacting target animals
CN117814176A
Capture device and method
CN118000185A
Lightweight target detection network design method for indoor semantic SLAM system
CN118097365A