Household pet identification method and system, terminal and medium

By adjusting the structure of YOLOv7 neural network, combining attention mechanism and decoupling head detection, the problem of poor recognition difficulty and adaptability in pet image recognition technology is solved, and efficient and accurate pet recognition effect is achieved.

CN120014669APending Publication Date: 2025-05-16SICHUAN ENRISING INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510069976.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When existing pet image recognition technology processes pet images from different environments and shooting angles, there are problems such as increased recognition difficulty, difficulty in feature extraction and matching, and poor model adaptability.

Method used

The adjusted YOLOv7 neural network is used as the basic model. By connecting the GAM attention mechanism layer at the BackBone network output, the SA attention mechanism layer is connected at the Neck network layer output, and the detection head is set as the Decoupled_Detect decoupling head, combining Shuffle_Block and GSConv to reduce the network complexity.

Benefits of technology

It improves the accuracy and efficiency of pet identification, adapts to the lightweight needs of deploying in home smart terminals, and enhances the model's ability to identify different pet breeds and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014669A_ABST
    Figure CN120014669A_ABST
Patent Text Reader

Abstract

The invention discloses a home pet recognition method and system, a terminal and a medium, and relates to the field of image recognition, and the key points of the technical scheme are as follows: obtaining a to-be-recognized pet image; recognizing a pet in the to-be-recognized pet image based on a pre-trained pet recognition model to obtain a recognition result of the home pet; the method comprises the following steps: performing deep learning training on image data of different types of pets to obtain a pet recognition model; wherein the deep learning training is completed through the YOLOv7 neural network after the network structure is adjusted, and the mode of adjusting the network structure of the YOLOv7 neural network comprises the steps that the output of a BackBone network of the YOLOv7 neural network is connected with a GAM attention mechanism layer, and the output of a Neck network layer of the YOLOv7 neural network is connected with an SA attention mechanism layer. A detection head of the YOLOv7 neural network is set as a DecoupledDeect decoupling head, and the SA attention mechanism layer is connected with the DecoupledDeect decoupling head.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and more specifically, to a method, system, terminal and medium for identifying household pets. Background Art

[0002] Pets have become an indispensable part of people's daily lives, and more and more people choose to keep pets. With the popularization of smart home technology and the continuous improvement of people's living standards, how to manage and care for pets more efficiently and intelligently has become an urgent problem to be solved. Therefore, the research and application of pet recognition technology has become particularly important, which can help people better interact with pets and provide more intelligent pet management solutions.

[0003] The existing pet image recognition technology has the following problems: First, due to the large difference in the distance between the pet and the camera, the scale of the pet in the image varies significantly, which makes recognition more difficult. In different environments and shooting angles, the appearance of the pet may change drastically, posing a great challenge to recognition; second, the actual scenes of pet image recognition are usually more complex, often with factors such as occlusion and variable target posture, which makes it difficult to extract and match the pet features in the image; and pet recognition technology usually requires a large amount of accurately labeled data for training, but existing public data sets often cannot cover all breeds, postures and environmental conditions, resulting in poor adaptability of the recognition model. Summary of the invention

[0004] The purpose of the present invention is to provide a household pet identification method, system, terminal and medium, which solves the scale change problem caused by the distance difference between the pet and the camera, thereby improving the accuracy of identification.

[0005] A first aspect of the present invention provides a method for identifying a household pet, the method comprising:

[0006] Obtain a pet image to be identified;

[0007] Based on a pre-trained pet recognition model, pets in the pet image to be recognized are recognized to obtain recognition results of household pets; wherein, deep learning training is performed on image data of different types of pets to obtain a pet recognition model; wherein, the deep learning training is completed by adjusting the YOLOv7 neural network after the network structure, wherein the method of adjusting the network structure of the YOLOv7 neural network includes: connecting a GAM attention mechanism layer to the output of the BackBone network of the YOLOv7 neural network, connecting an SA attention mechanism layer to the output of the Neck network layer of the YOLOv7 neural network, setting the detection head of the YOLOv7 neural network to a Decoupled_Detect decoupling head, and connecting the SA attention mechanism layer to the Decoupled_Detect decoupling head.

[0008] In one implementation, the method of adjusting the network structure of the YOLOv7 neural network also includes: replacing the BackBone network of the YOLOv7 neural network with a Shuffle_Block network, replacing the Conv layer in the Neck network layer of the YOLOv7 neural network with a GSConv layer, and connecting the GAM attention mechanism layer to the GSConv layer.

[0009] In one implementation, the Shuffle_Block network has multiple Shuffle Block modules stacked in sequence, and the output of the last Shuffle Block module serves as the input of the GAM attention mechanism layer.

[0010] In one implementation, the number of the Decoupled_Detect decoupling heads is 3, wherein the inputs of two Decoupled_Detect decoupling heads are the outputs of two C3 modules of the Neck network layer, and the input of one Decoupled_Detect decoupling head is the output of the SA attention mechanism layer.

[0011] A second aspect of the present invention provides a household pet identification system, the system comprising:

[0012] An image acquisition module, used for acquiring an image of a pet to be identified;

[0013] An image recognition module is used to identify pets in a pet image to be identified based on a pre-trained pet recognition model to obtain recognition results of household pets; wherein, deep learning training is performed on image data of different types of pets to obtain a pet recognition model; wherein the deep learning training is completed by adjusting the YOLOv7 neural network after the network structure is adjusted, wherein the method of adjusting the network structure of the YOLOv7 neural network includes: connecting a GAM attention mechanism layer to the output of the BackBone network of the YOLOv7 neural network, connecting an SA attention mechanism layer to the output of the Neck network layer of the YOLOv7 neural network, setting the detection head of the YOLOv7 neural network to a Decoupled_Detect decoupling head, and connecting the SA attention mechanism layer to the Decoupled_Detect decoupling head.

[0014] In one implementation, the method of adjusting the network structure of the YOLOv7 neural network also includes: replacing the BackBone network of the YOLOv7 neural network with a Shuffle_Block network, replacing the Conv layer in the Neck network layer of the YOLOv7 neural network with a GSConv layer, and connecting the GAM attention mechanism layer to the GSConv layer.

[0015] In one implementation, the Shuffle_Block network has multiple Shuffle Block modules stacked in sequence, and the output of the last Shuffle Block module serves as the input of the GAM attention mechanism layer.

[0016] In one implementation, the number of the Decoupled_Detect decoupling heads is 3, wherein the inputs of two Decoupled_Detect decoupling heads are the outputs of two C3 modules of the Neck network layer, and the input of one Decoupled_Detect decoupling head is the output of the SA attention mechanism layer.

[0017] According to a third aspect of the present invention, a smart pet terminal is provided, comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of a household pet identification method provided in the first aspect of the present invention are implemented.

[0018] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of a household pet identification method provided in the first aspect of the present invention are implemented.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] In a household pet identification method provided by the present invention, YOLOv7 is used as the basic neural network for pet identification, and the GAM and SA attention mechanisms are combined to optimize the recognition accuracy of the pet. At the same time, Shuffle_Block and GSConv are used to reduce the complexity of the YOLOv7 neural network, improve the recognition efficiency, and adapt to the lightweight requirements of deployment in home smart terminals. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:

[0022] Figure 1 A schematic diagram of a process flow of a household pet identification method provided by an embodiment of the present invention;

[0023] Figure 2 A schematic diagram of the structure of a YOLOv7 neural network after adjusting the network structure provided in an embodiment of the present invention;

[0024] Figure 3 A schematic diagram of the structure of the GAM attention mechanism layer provided by an embodiment of the present invention;

[0025] Figure 4 A schematic diagram of the structure of the SA attention mechanism layer provided in an embodiment of the present invention;

[0026] Figure 5 A schematic diagram of the structure of the Shuffle_Block module provided in an embodiment of the present invention;

[0027] Figure 6 A schematic diagram of the structure of the GSConv layer provided in an embodiment of the present invention;

[0028] Figure 7 A network structure diagram of Decoupled_Detect provided in an embodiment of the present invention;

[0029] Figure 8 A functional block diagram of a household pet identification system provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.

[0031] It should be noted that the terms "include" or "may include" used in various embodiments of the present application indicate the presence of the function, operation or element applied for, and do not limit the addition of one or more functions, operations or elements. In addition, as used in various embodiments of the present application, the terms "include", "have" and their cognates are only intended to indicate specific features, numbers, steps, operations, elements, components or a combination of the foregoing items, and should not be understood as first excluding the presence of one or more other features, numbers, steps, operations, elements, components or a combination of the foregoing items or the possibility of adding one or more features, numbers, steps, operations, elements, components or a combination of the foregoing items.

[0032] It should be understood that terms such as "first" and "second" are only used for descriptive purposes and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0033] Please refer to Figure 1 , Figure 1 A schematic diagram of a process for identifying a household pet provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes:

[0034] S101, obtaining a pet image to be identified.

[0035] In this embodiment, the pet image to be identified can be realized through a home pet terminal with a camera module, such as an automatic feeding terminal, a drinking water terminal, etc., which is currently a common and mature terminal, and is not specifically described in this embodiment. In the home pet terminal, a camera connected to an embedded device captures the video stream in the home environment, and the input frame rate of the camera is set to 35fps to balance the detection accuracy and energy consumption. In order to optimize the image quality, the video stream is first pre-processed by a hardware image signal processing unit (ISP) to enhance the clarity and details of the image, ensuring efficient detection and low-power operation, which is suitable for continuous monitoring in a home environment.

[0036] S102, based on the pre-trained pet recognition model, the pet in the pet image to be recognized is recognized to obtain the recognition result of the household pet; wherein, the image data of different types of pets are deep-learned and trained to obtain the pet recognition model; wherein, the deep-learning training is completed by adjusting the YOLOv7 neural network after the network structure, wherein the method of adjusting the network structure of the YOLOv7 neural network includes: connecting a GAM attention mechanism layer to the output of the BackBone network of the YOLOv7 neural network, connecting an SA attention mechanism layer to the output of the Neck network layer of the YOLOv7 neural network, setting the detection head of the YOLOv7 neural network to a Decoupled_Detect decoupling head, and connecting the SA attention mechanism layer to the Decoupled_Detect decoupling head.

[0037] In this embodiment, the crawler technology is used to obtain the original images of pets of different pet breeds in an indoor environment, and the collected image data is cleaned and processed, and the duplicate images and images that do not meet the requirements are eliminated, and then the image data is enhanced to expand the data set. The data of the original pet images is enhanced to improve the resolution, and the pet image data set is annotated according to the PascalVOC data format. The image format is unified as jpg format, and the input model resolution size is 640*640. There are a total of 8 types of household pets in the image, namely cats, dogs, fish, snakes, turtles, rabbits, hamsters, and parrots. After the annotation is completed, the data set is divided into a training set, a test set, and a verification set according to a ratio of 8:1:1. Among them, the PascalVOC image data uses the LABELIMG tool to select the pet image area to be detected in the shape of a rectangular recognition frame, annotate the position and breed of the pet in the recognition area, and save the corresponding annotation frame xml file based on this, and name the detection image as 000001.jpg and 000002.jpg. The image naming sequence increases in sequence and corresponds to the annotation frame xml file one by one. The cleaned and pre-processed images were enhanced with basic data enhancement work such as vertical flip, horizontal flip, translation, scaling, and cropping, and then Mosaic data enhancement was performed. A total of 180,384 pet images were obtained, of which 144,307 were used for training and 18,038 were used for testing. The main hyperparameters input before training were the sliding average decay rate of 0.9995, the judgment threshold of 0.65, the number of anchor boxes at each scaling ratio of 3, the number of sample batches BATCH_SIZE of 12, the initial learning rate of 0.0005, and the stable learning rate of 0.00001.

[0038] The main problem facing household pet target recognition now is that traditional image processing methods have poor adaptability to complex environments. Traditional methods usually rely on manually designed features such as edges, corners, and colors, which makes them easily fail in the case of lighting changes, occlusions, or complex backgrounds. For small pets that move quickly, the recognition accuracy of traditional algorithms is often not high and cannot accurately track the target. In addition, traditional algorithms have weak adaptability to different environments and have difficulty handling various transformations or changes in pet postures, resulting in poor recognition effects in dynamic environments. In contrast, deep learning-based methods can better cope with these challenges by automatically learning multi-level features.

[0039] Therefore, this embodiment uses the deep learning neural network YOLOv7 algorithm model to perform pet recognition. Figure 2 As shown in the figure, YOLOv7 includes BackBone network layer, Neck network layer and detection head. In order to meet the requirements of small target recognition of neural network, a GAM attention mechanism layer is added to the BackBone backbone network layer, and an SA attention mechanism layer is added to the Neck network layer to improve the recognition accuracy of the pet due to the long distance between the pet and the camera. It should be noted that for the remaining network structures of the BackBone network layer, Neck network layer and detection head that have not been replaced, they are consistent with the basic YOLOv7 neural network.

[0040] The network structure diagram of Decoupled_Detect is as follows Figure 7 As shown in the figure, the detection head is set to Decoupled_Detect, which can decompose the target detection task into two subtasks: target classification and target localization. The former is responsible for predicting the category information of the target, and the latter is responsible for predicting the bounding box position information of the target. The target classification branch usually uses a classification network to predict the category probability distribution of the target; the target localization branch uses a regression network to predict the bounding box position of the target.

[0041] The GAM attention mechanism layer is added to the BackBone backbone network. By introducing the global perception mechanism and multi-scale attention mechanism, the model's ability to perceive and distinguish small targets is enhanced. The structural diagram of the GAM attention mechanism layer is as follows: Figure 3As shown in the figure, the global perception module (Channel attention module) in the GAM attention mechanism layer is used to capture the global context information of the image. By introducing global average pooling and fully connected layers, the entire feature map is globally perceived, thereby improving the perception of small targets under the pet camera; the multi-scale attention module (Spatial attention module) uses multiple parallel attention branches to weight the feature maps of different scales respectively, multiply the feature maps of different scales with the corresponding attention weights, and obtain the weighted feature representation, which can better distinguish targets of different scales and enhance the detection ability of small targets.

[0042] Adding the SA attention mechanism to the Neck network layer can enhance the network's perception and accuracy of small targets. The SA attention mechanism introduces cross-attention in the channel and spatial dimensions, and specifically adjusts the channel relationship of the feature map, thereby improving the ability to locate and classify small targets. The SA attention mechanism network structure diagram is shown in the figure below. Figure 4 As shown, SA attention first converts the feature map Divide into G groups, X=[X1,...,X G ], Each X k Gradually capture specific semantic responses during training. Then generate corresponding coefficients for each sub-feature through the attention module. After the channel attention mechanism, the global average pooling (GAP) operation is used to embed global information. As shown in formula (1):

[0043] in, H×W represents the spatial dimension, X k1 Represents the reduction factor.

[0044] Then the sigmoid activation function is used to aggregate information to generate a compact feature map. As shown in formula (2):

[0045] X' k1 =σ(F c (s))·X k1 =σ(W1s+b1)·X k1 (2), where X′ k1 represents the final output of channel attention, and are parameters used for scaling and moving.

[0046] The spatial attention focuses on the object information "position". The spatial attention output X k2 Perform GroupNorm (GN) operation to obtain spatial data information. c (·) Strengthen The calculation process of this step is expressed as formula (3):

[0047] X' k2 =σ(W2·GN(X k2 )+b2)·X k2 (3), where X' k2 represents the final output of spatial attention, GN() represents the group norm used to obtain spatial statistics, and W2 and b2 are in the range of Parameters.

[0048] Finally, the information of these two branches is aggregated through channel fusion, as shown in formula (4): Among them, X′ k Represents the final output of the SA attention mechanism. After the above operations, all sub-features use the channel shuffle operation in the channel dimension to fuse cross-group information.

[0049] To meet the requirements of lightweight learning detection of neural networks, the backbone network is replaced with Shuffle_Block, the neck network Conv convolution layer is replaced with GSConv, and the detection head is set to Decoupled_Detect decoupled head detection to reduce network calculation parameters and reduce the size of the network model.

[0050] In some embodiments, the method of adjusting the network structure of the YOLOv7 neural network also includes: replacing the BackBone network of the YOLOv7 neural network with a Shuffle_Block network, replacing the Conv layer in the Neck network layer of the YOLOv7 neural network with a GSConv layer, and connecting the GAM attention mechanism layer to the GSConv layer.

[0051] Specifically, the main purpose of replacing the backbone network with Shuffle_Block is to enhance the network's feature expression ability, reduce the number of parameters and computational complexity. The network structure of the Shuffle_Block module is as follows: Figure 5 As shown in Figure 1, the core idea is to introduce a channel rearrangement operation, which realizes information interaction and cross-channel feature fusion by rearranging the channels in the input feature map. The improved network can enhance the diversity and richness of features and improve the accuracy and stability of small target detection.

[0052] like Figure 2 As shown in FIG, the Shuffle_Block network has multiple Shuffle Block modules stacked in sequence, and the output of the last Shuffle Block module is used as the input of the GAM attention mechanism layer. Figure 2The connection relationship shown in the three parts of BackBone, Neck and Head is the connection relationship of the basic YOLOv7 neural network itself. This embodiment only replaces the corresponding network and does not change the connection relationship of the existing basic YOLOv7 neural network.

[0053] like Figure 6 As shown in the figure, in small target detection, targets usually have small size and low intensity features. After replacing the Conv layer of the Neck network of YOLOv7 with GSConv, the diversity and richness of features can be increased, the model's ability to perceive small targets such as household pets at a long distance captured by the camera can be enhanced, and the detection accuracy and robustness of small targets can be improved.

[0054] like Figure 2 As shown, in some embodiments, the number of the Decoupled_Detect decoupling heads is 3, wherein the inputs of two Decoupled_Detect decoupling heads are respectively the outputs of two C3 modules of the Neck network layer, and the input of one Decoupled_Detect decoupling head is the output of the SA attention mechanism layer.

[0055] Secondly, the present invention combines deep learning, image processing and embedded technology to provide an efficient and accurate household pet recognition system with the following technical effects: pet images are acquired through crawler technology, and data cleaning and enhancement are performed to improve the diversity and robustness of the data set. The enhanced data improves the model's ability to recognize different pet breeds, especially in the case of poor image quality, it can still maintain a high recognition accuracy. YOLOv7 is used as the basic detection model for pet target recognition, and the recognition accuracy of small targets is optimized in combination with GAM and SA attention mechanisms. At the same time, Shuffle_Block and GSConv are used to reduce network complexity, improve detection speed and efficiency, and meet lightweight requirements.

[0056] By deploying the model on an embedded device and processing the video stream input from the camera in real time, efficient detection and low-power operation are ensured, which is suitable for continuous monitoring in a home environment. The pet intelligent recognition system provides three detection methods: local image, camera, and video stream. Users can easily set model parameters and view recognition results through an intuitive interface, which improves the ease of use and operation experience of the system. The system can monitor pet behavior in real time, detect abnormal activities and upload them to the cloud, provide real-time feedback, and enhance pet safety monitoring functions. The recognition results are transmitted back through the SRS push-pull stream server, which supports remote access and has high scalability, making it easy to integrate other smart devices.

[0057] Please refer to Figure 8 , Figure 8The present invention provides a functional block diagram of a household pet identification system, which includes:

[0058] An image acquisition module 810 is used to acquire an image of a pet to be identified;

[0059] The image recognition module 820 is used to identify the pet in the pet image to be identified based on a pre-trained pet recognition model to obtain the recognition result of the household pet; wherein, deep learning training is performed on the image data of different types of pets to obtain the pet recognition model; wherein, the deep learning training is completed by adjusting the YOLOv7 neural network after the network structure, wherein the method of adjusting the network structure of the YOLOv7 neural network includes: connecting a GAM attention mechanism layer to the output of the BackBone network of the YOLOv7 neural network, connecting an SA attention mechanism layer to the output of the Neck network layer of the YOLOv7 neural network, setting the detection head of the YOLOv7 neural network to a Decoupled_Detect decoupling head, and connecting the SA attention mechanism layer to the Decoupled_Detect decoupling head.

[0060] Accordingly, in a household pet identification system provided by an embodiment of the present invention, YOLOv7 is used as the basic neural network for pet identification, and the GAM and SA attention mechanisms are combined to optimize the pet identification accuracy. At the same time, Shuffle_Block and GSConv are used to reduce the complexity of the YOLOv7 neural network, improve the recognition efficiency, and adapt to the lightweight requirements of deployment in home smart terminals.

[0061] In some embodiments, the method of adjusting the network structure of the YOLOv7 neural network also includes: replacing the BackBone network of the YOLOv7 neural network with a Shuffle_Block network, replacing the Conv layer in the Neck network layer of the YOLOv7 neural network with a GSConv layer, and connecting the GAM attention mechanism layer to the GSConv layer.

[0062] In some embodiments, the Shuffle_Block network has multiple Shuffle Block modules stacked in sequence, and the output of the last Shuffle Block module serves as the input of the GAM attention mechanism layer.

[0063] In some embodiments, the number of the Decoupled_Detect decoupling heads is 3, wherein the inputs of two Decoupled_Detect decoupling heads are the outputs of two C3 modules of the Neck network layer respectively, and the input of one Decoupled_Detect decoupling head is the output of the SA attention mechanism layer.

[0064] The embodiment of the present application also provides a smart pet terminal. The smart pet terminal includes a processor, a memory, a communication interface, and at least one communication bus for connecting the processor, the memory, and the communication interface. The memory includes but is not limited to a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (PROM), or a portable CD-ROM, and the memory is used for related instructions and data.

[0065] The communication interface is used to receive and send data. The processor can be one or more CPUs. When the processor is a CPU, the CPU can be a single-core CPU or a multi-core CPU. The processor in the smart pet terminal is used to read one or more programs stored in the memory and perform the following operations: obtaining a pet image to be identified; identifying the pet in the pet image to be identified based on a pre-trained pet recognition model to obtain a recognition result of a household pet; wherein, deep learning training is performed on image data of different types of pets to obtain a pet recognition model; wherein, the deep learning training is completed by adjusting the network structure of the YOLOv7 neural network, wherein the method of adjusting the network structure of the YOLOv7 neural network includes: connecting a GAM attention mechanism layer to the output of the BackBone network of the YOLOv7 neural network, connecting an SA attention mechanism layer to the output of the Neck network layer of the YOLOv7 neural network, setting the detection head of the YOLOv7 neural network to a Decoupled_Detect decoupling head, and connecting the SA attention mechanism layer to the Decoupled_Detect decoupling head.

[0066] It should be noted that the specific implementation of each operation can be Figure 1 The corresponding description of the method embodiment shown, the smart pet terminal can be used to execute a household pet identification method of the above method embodiment of the present application, which will not be described in detail here.

[0067] In the embodiments of the present disclosure, a computer-readable storage medium is also provided, and the computer-readable storage medium is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of a household pet identification method in the above embodiment. It should be understood by those skilled in the art that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0068] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for identifying household pets, characterized in that the method include: Obtain a pet image to be identified; Based on a pre-trained pet recognition model, pets in the pet image to be recognized are recognized to obtain recognition results of household pets; wherein, deep learning training is performed on image data of different types of pets to obtain a pet recognition model; wherein, the deep learning training is completed by adjusting the YOLOv7 neural network after the network structure, wherein the method of adjusting the network structure of the YOLOv7 neural network includes: connecting a GAM attention mechanism layer to the output of the BackBone network of the YOLOv7 neural network, connecting an SA attention mechanism layer to the output of the Neck network layer of the YOLOv7 neural network, setting the detection head of the YOLOv7 neural network to a Decoupled_Detect decoupling head, and connecting the SA attention mechanism layer to the Decoupled_Detect decoupling head.

2. A household pet identification method according to claim 1, characterized in that: The method of adjusting the network structure of the YOLOv7 neural network also includes: replacing the BackBone network of the YOLOv7 neural network with a Shuffle_Block network, replacing the Conv layer in the Neck network layer of the YOLOv7 neural network with a GSConv layer, and connecting the GAM attention mechanism layer to the GSConv layer.

3. A household pet identification method according to claim 2, characterized in that: The Shuffle_Block network has multiple Shuffle Block modules stacked in sequence, and the output of the last Shuffle Block module is used as the input of the GAM attention mechanism layer.

4. A household pet identification method according to claim 2, characterized in that: The number of the Decoupled_Detect decoupling heads is 3, wherein the inputs of two Decoupled_Detect decoupling heads are the outputs of two C3 modules of the Neck network layer respectively, and the input of one Decoupled_Detect decoupling head is the output of the SA attention mechanism layer.

5. A household pet identification system, characterized in that: The system includes: An image acquisition module, used for acquiring an image of a pet to be identified; An image recognition module is used to identify pets in a pet image to be identified based on a pre-trained pet recognition model to obtain recognition results of household pets; wherein, deep learning training is performed on image data of different types of pets to obtain a pet recognition model; wherein the deep learning training is completed by adjusting the YOLOv7 neural network after the network structure is adjusted, wherein the method of adjusting the network structure of the YOLOv7 neural network includes: connecting a GAM attention mechanism layer to the output of the BackBone network of the YOLOv7 neural network, connecting an SA attention mechanism layer to the output of the Neck network layer of the YOLOv7 neural network, setting the detection head of the YOLOv7 neural network to a Decoupled_Detect decoupling head, and connecting the SA attention mechanism layer to the Decoupled_Detect decoupling head.

6. A household pet identification system according to claim 5, characterized in that: The method of adjusting the network structure of the YOLOv7 neural network also includes: replacing the BackBone network of the YOLOv7 neural network with a Shuffle_Block network, replacing the Conv layer in the Neck network layer of the YOLOv7 neural network with a GSConv layer, and connecting the GAM attention mechanism layer to the GSConv layer.

7. A household pet identification system according to claim 6, characterized in that: The Shuffle_Block network has multiple Shuffle Block modules stacked in sequence, and the output of the last Shuffle Block module is used as the input of the GAM attention mechanism layer.

8. A household pet identification system according to claim 6, characterized in that: The number of the Decoupled_Detect decoupling heads is 3, wherein the inputs of two Decoupled_Detect decoupling heads are the outputs of two C3 modules of the Neck network layer respectively, and the input of one Decoupled_Detect decoupling head is the output of the SA attention mechanism layer.

9. A smart pet terminal, characterized in that: The smart pet terminal includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of a household pet identification method as described in any one of claims 1 to 4 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of a household pet identification method as claimed in any one of claims 1 to 4 are implemented.

Citation Information

Cited By

  • Home remote monitoring and early warning system and method based on personalized threshold value

    CN120808582A