An image recognition system for recognizing pig images and its image recognition method
By using a deep high-resolution network model to simultaneously identify pig targets and their key points in the pig house, the problems of high computing resource consumption and high model management and maintenance costs in the prior art are solved, and more efficient identification and diagnostic effects are achieved.
Patent Information
- Application Number
- CN202111481158.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-06
AI Technical Summary
In the prior art, when detecting key points on infrared thermal imaging of pigs in pig houses, two separate models are required, resulting in increased computing resource consumption and high model management and maintenance costs.
A deep high-resolution network model is used to receive image data related to pig images and perform operations to simultaneously identify pig targets and their key points.
It reduces the consumption of computing resources, simplifies task management, and improves the accuracy of identification results for use in the diagnosis and treatment of pigs.
Smart Images

Figure CN114202771B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of image processing technology. More specifically, the present disclosure relates to an image recognition system, an image recognition method, and a computer-readable storage medium for recognizing pig images. Background Art
[0002] In the pig farming industry in pig houses, it is usually necessary to detect the temperature of key parts (keypoints) on the infrared thermal image of pigs to determine whether the pigs have a fever, so as to diagnose and treat the pigs. This can not only avoid potential contamination caused by direct contact between humans and pigs, but also save human resources. Currently, the "top-down" method is usually adopted to detect the keypoints on the infrared thermal image of pigs, that is, two separate models are used to detect the keypoints of pigs. For example, first, the target is determined via a target detector (such as the yolo series), and then the keypoints of the pigs are obtained via a keypoint detector (such as alphapose). However, using two models increases the consumption of computing resources and the cost of model management and maintenance. Summary of the Invention
[0003] In order to at least partially solve the technical problems mentioned in the background art, the solution of the present disclosure provides a solution for recognizing pig images. By using the solution of the present disclosure, the pig target and its keypoints can be obtained simultaneously by means of one model, thereby significantly reducing the consumption of computing resources and simplifying the task. For this purpose, the present disclosure provides solutions in the following aspects.
[0004] In one aspect, the present disclosure provides an image recognition system for recognizing pig images, including: one or more processors; a Deep High-Resolution Network (HRNet) model; and a computer-readable storage medium storing program instructions for implementing the Deep High-Resolution Network model, which, when the program instructions are run by the one or more processors, cause the Deep High-Resolution Network model to perform: receiving image data related to the pig image; and performing operations on the image data to identify the pig target in the image data and the keypoints in the pig target.
[0005] In one embodiment, the image data is image data obtained by preprocessing the original image data related to the pig image.
[0006] In another embodiment, the computer-readable storage medium further stores program instructions for preprocessing the original image data, which, when the program instructions are run by the one or more processors, perform the following steps: performing an image stretching operation on the original image data to obtain an initial transformed image; and performing a mean transformation operation on the initial transformed image to preprocess the original image data.
[0007] In yet another embodiment, the computer-readable storage medium further stores program instructions for performing a downsampling operation on the original image data. When the program instructions are run by the one or more processors, the downsampling operation is performed on the original image data to obtain downsampled image data.
[0008] In yet another embodiment, the deep high-resolution network model includes a plurality of convolutional layers and a feature fuser, wherein: the plurality of convolutional layers are configured to perform multi-layer convolutional processing on the image data to extract feature data and identify pig targets and key points of the pig targets in the image data; and the feature fuser is configured to merge the feature data with the downsampled image data to obtain fused feature data.
[0009] In yet another embodiment, the plurality of convolutional layers include a plurality of serially-connected convolutional layers and a parallel-connected convolutional layer, wherein: the plurality of serially-connected convolutional layers are configured to perform multi-layer convolutional processing on the image data to extract feature data; the parallel-connected convolutional layer is configured to perform convolutional processing on the fused feature data to identify pig targets and key points of the pig targets in the image data.
[0010] In yet another embodiment, an output end of a last convolutional layer of the serially-connected convolutional layer is connected to an input end of the feature fuser, and an input end of the parallel-connected convolutional layer is connected to an output end of the feature fuser.
[0011] In yet another embodiment, the key points of the pig target include the left ear root, right ear root, left groin, and right groin of the pig target.
[0012] In another aspect, the present disclosure further provides an image recognition method for recognizing a pig image, including: inputting image data related to the pig image into a deep high-resolution network model; and using the deep high-resolution network model to perform operations on the image data to identify pig targets and key points in the pig targets in the image data.
[0013] In yet another aspect, the present disclosure further provides a computer-readable storage medium storing computer-readable instructions of program instructions for recognizing a pig image. When the computer-readable instructions are executed by one or more processors, the foregoing multiple embodiments are implemented.
[0014] Through the solution of the present disclosure, by performing operations on image data related to pig images through a deep high-resolution network model, the pig target and its key points can be obtained simultaneously, thereby reducing the consumption of computing resources. Further, in the embodiments of the present disclosure, the feature data is merged with the downsampled image data of the original image data, which can improve the resolution of the fused feature data, so as to obtain accurate recognition results for subsequent diagnosis and treatment of pigs. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understandable. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0016] Figure 1 is an exemplary structural block diagram of an image recognition system for recognizing pig images according to an embodiment of the present disclosure;
[0017] Figure 2 is an exemplary schematic diagram of an operation block of a deep high-resolution network model according to an embodiment of the present disclosure;
[0018] Figure 3 is an exemplary schematic diagram of an operation block of a parallel-connected convolutional layer according to an embodiment of the present disclosure;
[0019] Figure 4 is an exemplary flowchart of an image recognition method for recognizing pig images according to an embodiment of the present disclosure; and
[0020] Figure 5 is a block diagram of a computing device for recognizing pig images according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the embodiments described in this specification are only some embodiments provided by the present disclosure for the convenience of clear understanding of the solution and to meet the requirements of the law, rather than all the embodiments that can implement the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments disclosed in this specification without creative efforts belong to the scope of protection of the present disclosure.
[0022] Figure 1 is an exemplary structural block diagram of an image recognition system 100 for recognizing pig images according to an embodiment of the present disclosure. As Figure 1As shown, the image recognition system 100 may include a processor 101, a deep high-resolution network model 102, and a computer-readable storage medium 103. In one implementation scenario, the aforementioned deep high-resolution network model 102 may be implemented as computer program instructions stored (or resident) on a computer-readable storage medium, such as binary instruction codes.
[0023] In one embodiment, the aforementioned processor 101 may be one or more, and thus the present disclosure does not limit the number of processors. In some implementation scenarios, the processor may be a general-purpose processor ("CPU") or a dedicated graphics processing unit ("GPU"). In some other implementation scenarios, a combination of a CPU and a GPU may also be used, such as in some heterogeneous architecture systems. When using the aforementioned heterogeneous architecture system, computer program instructions regarding the embodiments of the present disclosure may be compiled at the CPU to obtain an executable program. Then, the executable program may be transmitted to the GPU through the driver interface between the CPU and the GPU to execute the program to perform recognition on the input image data related to the pig image.
[0024] Although the computer-readable storage medium is shown as a single box in the figure, the number thereof may also be multiple, such as various storage media capable of storing computer program instructions. As described above, the program instructions may include program instructions for implementing the deep high-resolution network model. For example, when the processor executes the aforementioned one or more program instructions, the deep high-resolution network model of the present disclosure may be configured to receive image data related to the pig image and perform operations on the image data to identify the pig target in the image data and the key points in the pig target. In some embodiments, the aforementioned image data related to the pig image may be image data obtained by preprocessing the original image data related to the pig image. The aforementioned key points may include, for example, the left ear root, the right ear root, the left groin, and the right groin of the pig target.
[0025] In an application scenario, the above-mentioned original image data related to pig images can be collected by, for example, an infrared camera. Based on the above description, the above-mentioned computer-readable storage medium also stores program instructions for preprocessing the original image data. When the program instructions are run by one or more processors, the following steps are executed to preprocess the aforementioned original image data. Specifically, first, an image stretching operation is performed on the original image data to obtain an initial transformed image. For example, the original image data is stretched into an initial transformed image with a size of 512*512. Then, a mean transformation operation is performed on the initial transformed image to preprocess the original image data. It can be understood that the mean transformation refers to calculating the mean and standard deviation of the initial transformed image respectively, and performing data normalization on the obtained mean and standard deviation, so that the preprocessed image data conforms to the standard normal distribution (i.e., the mean is 0 and the standard deviation is 1). That is, the image data related to pig images received by the deep high-resolution network model of the embodiments of the present disclosure conforms to the standard normal distribution.
[0026] In one embodiment, the above-mentioned computer-readable storage medium also stores program instructions for performing a downsampling operation on the original image data. When the program instructions are run by one or more processors, a downsampling operation is performed on the original image data to obtain downsampled image data. It can be understood that the aforementioned downsampling (i.e., shrinking the image) operation can make the image conform to the size of the display area and generate a thumbnail of the corresponding image. For example, in the embodiments of the present disclosure, the resolution of the original image data can be downsampled by a factor of 4. For example, downsampling the resolution of the original image data with a size of 512*512 by a factor of 4 can obtain downsampled image data with a size of 128*128. The downsampled image data can be merged with subsequent feature data to improve the resolution of the fused feature data, so as to obtain an accurate recognition result.
[0027] In an implementation scenario, the deep high-resolution network model of the embodiments of the present disclosure may include multiple convolutional layers and a feature fuser. Among them, the multiple convolutional layers may be configured to perform multi-layer convolutional processing on image data to extract feature data and identify pig targets and key points of the pig targets in the image data. The aforementioned feature fuser may be configured to merge the aforementioned feature data with the above downsampled image data to obtain fused feature data. In some embodiments, the aforementioned multiple convolutional layers may include multiple serially connected convolutional layers and a parallelly connected convolutional layer, and the output end of the last convolutional layer of the serially connected convolutional layers is connected to the input end of the feature fuser, while the input end of the parallelly connected convolutional layer is connected to the output end of the feature fuser. In the implementation scenario, the aforementioned serially connected multiple convolutional layers may be configured to perform multi-layer convolutional processing on image data to extract feature data. The aforementioned parallelly connected convolutional layer may be configured to perform convolutional processing on the fused feature data to identify pig targets and key points of the pig targets in the image data.
[0028] Combined with the above description, it can be seen that the embodiments of the present disclosure can simultaneously obtain pig targets and their key points by performing operations on image data related to pig images through a deep high-resolution network model. Compared with the existing method of using two models to sequentially obtain pig targets and key points, this reduces the consumption of computing resources and lowers the cost of model management and maintenance.
[0029] As mentioned above, the image data related to pig images in the embodiments of the present disclosure may be image data obtained by preprocessing the original image data related to pig images. Further, the preprocessed image data is processed using the deep high-resolution network model to identify pig targets and key points in the image data. The following will be combined with Figure 2 to describe in detail the deep high-resolution network model of the embodiments of the present disclosure.
[0030] Figure 2 is an exemplary schematic diagram showing the operation boxes of the deep high-resolution network model according to the embodiments of the present disclosure. It should be understood that Figure 2 the deep high-resolution network model 102 shown is Figure 1 a specific implementation of the deep high-resolution network model 102 in the image recognition system 100 shown. Thus, regarding Figure 1 the relevant details and features of the described image recognition system 100 also apply to Figure 2 the description.
[0031] As Figure 2As shown, the depth high-resolution network model 102 may include a plurality of convolutional layers 201 and a feature fuser 202. For example, four convolutional layers are exemplarily shown in the figure, and the four convolutional layers are, in sequence, convolutional layer 201-1 to convolutional layer 201-4, where convolutional layers 201-1 to 201-3 are serially connected, and convolutional layer 201-4 is connected in parallel with convolutional layers 201-1 to 201-3. Further, the output end of the last serially connected convolutional layer 201-3 is connected to the input end of the feature fuser 202, and the input end of the parallel-connected convolutional layer 201-4 is connected to the output end of the feature fuser 202. In an implementation scenario, the depth high-resolution network model 102 first sequentially passes the received image data related to the pig image through convolutional layers 201-1 to 201-3 to extract feature data. Then, the feature fuser 202 combines the feature data and the above-mentioned downsampled image data to obtain fused feature data. Finally, the convolutional layer 201-4 processes the fused feature data to simultaneously obtain the pig target and the key points of the pig target.
[0032] In sequentially passing the received image data related to the pig image through convolutional layers 201-1 to 201-3 to extract feature data, first, the received image data related to the pig image is subjected to a convolution operation through convolutional layer 201-1. Under this convolutional layer 201-1, a convolutional kernel with a size of 1*1 is used, and the sliding step of the convolutional kernel is 1 to obtain feature image data conv1. It can be understood that the resolution of the feature image data conv1 remains unchanged compared to the image data. Next, the previously extracted feature image data conv1 is subjected to a convolution operation through convolutional layer 201-2. Under this convolutional layer 201-2, a convolutional kernel with a size of 3*3 is used, and the sliding step of the convolutional kernel is 4 to obtain feature image data conv2. It can be understood that the resolution of the feature image data conv2 processed by convolutional layer 201-2 is reduced by 4 times. Further, the feature image data conv2 is subjected to a convolution operation through convolutional layer 201-3. Under this convolutional layer 201-3, a convolutional kernel with a size of 1*1 is used, and the sliding step of the convolutional kernel is 1 to obtain feature image data conv3. That is, the resolution of the feature image data conv3 remains unchanged, and the feature image data conv3 is the feature data extracted from the image data sequentially through convolutional layers 201-1 to 201-3.
[0033] Based on the above-extracted feature data, the feature data is merged with the above downsampled data through the feature fuser 202 to obtain fused feature data. That is, the feature image data conv3 is, for example, added to the above downsampled data to obtain fused feature data. The figure further shows that the previously obtained fused feature data is input into the convolutional layer 201-4 to perform a convolutional operation on the fused feature data to obtain multiple channels. For example, in an exemplary scenario, 17 channels can be obtained by processing the image data related to the pig image, and they are divided into 5 channel layers. The 5 channel layers correspond to the classification layer, the width and height layer, the key point layer, the regression layer, and the heat map layer of the key points in sequence. For example Figure 3 as shown.
[0034] Figure 3 is an exemplary schematic diagram showing the operation box of the parallel-connected convolutional layers according to an embodiment of the present disclosure. It should be understood that Figure 3 the parallel-connected convolutional layers shown Figure 2 are a specific implementation of the deep high-resolution network model 102 shown. Thus, regarding Figure 2 the relevant details and features of the described deep high-resolution network model 102 also apply to Figure 3 the description.
[0035] As Figure 3 shown, the parallel-connected convolutional layer 201-4 receives the fused feature data output by the above feature fuser, performs a convolutional operation on the fused feature data to obtain 17 channels, and then divides the 17 channels into 5 channel layers. The 5 channel layers correspond to the classification layer 301, the width and height layer 302, the key point layer 303, the regression layer 304, and the heat map layer 305 of the key points in sequence. Among them, the classification layer 301 contains 1 layer, that is, one channel. Similarly, the aforementioned width and height layer 302, key point layer 303, regression layer 304, and heat map layer 305 of the key points respectively contain 2 layers, 8 layers, 2 layers, and 4 layers.
[0036] In some embodiments, the aforementioned classification layer corresponds to the classification of the target, that is, whether the target is a pig. In one implementation scenario, the value corresponding to the classification layer is "1" or "0". When the value corresponding to the classification layer is "1", the target is a pig. Conversely, when the value corresponding to the classification layer is "0", the target is not a pig. Further, the aforementioned width and height layer corresponds to the width and height of the target position, the key point layer corresponds to the coordinates of four key points, the regression layer corresponds to the offset of the center point of the border, and the heat map layer of the key points corresponds to the heat map prediction of the key points.
[0037] It should be understood that the peaks in the obtained heatmap can correspond to the centers of the pig targets (or bounding boxes), and the image features at each peak can be used to predict the height and weight of the bounding box. For key points, their positions can be regarded as the offset values from the center of the bounding box, and regression analysis can be performed at the center of the bounding box to obtain the key points of the pig target. Based on this, decoding the above five channel layers (classification layer, width-height layer, key point layer, regression layer, and heatmap layer of key points) to obtain the probabilities corresponding to the bounding box and the probabilities corresponding to the positions of key points can identify the pig target and the key points in the pig target. Specifically, by decoding the classification layer 301, width-height layer 302, and regression layer 304, the probabilities corresponding to the bounding box can be obtained, and thus the pig target 306 can be obtained. By decoding the key point layer 303 and the heatmap layer 305 of key points, the probabilities corresponding to the positions of key points can be obtained, and thus the key points 307 in the pig target can be obtained simultaneously. Temperature detection is performed based on the identified key points in the pig target to determine whether the pig has a fever, so as to conduct quality and diagnosis on the pig.
[0038] Figure 4 FIG. 4 is an exemplary flowchart of an image recognition method 400 for recognizing a pig image according to an embodiment of the present disclosure. As Figure 4 shown, at step S402, the image data related to the pig image is input into the Deep High-Resolution Network model. As can be seen from the foregoing, the image data related to the pig image is the image data obtained by preprocessing the original image data related to the pig image. In one embodiment, the original image data can first be subjected to, for example, an image stretching operation to obtain an initial transformed image. Then, a mean transformation operation can be performed on the initial transformed image to preprocess the original image data. At step S404, the Deep High-Resolution Network model is used to perform operations on the image data to identify the pig target and the key points in the pig target in the image data.
[0039] In one implementation scenario, the above Deep High-Resolution Network model can include multiple convolutional layers and a feature fusion unit. For example, it can include four convolutional layers and one feature fusion unit. Among them, the four convolutional layers include three serially connected convolutional layers and one parallelly connected convolutional layer, and the output end of the last convolutional layer of the three serially connected convolutional layers is connected to the input end of the feature fusion unit, and the input end of the one parallelly connected convolutional layer is connected to the output end of the feature fusion unit. In some embodiments, the multiple serially connected convolutional layers can be used to perform multi-layer convolutional processing on the image data to extract feature data, and the one parallelly connected convolutional layer can be used to perform convolutional processing on the fused feature data to identify the pig target and the key points of the pig target in the image data. For the operations of the Deep High-Resolution Network model, reference can be made to the above Figures 2-3The content described above will not be elaborated herein in the present disclosure.
[0040] Although the training process of the deep high-resolution network model of the present disclosure is not mentioned above, based on the content of the present disclosure, those skilled in the art can understand that the deep high-resolution network model of the present disclosure can be trained with training data to obtain a deep high-resolution network model with high precision. For example, in the forward propagation process of neural network training, the present disclosure can utilize the combined Figures 2-3 obtained to include fused feature data to train the deep high-resolution network model of the present disclosure, and compare the training result with the expected result (or true value) to obtain the corresponding loss function. Further, in the backpropagation process of neural network training, the present disclosure utilizes the obtained loss function and updates the weights based on, for example, the gradient descent algorithm to reduce the error between the output and the true value.
[0041] Combined with the above description, using the image recognition system of the present disclosure, the recognition result can be obtained by processing the image data related to the pig image through the deep high-resolution network model. For example, the image data related to the pig image collected can be input into the image recognition system of the present disclosure, and the pig target in the image data and the key points in the pig target can be recognized simultaneously. By detecting the temperature of the key points, the pig target can be diagnosed and treated.
[0042] Figure 5 is a block diagram showing a computing device for recognizing a pig image according to an embodiment of the present disclosure. As Figure 5 shown, the computing device 500 may include a central processing unit (“CPU”), which may be a general-purpose CPU, a dedicated CPU, or other information processing and program execution units. Further, the computing device 500 may also include a mass storage 512 and a read-only memory (“ROM”) 513, where the mass storage 512 may be configured to store various types of data, such as various image data related to pig images, algorithm data, intermediate results, and various programs required to run the computing device 500. The read-only memory 513 may be configured to store data required for the power-on self-test of the computing device 500, the initialization of each functional module in the system, the driver for the basic input / output of the system, and the data required to boot the operating system.
[0043] Optionally, the computing device 500 may further include other hardware platforms or components, such as the illustrated TPU (Tensor Processing Unit) 514, GPU (Graphics Processing Unit or Graphics Processor) 515, FPGA (Field Programmable Gate Array) 516, and MLU (Machine Learning Unit) 517. It can be understood that although a variety of hardware platforms or components are illustrated in the computing device 500, these are merely exemplary and not restrictive, and those skilled in the art can add or remove corresponding hardware according to actual needs. For example, the computing device 500 may include only a CPU to implement the image recognition system of the present disclosure.
[0044] The computing device 500 of the present disclosure further includes a communication interface 518, so that it can be connected to a local area network / wireless local area network (LAN / WLAN) 505 through the communication interface 518, and further can be connected to a local server 506 or connected to the Internet (“Internet”) 507 through the LAN / WLAN. Alternatively or additionally, the computing device 500 of the present disclosure can also be directly connected to the Internet or a cellular network based on wireless communication technology through the communication interface 518, such as based on the 3rd generation (“3G”), 4th generation (“4G”), or 5th generation (“5G”) wireless communication technology. In some application scenarios, the computing device 500 of the present disclosure can also access a server 508 and a database 509 of an external network as needed, so as to obtain various known image models, data, and modules, and can remotely store various data, such as various types of data for presenting pig images.
[0045] The peripheral devices of the computing device 500 may include a display device 502, an input device 503, and a data transmission interface 504. In one embodiment, the display device 502 may include, for example, one or more speakers and / or one or more visual displays, which are configured to provide voice prompts and / or image / video displays for the operation process or final result of displaying pig images of the present disclosure. The input device 503 may include, for example, a keyboard, a mouse, a microphone, a gesture capture camera, and other input buttons or controls, which are configured to receive the input of pig image data and / or user instructions. The data transmission interface 504 may include, for example, a serial interface, a parallel interface, or a universal serial bus interface (“USB”), a small computer system interface (“SCSI”), serial ATA, FireWire, PCI Express, and a high-definition multimedia interface (“HDMI”), etc., which are configured for data transmission and interaction with other devices or systems. According to the solution of the present disclosure, the data transmission interface 504 can receive image data related to pig images collected by an infrared camera and transmit image data related to pig images or various other types of data and results to the computing device 500.
[0046] The above-mentioned CPU 511, mass storage 512, read-only memory ROM 513, TPU 514, GPU 515, FPGA 516, MLU 517, and communication interface 518 of the computing device 500 of the present disclosure can be interconnected via a bus 519, and data interaction with peripheral devices can be achieved through this bus. In one embodiment, through this bus 519, the CPU 511 can control other hardware components and their peripheral devices in the computing device 500.
[0047] The above combination Figure 5 has described a computing device that can be used to execute the image recognition system of the present disclosure. It should be understood that the structure of the computing device here is merely exemplary, and the implementation manner and implementation entity of the present disclosure are not limited by it, but can be changed without departing from the spirit of the present disclosure.
[0048] It should also be understood that any module, unit, component, server, computer, terminal, or device that executes instructions in the examples of the present disclosure can include or otherwise access a computer-readable medium, such as a storage medium, a computer storage medium, or a data storage device (removable) and / or non-removable), such as a magnetic disk, an optical disk, or a magnetic tape. The computer storage medium can include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. When the foregoing computer-readable instructions, data structures, program modules, or other data are executed, the image recognition method for identifying pig images described in the present disclosure in combination with the attached Figure 4 can be implemented.
[0049] It should be noted that although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be changed in the order of execution. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0050] It should be understood that when terms such as "first", "second", "third", and "fourth" are used in the claims, the description, and the drawings of the present disclosure, they are only used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the description and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0051] It should also be understood that the terms used in this disclosure are merely for the purpose of describing specific embodiments and are not intended to limit the disclosure. As used in this disclosure and the claims, unless the context clearly indicates otherwise, the singular forms "a," "an," and "the" are intended to include the plural forms. It should also be further understood that the term "and / or" used in this disclosure and the claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0052] Although the embodiments of the present disclosure are as described above, the above content is only an example for facilitating the understanding of the present disclosure and is not intended to limit the scope and application scenarios of the present disclosure. Any person skilled in the art within the technical field of the present disclosure may make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed by the present disclosure. However, the scope of patent protection of the present disclosure shall still be subject to the scope defined by the appended claims.
Claims
1. An image recognition system for recognizing pig images, comprising: one or more processors; a deep high-resolution network model; and a computer-readable storage medium storing program instructions for implementing the deep high-resolution network model, which, when run by the one or more processors, cause the deep high-resolution network model to perform: receiving image data related to the pig image; and performing operations on the image data to identify pig targets in the image data and key points in the pig targets; wherein the deep high-resolution network model includes a plurality of convolutional layers and a feature fusion unit, wherein: the plurality of convolutional layers are configured to perform multi-layer convolutional processing on the image data to extract feature data and identify pig targets in the image data and key points of the pig targets; the feature fusion unit is configured to merge the feature data with downsampled image data to obtain fused feature data; the plurality of convolutional layers include a plurality of serially connected convolutional layers and a parallelly connected convolutional layer, wherein: the plurality of serially connected convolutional layers are configured to perform multi-layer convolutional processing on the image data to extract feature data; the parallelly connected convolutional layer is configured to perform convolutional processing on the fused feature data to identify pig targets in the image data and key points of the pig targets; wherein an output end of a last convolutional layer of the plurality of serially connected convolutional layers is connected to an input end of the feature fusion unit, and an input end of the parallelly connected convolutional layer is connected to an output end of the feature fusion unit; the parallelly connected convolutional layer obtains a plurality of channels after performing a convolutional operation on the fused feature data, and divides them into a classification layer, a width-height layer, a key-point layer, a regression layer, and a heat map layer of key points.
2. The image recognition system according to claim 1, wherein the image data is image data obtained by preprocessing the original image data related to the pig image.
3. The image recognition system according to claim 2, wherein the computer-readable storage medium further stores program instructions for preprocessing the original image data, which, when run by the one or more processors, perform the following steps: performing an image stretching operation on the original image data to obtain an initial transformed image; and performing a mean transformation operation on the initial transformed image to preprocess the original image data.
4. The image recognition system according to claim 2, wherein the computer-readable storage medium further stores program instructions for performing a downsampling operation on the original image data, which, when run by the one or more processors, perform a downsampling operation on the original image data to obtain downsampled image data.
5. The image recognition system according to claim 1, wherein the key points of the pig target include the left ear root, the right ear root, the left groin, and the right groin of the pig target.
6. An image recognition method for recognizing pig images, comprising: Input the image data related to the pig image into the deep high-resolution network model; And Use the deep high-resolution network model to perform operations on the image data to identify the pig target in the image data and the key points in the pig target; Wherein the deep high-resolution network model includes a plurality of convolutional layers and a feature fuser, wherein: the plurality of convolutional layers are configured to perform multi-layer convolutional processing on the image data to extract feature data and identify the pig target in the image data and the key points of the pig target; the feature fuser is configured to merge the feature data with the downsampled image data to obtain fused feature data; The plurality of convolutional layers include a plurality of serially connected convolutional layers and a parallel-connected convolutional layer, wherein: the plurality of serially connected convolutional layers are configured to perform multi-layer convolutional processing on the image data to extract feature data; the parallel-connected convolutional layer is configured to perform convolutional processing on the fused feature data to identify the pig target in the image data and the key points of the pig target; Wherein the output end of the last convolutional layer of the plurality of serially connected convolutional layers is connected to the input end of the feature fuser, and the input end of the parallel-connected convolutional layer is connected to the output end of the feature fuser; the parallel-connected convolutional layer obtains a plurality of channels after performing convolutional operations on the fused feature data, and divides them into a classification layer, a width-height layer, a key-point layer, a regression layer, and a heat map layer of the key points.
7. A computer-readable storage medium, which includes program instructions for identifying a pig image, and when the program instructions are executed by one or more processors, the method according to claim 6 is implemented.
Citation Information
Patent Citations
Abdomen multi-organ nuclear magnetic resonance image segmentation method and system based on FCN and medium
CN110705555A
Human body key point identification method and device, intelligent terminal and storage medium
CN112712015A
Human body posture prediction method and system based on improved high-resolution network
CN113076891A