Cattle face recognition method, device, electronic device and medium integrating deep learning
Through the cow face recognition method that integrates depth information and image information, the binocular camera and twin network are used to solve the robustness of cow face recognition under light and posture changes, and achieve higher recognition accuracy and stability.
Patent Information
- Application Number
- CN202211460737.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-11-17
AI Technical Summary
The existing cow face recognition technology is not robust enough in different scenarios, especially affected by the changes in the cow's posture and lighting, making it difficult to accurately identify the cow's identity.
The method of fusion of depth information and image information is adopted to obtain the RGB and depth maps of the cow's face through a binocular camera, and the training is combined with the depth image segmentation algorithm and twin network. The outline information and lighting invariance of the depth map are used to improve the recognition accuracy.
It improves the robustness of cattle face recognition under different lighting and posture changes, enhances the accuracy and stability of recognition, and performs well in complex environments such as ranches.
Smart Images

Figure CN115909401B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual recognition, and particularly relates to a cow face recognition method, device, electronic device and medium integrating deep learning. Background Art
[0002] An image segmentation algorithm combining RGB images and depth images. A cow face segmentation algorithm is proposed, which combines RGB images and depth images for cow face segmentation. By using the characteristic that the pixel values of the depth image represent the distances between points of the photographed object and the camera and are not affected by the object color, the depth image is converted into an HSV space picture, and threshold segmentation is performed using the threshold of V (brightness). A method for obtaining the segmentation threshold is proposed. According to the characteristic that the last large waveform of the depth image histogram in the cow face recognition scenario is the waveform of the cow face, the last large wave valley of the depth image histogram is used as the segmentation threshold for threshold segmentation. The RGB image segmentation algorithm GrabCut is used for auxiliary segmentation. The cow face segmentation algorithm avoids the adverse effects of the similarity or identity between the cow face color and the background color on segmentation. After segmentation, the foreground and background of the cow face picture are separated, avoiding interference in subsequent operations on the cow face picture. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a cow face recognition method, device, electronic device and readable storage medium integrating deep learning for the deficiencies of the above-mentioned prior art. This method improves the cow face recognition technology by using the method of fusing depth information and image information, reduces the influence of changes in cow posture and illumination on cow face recognition, and improves the robustness of cow face recognition in different scenarios.
[0004] To solve the above technical problem, the technical solution adopted by the present invention is: A cow face recognition method integrating deep learning, characterized by including:
[0005] Obtain paired RGB images and depth images of the sample cow's face, and make an image data set of the sample cow;
[0006] Input the image data set of the sample cow into the cow face segmentation algorithm combining RGB images and depth images, segment the cow face from the background of the RGB image and the depth image to obtain an RGB cow face image and a depth cow face image, and then form a cow face picture pair according to the RGB cow face image and the depth cow face image of the sample cow. A sample cow face data set is constructed from the sample cow face picture pairs. The sample cow face data set includes cow face RGB images and depth images of multiple sample pairs. The cow face RGB images and depth images of each sample pair include cow face RGB images and depth images of two samples. The two sample cows come from different cows or the same cow. The cow face RGB images and depth images of each sample pair are labeled, and the label is used to classify whether the two sample cows are from the same cow;
[0007] Input the sample cattle face dataset into the cattle face recognition network that fuses depth information and image information for training until the cattle face recognition network distinguishes the pictures of different sample cattle in the sample cattle face dataset from each other, end the training, and obtain the trained cattle face recognition network;
[0008] First, register the paired RGB images and depth images of each cattle in the cattle farm. Select one cattle as the cattle to be recognized, obtain the paired RGB image and depth image of the cattle to be recognized, segment them into the RGB cattle face image and depth cattle face image of the cattle to be recognized through the cattle face segmentation algorithm that combines the RGB image and the depth image, and then input them into the trained cattle face recognition network to obtain the identity information of the cattle to be recognized.
[0009] Further, in S1, the depth camera is a binocular camera, and the depth image and RGB image of the same picture of the sample cattle face are collected, that is, the paired RGB image and depth image of the sample cattle face.
[0010] Further, in S2, input the image dataset of the sample cattle into the cattle face segmentation algorithm that combines the RGB image and the depth image to segment the cattle face from the background of the RGB image and the depth image to obtain the RGB cattle face image and the depth cattle face image; specifically including:
[0011] S201. Cluster and display the pixels from dark to bright in the depth image of the image dataset of the sample cattle to obtain the histogram of the depth image;
[0012] S202. Use the algorithm for obtaining wave troughs based on continuous wavelet transform to segment the wave troughs of the histogram of the depth image in S201, calculate the coordinate array of the cattle face rectangular frame in the depth image according to the threshold of the last wave trough, and finally obtain the cattle face rectangular frame in the RGB image according to the coordinate array of the cattle face rectangular frame in the depth image;
[0013] S203. Perform the GrabCut segmentation algorithm on the RGB image of the cattle face within the cattle face rectangular frame in the RGB image in S202 to obtain the initially segmented cattle face image A;
[0014] S204. Based on the brightness of the depth image where the last wave trough is located in S202 as the threshold, segment the depth image, keep the part greater than the threshold, remove the part less than the threshold and replace it with black as the background color; then obtain the coordinates of the foreground pixel points of the depth image, and segment the RGB image according to the coordinates to obtain the initially segmented cattle face image B.
[0015] Further, the sample pair in S2 is a positive sample pair or a negative sample pair. When the sample pair is a positive sample pair, it means that the two sample cattle of the sample pair come from the same cattle. When the sample pair is a negative sample pair, it means that the two sample cattle of the sample pair come from different cattle.
[0016] Further, the form of the label in S2 is an array, where the subscript of the array is the sample pair number, and the value of the array is 0 or 1. Specifically, assume that X1 and X2 are sample pairs of two sample cows respectively, and Y is the label of the sample pair. When the sample pair composed of X1 and X2 comes from the same cow, the sample pair is matched and is a positive sample pair, and the label Y is set to 1, indicating that they come from the same sample. When the sample pair composed of X1 and X2 comes from different cows, the sample pair is not matched and is a negative sample pair, and the label Y is set to 0, indicating that they come from different samples.
[0017] Further, the cow face recognition network that fuses depth information and image information in S3 uses a siamese network as the backbone network for cow face recognition. The two sub-networks with shared weights are important components of the siamese network. The siamese network uses a convolutional neural network to map the original image to a high-dimensional feature space. Weight sharing means that in the two convolutional neural networks, the weights of the convolutional kernels in the convolutional layer, the bias of this channel in the convolutional layer, the weights in the fully connected layer, and the bias in the fully connected layer and other trainable parameters are synchronously updated as the number of training epochs increases.
[0018] Further, when the siamese network processes picture information, the input samples x1 and x2 are both RGB images. When the siamese network processes three-dimensional modalities, the input samples x1 and x2 are both depth images. f(x) includes a convolutional layer, a pooling layer, a dropout layer, and a fully connected layer. During the training process, parameter sharing is used to obtain the distance between each pair of samples. The parameter is represented by w, and both f(x1) and f(x2) use w as the parameter. Use distance<f(x1), f(x2)> to represent the distance between the two samples finally obtained after being processed by the network. According to the pre-labeled tags, reduce the loss of samples in the same category and increase the loss of samples in different categories.
[0019] Further, the RGB cow face image and the depth cow face image are separately trained using a siamese network. After extracting features, multiply by the corresponding weights. Finally, use the late decision-level fusion method to obtain the prediction result by taking the average of the sample distances distance<f(x1), f(x2)> output by the network.
[0020] The weights are calculated by the following method: Convert the RGB cow face image to the HSV space, use the lightness as the basis for judging the illumination intensity, and select the lightness with appropriate illumination and the clarity of the photo not affected by illumination as the standard value. The weights are represented by the following formula:
[0021]
[0022] W d =1 - W R
[0023] In the formula, W RRepresents the weight of the RGB cow face image, V S Represents the standard value of lightness, V P Represents the lightness value of the RGB cow face image, W d Represents the weight of the depth map;
[0024] The fusion adopts late decision-level fusion, and the voting method in deep learning is used to combine the weights for fusion. The voting method is mostly used in classical classification and recognition networks. The last fully connected layer outputs the probability that the sample belongs to each class. The voting method averages the probabilities of different modalities to obtain the final result. In this method, the siamese network judges whether the sample pair belongs to the same cow by distance. Therefore, the decision-level fusion obtains the prediction result by averaging the sample distances output by the network. First, the siamese network is used to train the data of the two modalities separately, and then the fusion is performed during decision-making.
[0025] The present invention also discloses a cow face recognition device integrating deep learning, including:
[0026] An acquisition module, configured to obtain paired RGB images and depth images of the sample cow's face, and make an image data set of the sample cow;
[0027] A processing module, configured to input the image data set of the sample cow into a cow face segmentation algorithm that combines the RGB image and the depth image, segment the cow face from the background of the RGB image and the depth image to obtain an RGB cow face image and a depth cow face image, and then form a cow face picture pair according to the RGB cow face image and the depth cow face image of the sample cow, and construct a sample cow face data set from the cow face picture pairs of the samples. The sample cow face data set includes the cow face RGB images and depth images of multiple sample pairs. The cow face RGB images and depth images of each sample pair include the cow face RGB images and depth images of two samples. The two sample cows come from different cows or the same cow. The cow face RGB images and depth images of each sample pair are labeled, and the label is used to classify whether the two sample cows come from the same cow;
[0028] A training module, configured to input the sample cow face data set into a cow face recognition network that fuses depth information and image information for training until the cow face recognition network distinguishes the pictures of different sample cows in the sample cow face data set, ends the training, and obtains a trained cow face recognition network;
[0029] An identification module, configured to first register the paired RGB images and depth images of each cow in the cattle farm, select one cow as the cow to be identified, obtain the paired RGB images and depth images of the cow to be identified, segment them into the RGB cow face image and the depth cow face image of the cow to be identified through a cow face segmentation algorithm that combines the RGB image and the depth image, and then input them into the trained cow face recognition network to obtain the identity information of the cow to be identified.
[0030] The present invention also discloses an electronic device, characterized in that the electronic device includes:
[0031] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a cattle face recognition method integrating deep learning as described above.
[0032] The present invention also discloses a computer-readable storage medium storing computer instructions, and the computer instructions are operated to execute the above cattle face recognition method integrating deep learning.
[0033] The present invention has the following advantages compared with the prior art:
[0034] 1. In the process of cattle individual recognition by the cattle face recognition algorithm of the present invention in the technology of integrating depth information and image information for cattle face recognition, features such as the contour and color of the cattle face are extracted through the two-dimensional image recognition process as the basis for recognition, and then the depth information of the cattle face is used to obtain the length information of the cattle face in the direction from the cattle forehead to the cattle nose, and this information is used as an extra extracted feature to help improve the recognition effect of two-dimensional cattle face recognition. In addition, the contour information of the cattle face in the depth map is more complete and is not interfered by the body image of the same color as the face. When the natural light conditions change drastically, the depth map is not affected. At the same time, this method is still effective when the light changes and has extremely high robustness.
[0035] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of the cattle face recognition method integrating depth information provided in Embodiment 1 of the present invention.
[0037] Figure 2 is a reference diagram of paired cattle face RGB images and depth maps provided in Embodiment 1 of the present invention.
[0038] Figure 3 is a flowchart of the cattle face segmentation algorithm combining RGB images and depth maps provided in Embodiment 1 of the present invention.
[0039] Figure 4 is a corresponding schematic diagram of the histogram of the depth map provided in Embodiment 1 of the present invention.
[0040] Figure 5 is a diagram showing the meaning of different waves in the histogram of the depth map provided in Embodiment 1 of the present invention.
[0041] Figure 6It is a schematic diagram of the histogram after smoothing processing provided by Embodiment 1 of the present invention.
[0042] Figure 7 It is a comparison chart of different threshold segmentations provided by Embodiment 1 of the present invention.
[0043] Figure 8 It is a structure diagram of the Siamese network provided by Embodiment 1 of the present invention.
[0044] Figure 9 It is a flowchart for the recognition algorithm to make a decision on fusing the RGB image and the depth image.
[0045] Figure 10 It is a comparison chart of the early fusion and late decision fusion effects of the recognition network provided by the present invention.
[0046] Figure 11 It is a comparison chart of the recognition networks of different methods provided by the present invention. Detailed implementation manners
[0047] Embodiment 1
[0048] As Figure 1 shown, a cow face recognition method integrating depth information provided by an embodiment of the present invention includes:
[0049] S1. Obtain a pair of RGB images and depth images of multiple sample cow faces, and make an image data set of the sample cows;
[0050] Further, it is collected by a depth camera, specifically a binocular camera, and two cameras with extremely close distances on it are used to capture simultaneously, which is convenient for obtaining the depth image and the RGB image of the same picture at the same moment, as Figure 2 shown;
[0051] S2. Input the image data set of the sample cows into a cow face segmentation algorithm that combines the RGB image and the depth image, segment the cow face from the background of the RGB image and the depth image to obtain an RGB cow face image and a depth cow face image, then form a cow face picture pair according to the RGB cow face image and the depth cow face image of the sample cows, and construct a sample cow face data set from the cow face picture pairs of the samples. The sample cow face data set includes the cow face RGB images and depth images of multiple sample pairs. The cow face RGB images and depth images of each sample pair include the cow face RGB images and depth images of two samples. The two sample cows come from different cows or the same cow. The cow face RGB images and depth images of each sample pair are labeled, and the label is used to classify whether the two sample cows come from the same cow;
[0052] S3. Input the sample bovine face dataset into the bovine face recognition network that fuses depth information and image information for training until the bovine face recognition network can distinguish the pictures of different sample bovines in the sample bovine face dataset from each other, end the training, and obtain the trained bovine face recognition network;
[0053] S4. First, register the paired RGB images and depth images of each bovine in the cattle farm. Select one bovine as the bovine to be recognized, obtain the paired RGB image and depth image of the bovine to be recognized, segment them into the RGB bovine face image and depth bovine face image of the bovine to be recognized through the bovine face segmentation algorithm that combines the RGB image and the depth image, and then input them into the trained bovine face recognition network to obtain the identity information of the bovine to be recognized.
[0054] In this embodiment, the binocular camera is fixed in the breeding base to collect RGB images and depth images. Before collecting data, it is necessary to enter the base at an appropriate time with the permission of the management personnel of the breeding base and disinfect to prevent bacteria or viruses from infecting the cattle. There are multiple sample bovines, and for each bovine, bovine face RGB images and depth images are collected under multiple postures and various lighting conditions.
[0055] In this embodiment, inputting the image dataset of the sample bovine into the bovine face segmentation algorithm that combines the RGB image and the depth image to segment the bovine face from the background of the RGB image and the depth image to obtain the RGB bovine face image and the depth bovine face image specifically includes:
[0056] First, use the image segmentation algorithm based on the RGB image to segment the RGB image. After segmenting the ideal area, replace the background with black and keep the foreground part. Then, use the segmentation algorithm based on the depth image for segmentation to further remove the inaccurate parts caused by only relying on color segmentation in the RGB image segmentation algorithm.
[0057] The segmentation algorithm based on the depth image is the main body of the bovine face segmentation algorithm that combines the RGB image and the depth image. The depth image segmentation algorithm relies on depth image threshold segmentation, effectively avoiding mis-segmentation caused by the same color of the object and the background color. In the segmentation algorithm based on the depth image, the method of obtaining the threshold using the histogram can accurately obtain the threshold in the case where the object color is the same as or similar to the background color, which is suitable for the bovine face segmentation scenario.
[0058] The bovine face segmentation algorithm that combines the RGB image and the depth image includes: obtaining a rectangular frame through the segmentation algorithm based on the depth image, using the rectangular frame for the segmentation algorithm based on the RGB image, and using the segmentation algorithm based on the depth image again after the segmentation algorithm based on the RGB image, and fusing the results of the two algorithms. As Figure 3 shown, the detailed steps are as follows:
[0059] S201. Cluster the pixels from dark to bright in the depth map D in the image dataset of the sample cows:
[0060] f1(D) = H
[0061] The f1 function represents the clustering method, and H represents the array data of bright and dark pixels obtained after clustering. A histogram of the depth map is obtained based on the clustering data, as Figure 4 shown, Figure 4 The left side is the depth map, and the right side is the corresponding histogram;
[0062] S202. Use the valley acquisition algorithm based on continuous wavelet transform to segment the valleys of the histogram of the depth map. Calculate the coordinate array of the cow face rectangular frame in the depth map according to the threshold of the last valley, and finally obtain the cow face rectangular frame in the RGB map according to the cow face rectangular frame coordinate array;
[0063] In the cow face depth map, the target is the cow face closest to the depth camera. In the histogram of the depth map, due to the complexity of the shooting scene, there are multiple waves, and objects with continuous depth will form a single wave separately. In the shooting environment of the cow face, each point of the cow face has continuous depth and is discontinuous with the background information such as the cow body, cowshed or pasture. Therefore, the cow face appears as a single waveform in the histogram. Also, because the cow's head is closer to the depth camera compared to the background information such as the cow body, cowshed or pasture; therefore, the last wave in the histogram of the depth map is the wave of the cow face, and the position of the last larger valley is the threshold for cow face segmentation. As Figure 5 shown, the green is the background and the red is the cow face.
[0064] In actual applications, it is found that the distribution of peaks and valleys in the depth map is not obvious and there are noise points. The noise points are higher or lower than the surrounding pixel points, forming small peaks or small valleys. The valley acquisition algorithm based on continuous wavelet transform will be affected and recognize the noise points as large valleys. Therefore, use the interpolation smoothing method to smooth the histogram curve:
[0065] First, set the step size covered by each smoothing operation, and then for a certain point x on the histogram, calculate its smoothed value temp:
[0066]
[0067] The smoothed histogram is as Figure 6 shown. The frequency value of each pixel point in the smoothed histogram will be affected and deviate from the true value. However, the histogram is only used to obtain the threshold, so the deviation from the true value will not affect image segmentation. After smoothing, the larger peaks and valleys are retained, and the noise points and small valleys are covered;
[0068] Generally, when determining peaks and valleys, the background and targets are distinguished by the method of extreme values. However, in the application here, the last valley is not necessarily the minimum value of the entire histogram, and the last peak is not necessarily the maximum value. Therefore, the method of extreme values is not applicable to this scenario. Here, a peak detection method based on continuous wavelet transform is used instead:
[0069] C(a, b) = ∫s(t)ψ a,b (t)dt,
[0070]
[0071] In the formula, ψ a,b (t) represents the scaled and transformed wavelet, ψ(t) is the mother wavelet, a ∈ R + is the scale of the mother wavelet scaling, b ∈ R is the distance of the mother wavelet translation, s(t) represents the signal, and C represents the two-dimensional matrix of wavelet coefficients.
[0072] In the detection of peaks, in order to obtain better performance, the wavelet should have the basic characteristics of a peak, including approximate symmetry and a main positive peak. The Mexican Hat wavelet is used as the mother wavelet. In 3D space, taking the amplitude of the CWT coefficients as the third dimension to visualize the 2D CWT coefficients, the problem of peak detection can be transformed into the problem of finding the ridge line on the 2D CWT coefficient matrix.
[0073] The ridge line finding algorithm is an algorithm for obtaining peaks, that is: initialize the ridge line. For the ridge lines with a spacing less than a certain threshold, find the maximum point of the next adjacent scale in the (n - 1)-th row of the coefficient matrix. If not found, the number of gaps in the ridge line is incremented by 1. Save the ridge lines with a spacing greater than the threshold and delete them from the search list. Take the maximum point that is not connected to the upper layer point as the new ridge line. Repeat the above steps. If the length of the ridge line is greater than a certain threshold, and the scale corresponding to the maximum amplitude on the ridge line should be within a certain range, and the scale is proportional to the width of the peak, it is determined as a peak.
[0074] In this embodiment, instead of obtaining peaks, each valley needs to be obtained. Therefore, the algorithm for obtaining peaks is improved to an algorithm for obtaining valleys based on continuous wavelet transform. The specific idea is to invert the histogram, and then use the ridge line finding method to obtain the peaks after inversion, so that the valleys in the original histogram can be obtained. The specific approach is to take the negative value of the number of pixels for each brightness value in the histogram, and after smoothing, use the ridge line finding method to obtain the coordinates of each valley in the histogram;
[0075] After obtaining the valleys, calculate the coordinate array of the cow face rectangular frame in the RGB image according to the threshold of the last valley, which is specifically expressed as:
[0076] f2(H1) = arr_rectangle[4] = [x1, y1, x2, y2]
[0077] The function f2 represents the above-mentioned algorithm for calculating the threshold of the histogram valley. H1 is the H data after smoothing. This step is mainly to obtain the coordinates arr_rectangle of the bovine face rectangle in the image, that is, to obtain the upper-left position coordinates (x1, y1) and the lower-right position coordinates (x2, y2) of the rectangle;
[0078] In the threshold calculation method, for the histogram as shown Figure 5 below, the OTSU method is compared with the method of this paper. Among these three comparison methods, the bovine face depth map histogram dataset is used to calculate the threshold. Taking Figure 5 as an example, after measuring and analyzing the original depth map, the threshold between 210 and 220 is the appropriate segmentation boundary for the bovine face. The threshold results of different methods are shown in Table 1:
[0079] Table 1 Thresholds Obtained by Different Algorithms
[0080] Segmentation algorithm Threshold OTSU algorithm 126 The algorithm in this paper without using the smoothing algorithm 231 The algorithm in this paper 216
[0081] The last large wave in the histogram is the wave where the bovine face is located. The threshold obtained by the OTSU method is less than the threshold where the bovine face is located. The above three thresholds are segmented, and the results are as follows Figure 7 shown. When the position of the threshold is before the last large wave valley, too much background information is included in the foreground after segmentation. When the threshold position is after the last large wave valley, an "over-segmentation" phenomenon will occur. The threshold of this method is at the position of the last large wave valley, and the segmentation effect is the best.
[0082] S203. Perform the GrabCut segmentation algorithm on the bovine face in the RGB image within the bovine face rectangle to obtain the preliminarily segmented bovine face image A; specifically including:
[0083] First, model the color data in the RGB image R, then use the method of iterative energy minimization for segmentation, and obtain the framed target rectangle according to the interaction with the user; then replace the framed target rectangle with the obtained bovine face rectangle art_rectangle; finally, replace all the background with black according to the f3 function of the GrabCut segmentation algorithm, specifically expressed as:
[0084] f3(R, art_rectangle) = R1 ∪ R3
[0085] Both R1 and R3 are sub-regions of R. R1 represents the bovine face region in the image, and R3 represents the non-bovine face region segmented due to the limitations of the GrabCut segmentation algorithm.
[0086] S204. Using the brightness of the depth map D at the last trough in S202 as a threshold, segment the depth map. The part greater than the threshold is retained, and the part less than the threshold is removed and replaced with black as the background color, which is specifically expressed as:
[0087] f4(D) = D1 ∪ D2
[0088] The f4 function is a depth map segmentation algorithm. D1 and D2 are subsets of the depth map D. D1 represents the cow face area in the depth map, and D2 represents the non-cow face area segmented due to the limitations of the depth map segmentation algorithm.
[0089] Then, obtain the coordinates of the foreground pixel points of the depth map, and segment the RGB map based on the coordinates to obtain the preliminarily segmented cow face map B, which is specifically expressed as:
[0090] f5(D1 ∪ D2) = R1 ∪ R2
[0091] The f5 function will segment the RGB map according to the cow face coordinates obtained by the depth map segmentation algorithm. R2 is a sub-region of R, representing the non-cow face area redundantly segmented in the RGB map according to the depth map segmentation algorithm.
[0092] S205. Take the intersection of the preliminarily segmented cow face map A in S203 and the preliminarily segmented cow face map B in S204.
[0093] (R1 ∪ R3) ∩ (R1 ∪ R2) = R1
[0094] Obtain the complete RGB cow face image R1, obtain the coordinates of the foreground pixel points of the complete RGB cow face image, and segment the depth map based on the coordinates to obtain the complete depth cow face image D1.
[0095] To verify the effect of the cow face segmentation algorithm combining RGB map and depth map proposed in this embodiment in the scenario where the object is close to the camera, and it is necessary to focus on specific targets in the image, and segment the background and foreground to remove the influence of the background. In the cow face segmentation scenario, it is compared with the contour detection segmentation algorithm and the OTSU algorithm. Among them, the contour detection segmentation algorithm only uses the RGB image data in this dataset for segmentation, and the OTSU algorithm only uses the depth image data in this dataset for segmentation.
[0096] To objectively evaluate the segmentation effect of the algorithm, PA (Pixel Accuracy) is used as an evaluation index. The cow face data of this embodiment is used for comparison. Since there is no correct image segmentation standard for cow face pictures, the pictures are manually segmented, the background part is removed, and the foreground part is retained as the standard segmentation image. Then, the pictures segmented by the cow face segmentation algorithm that combines the RGB image and the depth image are compared with the standard segmentation image. Calculate PA, that is, the proportion of correctly segmented pixels in the total pixels. The comparison results are shown in Table 2:
[0097] Table 2 Comparison of segmentation results of different segmentation algorithms
[0098] Segmentation algorithm PA Contour detection segmentation algorithm 60.21% OTSU algorithm 53.39% Cattle face segmentation algorithm combining RGB image and depth image 97.13%
[0099] As can be seen from Table 2, the existing image segmentation methods cannot effectively segment cow faces and cannot accurately judge background and foreground information. The segmentation effect of this embodiment is closer to the standard segmentation and can remove most of the cow face background.
[0100] In this embodiment, labels are assigned to the cow face data set. The label form is an array. The array subscript is the sample pair number, and the array value is 0 or 1. Assume that X1 and X2 are two sample pairs to be learned respectively, and Y is the label of the sample pair. When the sample pair composed of X1 and X2 comes from the same cow, the sample pair is matched, and the label Y is set to 1, indicating that they come from the same sample. When the sample pair composed of X1 and X2 comes from different cows, the sample pair is not matched, and the label Y is set to 0, indicating that they come from different samples.
[0101] The sample cow face data set has a total of C different cows, and each cow has E pairs of RGB images and depth images. When taking positive sample pairs, different pairs of RGB images and depth images (picture groups) of the same cow are numbered, and sample pairs are formed between different pairs of RGB images and depth images of the same cow based on the idea of "combination" in "permutation and combination". When taking negative sample pairs, different sample cows are numbered, and the negative sample pairs are made to come from different cows based on the idea of "combination".
[0102] In this embodiment, common image recognition uses two-dimensional images for recognition. The scenario of cow face recognition is different from that of human face recognition. The application scenarios of human face recognition are mostly at the entrances and exits of buildings such as railway stations, dormitories, and office buildings, and the scenarios are mostly indoors. For outdoor human face recognition, there are facilities that block light, so that human face recognition is not affected by light changes. However, ranches are usually outdoors. When the sun is strong during the day, it is extremely vulnerable to sunlight, making the RGB image too bright and affecting recognition. When sunlight decreases at dusk or at night, the RGB image is too dark, and the lighting in the cowshed is relatively dim with a small lighting range. Therefore, there is generally a problem of reduced light at night. In two-dimensional cow face recognition, it is affected by factors such as cow posture and lighting. The effect of two-dimensional cow face recognition in a complex environment is worse than that in a normal environment. Two-dimensional cow face pictures can only reflect partial information of the cow face. When the cow face undergoes posture changes or lighting changes, the information in the two-dimensional pictures of the same cow will change significantly. In the process of cow individual recognition using the cow face recognition technology that fuses depth information and image information, the cow face contour and color and other features are extracted through the two-dimensional image recognition process as the basis for recognition. Then, using the depth information of the cow face, the length information of the cow face in the direction from the cow's forehead to the cow's nose is obtained. This information is used as an additional extracted feature to help improve the recognition effect of two-dimensional cow face recognition. In addition, the contour information of the cow face in the depth map is more complete and is not interfered by the body image with the same color as the face. When the natural light conditions change drastically, the depth map will not be affected. At the same time, this method is still effective when the lighting changes and has extremely high robustness.
[0103] In this embodiment, the cow face pictures in the sample cow face dataset collected with various cow face postures and various lighting conditions are recognized. 75% of the cow face data is randomly selected as the training set, and 25% of the cow face data is used as the test set;
[0104] The RGB cow face image and the depth cow face image of the sample cow after background segmentation are formed into a cow face picture pair and sent into the cow face recognition network for recognition. A siamese network is used as the backbone network for cow face recognition. The two sub-networks with shared weights are important components of the siamese network. The siamese network uses a convolutional neural network to map the original image to a high-dimensional feature space, which can reduce the influence brought by geometric distortion. The structure of the siamese network is as Figure 8 shown.
[0105] The siamese network takes two inputs, each of which receives a sample. Therefore, the data received by the siamese network is a pair of samples. When the sample pair is a positive sample pair, the two samples come from the same cow. When the sample pair is a negative sample pair, the two samples come from different cows. A pair of samples enter two networks for training respectively. Since the two network structures are the same, they are both represented by the function f(x). The weight sharing of the two networks in the siamese network ensures that two similar images will not be mapped to positions that are very far apart in the high-dimensional space after passing through their respective networks. A convolutional neural network is used as the two sub-networks of the siamese network. Weight sharing means that in the two convolutional neural networks, the weights of the convolutional kernels in the convolutional layer, the biases of this channel in the convolutional layer, the weights in the fully connected layer, and the biases in the fully connected layer and other trainable parameters are updated synchronously as the number of training epochs increases.
[0106] When the siamese network processes image information, the input samples x1 and x2 are both RGB images; when the siamese network processes three-dimensional modalities, the input samples x1 and x2 are both depth images. f(x) includes a convolutional layer, a pooling layer, a dropout layer, and a fully connected layer. During the training process, parameter sharing is used to obtain the distance between each pair of samples. The parameter is represented by w, and both f(x1) and f(x2) use w as the parameter. Use distance<f(x1), f(x2)> to represent the distance between the two samples finally obtained after being processed by the network. According to the pre-labeled tags, reduce the loss of samples in the same category and increase the loss of samples in different categories;
[0107] The fused data of the image information and the depth information are the RGB cow face image and the depth cow face image of the sample cow. The RGB cow face image and the depth cow face image are separately trained using the siamese network. After extracting features, multiply by the corresponding weights. Use the late decision-level fusion method to judge whether the input sample pair is the same cow after fusion. The training fusion process is shown by Figure 9 as shown.
[0108] Since the depth image is not affected by the change of illumination, the weight is determined by the illumination intensity. Therefore, the weight is calculated by the following method: convert the RGB cow face image to the HSV space, and use the value (V) as the basis for judging the illumination intensity. Select the value when the illumination is appropriate and the clarity of the photo is not affected by the illumination as the standard value. When the illumination intensity is too large and the light is too bright, the value is higher than the standard value; when the illumination intensity is too small and the light is too dim, the value is lower than the standard value. The specific weight representation is shown in the following formula:
[0109]
[0110] W d =1 - W R
[0111] In the formula, W RRepresents the weight of the RGB cattle face image, V S Represents the standard value of brightness, V P Represents the brightness value of the RGB cattle face image, W d Represents the weight of the depth map. The greater the deviation of the brightness of the picture from the standard value, the greater the impact of light on the RGB image. Therefore, the weight of the RGB image is smaller, and the weight of the depth map is larger. At this time, the recognition effect relying on the RGB image is affected, and the depth map supplements the recognition of the RGB image under poor lighting conditions. On the contrary, the closer the brightness of the picture is to the standard value, the smaller the impact of light on the RGB image. Therefore, the weight of the RGB image is larger, the recognition weight of the depth map decreases, and the depth map provides more three-dimensional spatial feature information as an auxiliary.
[0112] The fusion adopts late decision-level fusion, and uses the voting method in deep learning to combine the above weights for fusion. The voting method is mostly used in classical classification and recognition networks to output the probability that a sample belongs to each class at the last fully connected layer. In this embodiment, the siamese network judges whether a sample pair belongs to the same cow based on the feature distance between samples. Therefore, the siamese network is first used to train the data of the two modalities separately, and then the voting method is used during decision-making to obtain the prediction result by weighted averaging the feature distance of different modalities output by the network, judge whether the sample pair belongs to the same cow, and then train the network according to the prediction result.
[0113] Through the fusion of deep learning and image information, when there is a deviation in the recognition prediction of a single modality, the other modality can correct the deviation and improve the overall recognition rate to complete the complementarity between modalities. The two-dimensional cattle face has a prediction deviation because it cannot obtain depth information, and the prediction of the depth map cattle face can correct the deviation. When the lighting condition is not ideal, the two-dimensional cattle face cannot obtain enough features and misrecognition occurs, and the prediction of the depth map cattle face can correct the deviation;
[0114] In this embodiment, the effects of using early fusion and late decision fusion are compared. Early fusion is to fuse the two-dimensional image and the three-dimensional image by adding weights to the images, and then use a neural network for training and recognition. The early fusion method used in this example is to directly fuse the RGB cattle face image and the depth cattle face image, and input the fused image into the siamese network for training to judge whether the sample pair comes from the same cow. The comparison results are as Figure 10 shown. The recognition rate of early fusion is 89.682%, and the recognition rate of late decision fusion is 93.619%. Early fusion does not perform feature extraction, and direct fusion causes pixel superposition and feature coverage, making it difficult to extract the features belonging to the cattle face in a single modality in subsequent feature extraction, causing certain interference in the scenario of cattle face recognition;
[0115] Compare the recognition rates of three recognition algorithms. The first is to use only the RGB images collected in S1 for cow face recognition, which belongs to the current conventional method; the second is to use the image dataset separated by the cow face segmentation algorithm that does not combine the RGB image and the depth image in S2, label it, and use it as the sample cow face dataset in S3 to train the cow face recognition network that fuses depth information and image information for cow recognition; the third is the recognition method disclosed in this embodiment; the comparison of the recognition results is as Figure 11 shown. Compared with the recognition using only RGB images, after introducing the recognition method that combines the depth image and the RGB image, the recognition rate has increased by 2.403%. Compared with the recognition method that combines the depth image and the RGB image, first perform image segmentation preprocessing and then perform the recognition that combines the depth image and the RGB image, the recognition rate has increased by 0.472%. Whether the recognition rate increases or remains unchanged, removing the background is necessary. Removing the influence of the background can make the neural network only focus on the features of the cow face in the picture and restore the recognition that only relies on the differences in cow face features. And in the application stage, the segmentation algorithm that removes the background is still used to ensure that the model is still effective when the background in the pasture changes unpredictably. In addition, the image segmentation method proposed in this paper is very suitable for application in scenarios where the object is close to the camera and the background needs to be ignored. For the subsequent image recognition field, it provides a new idea of introducing a depth image for segmentation and using a histogram to obtain the threshold, which is of great significance.
[0116] Embodiment 2
[0117] As Figure 2 shown, this embodiment also provides a cow face recognition device that fuses depth information, including: an acquisition module for obtaining paired RGB images and depth images of the sample cow face and making an image dataset of the sample cow;
[0118] a processing module for inputting the image dataset of the sample cow into the cow face segmentation algorithm that combines the RGB image and the depth image, segmenting the cow face from the background of the RGB image and the depth image to obtain an RGB cow face image and a depth cow face image, then forming a cow face picture pair according to the RGB cow face image and the depth cow face image of the sample cow, and constructing a sample cow face dataset from the cow face picture pairs of the samples. The sample cow face dataset includes cow face RGB images and depth images of multiple sample pairs. Each sample pair of cow face RGB images and depth images includes cow face RGB images and depth images of two samples. The two sample cows come from different cows or the same cow. Each sample pair of cow face RGB images and depth images is labeled, and the label is used to classify whether the two sample cow pairs come from the same cow;
[0119] a training module for inputting the sample cow face dataset into the cow face recognition network that fuses depth information and image information for training until the cow face recognition network distinguishes the pictures of different sample cows in the sample cow face dataset from each other, ends the training, and obtains a trained cow face recognition network;
[0120] An identification module, which is used to first register the paired RGB images and depth images of each cow in the cattle farm, select one cow as the cow to be identified, obtain the paired RGB image and depth image of the cow to be identified, segment them into the RGB cow face image and depth cow face image of the cow to be identified through a cow face segmentation algorithm that combines the RGB image and the depth image, and then input them into the trained cow face recognition network to obtain the identity information of the cow to be identified.
[0121] Embodiment 3
[0122] This embodiment provides an electronic device, which includes:
[0123] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the cow face recognition method for fusing depth information described in Embodiment 1.
[0124] Embodiment 4
[0125] This embodiment provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are operated to execute the cow face recognition method for fusing depth information described in Embodiment 1.
[0126] As mentioned above, these are only the preferred embodiments of the present invention and do not impose any limitations on the present invention. Any simple modification, change, and equivalent change made to the above embodiments according to the technical essence of the invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A cow face recognition method integrating deep learning, characterized in that, Including: S1. Obtain paired RGB images and depth images of the sample cow's face, and make an image dataset of the sample cow; S2. Input the image dataset of the sample cow into a cow face segmentation algorithm that combines RGB images and depth images, segment the cow face from the background of the RGB image and the depth image to obtain an RGB cow face image and a depth cow face image, and then form a cow face picture pair according to the RGB cow face image and the depth cow face image of the sample cow. Construct a sample cow face dataset from the sample cow face picture pairs. The sample cow face dataset includes RGB images and depth images of multiple sample pairs of cow faces. The RGB images and depth images of each sample pair of cow faces include RGB images and depth images of two sample cow faces. The two sample cows come from different cows or the same cow. The RGB images and depth images of each sample pair of cow faces are labeled, and the label is used to classify whether the two sample cows are from the same cow; S3. Input the sample cow face dataset into a cow face recognition network that fuses depth information and image information for training until the cow face recognition network distinguishes the pictures of different sample cows in the sample cow face dataset, and end the training to obtain a trained cow face recognition network; S4. First, register the paired RGB images and depth images of each cow in the cattle farm, select one cow as the cow to be recognized, obtain the paired RGB images and depth images of the cow to be recognized, segment them into an RGB cow face image and a depth cow face image of the cow to be recognized through a cow face segmentation algorithm that combines RGB images and depth images, and then input them into the trained cow face recognition network to obtain the identity information of the cow to be recognized; When inputting the image dataset of the sample cow into a cow face segmentation algorithm that combines RGB images and depth images to segment the cow face from the background of the RGB image and the depth image to obtain an RGB cow face image and a depth cow face image, it specifically includes: S201. Cluster and display the pixels from dark to bright in the depth image of the image dataset of the sample cow to obtain a histogram of the depth image; S202. Use an algorithm for obtaining wave valleys based on continuous wavelet transform to segment the wave valleys of the histogram of the depth image in S201, calculate the coordinate array of the cow face rectangular frame in the depth image according to the threshold of the last wave valley, and finally obtain the cow face rectangular frame in the RGB image according to the coordinate array of the cow face rectangular frame in the depth image; S203. Perform the GrabCut segmentation algorithm on the RGB image of the cow face within the cow face rectangular frame in the RGB image in S202 to obtain a preliminarily segmented cow face picture A; S204. Based on the brightness of the depth image where the last wave valley is located in S202 as the threshold, segment the depth image, retain the part greater than the threshold, remove the part less than the threshold and replace it with black as the background color; then obtain the coordinates of the foreground pixel points of the depth image, and segment the RGB image according to the coordinates to obtain a preliminarily segmented cow face picture B; S205. Take the intersection of the bovine face image A preliminarily segmented in S203 and the bovine face image B preliminarily segmented in S204 to obtain a complete RGB bovine face image R1. Obtain the coordinates of the foreground pixel points of the complete RGB bovine face image, and segment the depth map according to the coordinates to obtain a complete depth bovine face image D1.
2. The bovine face recognition method integrating deep learning according to claim 1, wherein In S1, the depth camera is a binocular camera, and the depth map and RGB map of the same picture of the sample bovine face are collected, that is, the paired RGB map and depth map of the sample bovine face.
3. The bovine face recognition method integrating deep learning according to claim 1, characterized in that, In S2, the sample pair is a positive sample pair or a negative sample pair. When the sample pair is a positive sample pair, it means that the two sample bovines of the sample pair come from the same cow. When the sample pair is a negative sample pair, it means that the two sample bovines of the sample pair come from different cows; The form of the label described in S2 is an array, where the subscript of the array is the sample pair number and the array value is 0 or 1; specifically, assume and are respectively sample pairs of two sample cows, is the label of the sample pair, and When the sample pair composed of comes from the same cow, the sample pair is matched and is a positive sample pair, and the label is set to 1, indicating from the same sample; and When the sample pair composed of comes from different cows, the sample pair is not matched and is a negative sample pair, and the label is set to 0, indicating from different samples.
4. A bovine face recognition method integrating deep learning according to claim 3, characterized in that, In S3, the bovine face recognition network that fuses depth information and image information uses a siamese network as the backbone network for bovine face recognition. The two sub-networks with shared weights are important components of the siamese network. The siamese network uses a convolutional neural network to map the original image to a high-dimensional feature space; Weight sharing means that in the two convolutional neural networks, the weights of the convolutional kernels in the convolutional layer, the biases of the channels in the convolutional layer, the weights in the fully connected layer, and the biases in the fully connected layer, the trainable parameters are synchronously updated as the number of training epochs increases.
5. The bovine face recognition method integrating deep learning according to claim 4, wherein When the Siamese network processes image information, the input samples and are both RGB images; when the Siamese network processes three-dimensional modalities, the input samples and are both depth images; It includes convolutional layers, pooling layers, dropout layers, and fully connected layers; during the training process, parameter sharing is used to obtain the distance between each pair of samples, and the parameter is represented by w, and both use w as the parameter; use to represent the distance between the two samples finally obtained after being processed by the network; according to the pre-labeled tags, reduce the loss of samples in the same category and increase the loss of samples in different categories.
6. The bovine face recognition method integrating deep learning according to claim 5, wherein The RGB cattle face image and the depth cattle face image are separately trained using a siamese network. After extracting features, corresponding weights are multiplied. Finally, a late decision-level fusion method is used to obtain the prediction result by taking the average of the sample distances output by the network. The average value is calculated to obtain the prediction result. The weight is calculated by the following method: convert the RGB bovine face image to the HSV space, use the lightness as the basis for judging the light intensity, and select the lightness with appropriate light and the clarity of the photo not affected by the light as the standard value. The weight is expressed by the following formula: ; ; In the formula, represents the weight of the RGB cattle face image, represents the standard value of lightness, represents the lightness value of the RGB cattle face image, represents the weight of the depth map.
7. A bovine face recognition device integrating deep learning, characterized in that, Including: An acquisition module for obtaining the paired RGB map and depth map of the sample bovine face and making an image data set of the sample bovine; A processing module for inputting the image data set of the sample bovine into a bovine face segmentation algorithm that combines the RGB map and the depth map, segmenting the bovine face from the background of the RGB map and the depth map to obtain an RGB bovine face image and a depth bovine face image, and then forming a bovine face picture pair according to the RGB bovine face image and the depth bovine face image of the sample bovine, constructing a sample bovine face data set from the bovine face picture pairs of the samples. The sample bovine face data set includes the bovine face RGB maps and depth maps of multiple sample pairs. The bovine face RGB maps and depth maps of each sample pair include the bovine face RGB maps and depth maps of two samples. The two sample bovines come from different cows or the same cow. The bovine face RGB maps and depth maps of each sample pair are labeled, and the label is used to classify whether the two sample bovine pairs come from the same cow; A training module for inputting the sample bovine face data set into the bovine face recognition network that fuses depth information and image information for training until the bovine face recognition network distinguishes the pictures of different sample bovines in the sample bovine face data set, ends the training, and obtains a trained bovine face recognition network; An identification module is used to first register the paired RGB images and depth images of each cow in the cattle farm, select one cow as the cow to be identified, obtain the paired RGB image and depth image of the cow to be identified, segment them into the RGB cow face image and depth cow face image of the cow to be identified through a cow face segmentation algorithm that combines the RGB image and the depth image, and then input them into the trained cow face recognition network to obtain the identity information of the cow to be identified; The processing module is specifically used to perform the following steps: S201. Cluster and display the pixels from dark to bright in the depth image in the image dataset of the sample cow to obtain the histogram of the depth image; S202. Use the valley acquisition algorithm based on continuous wavelet transform to segment the valleys of the histogram of the depth image in S201, calculate the coordinate array of the cow face rectangle frame in the depth image according to the threshold of the last valley, and finally obtain the cow face rectangle frame in the RGB image according to the coordinate array of the cow face rectangle frame in the depth image; S203. Perform the GrabCut segmentation algorithm on the cow face within the cow face rectangle frame in the RGB image in S202 to obtain the preliminarily segmented cow face image A; S204. Based on the brightness of the depth image where the last valley is located in S202 as the threshold, segment the depth image, keep the part greater than the threshold, remove the part less than the threshold and replace it with black as the background color; then obtain the coordinates of the foreground pixel points of the depth image, and segment the RGB image according to the coordinates to obtain the preliminarily segmented cow face image B; S205. Take the intersection of the preliminarily segmented cow face image A in S203 and the preliminarily segmented cow face image B in S204 to obtain the complete RGB cow face image R1, obtain the coordinates of the foreground pixel points of the complete RGB cow face image, and segment the depth image according to the coordinates to obtain the complete depth cow face image D1.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a cow face recognition method integrating deep learning as described in any one of claims 1-6.
9. A computer-readable storage medium stores computer instructions, and the computer instructions are operated to execute a cow face recognition method integrating deep learning as described in any one of claims 1-6.
Citation Information
Patent Citations
Face recognition method and device
CN113139465A