A method for target segmentation, completion and recognition in an occluded environment

Through the improved interpolation algorithm and multi-eigen convolution neural network model, the problem of slow completion and recognition speed and poor accuracy of picking targets in the occlusion environment is solved, and fast and accurate picking target recognition is achieved.

CN115984558BActive Publication Date: 2025-06-27NANJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211683146.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-06-27
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

The prior art problems of slow speed and poor accuracy in the segmentation and completion of picking targets under the occlusion environment.

Method used

The improved interpolation algorithm is used to complete the picking targets, and the picking targets in the occlusion environment are identified by combining the multi-feature convolutional neural network model.

Benefits of technology

It realizes the rapid and accurate identification of picking targets in the occlusion environment, and improves the accuracy and efficiency of picking target segmentation and completion and identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984558B_ABST
    Figure CN115984558B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for target segmentation completion and recognition in an occluded environment, including: collecting two consecutive video sequence images in an occluded environment, converting the images from the RGB color space to the YCbCr color space, extracting the Y component and performing normalization processing; performing mean clustering segmentation to obtain two segmentation images in the occluded environment; performing edge extraction and edge tracking, implementing contour reconstruction of the picking target through an improved interpolation algorithm, and then calculating and generating a fitting rectangle according to the segmentation images with contour reconstruction in the two occluded environments; calculating the relative accuracy, identifying the occluded repair image, and obtaining the segmentation completion and recognition image of the picking target in the occluded environment. The present invention provides an effective method for target segmentation completion and recognition in an occluded environment, and innovatively uses an improved interpolation algorithm to complete the contour of the picking target, having the advantages of strong practicability, high contour completion accuracy, and strong anti-background environment interference ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of agricultural picking, and relates to object recognition, instance segmentation and deep learning technologies, and particularly relates to a method for object segmentation completion and recognition in an occlusion environment. Background Art

[0002] Currently, most agricultural pickings are carried out manually one by one. The picking process takes a long time, has a large labor intensity and high labor costs. Using machines to replace manual picking has become an urgent problem to be solved in agricultural production. In recent years, with the development of computer vision technology, some scientific researchers have proposed a picking object segmentation completion and recognition system based on image processing and visual analysis, and it has been widely used in the agricultural field.

[0003] Most of the current object segmentation completion and recognition methods in an occlusion environment are for the segmentation completion and recognition of a single specific picking object, and do not involve the segmentation completion and recognition of multiple general objects. Such a single specific picking object segmentation completion and recognition method has many problems in practical applications, such as slow picking object completion and recognition and poor recognition accuracy.

[0004] Therefore, a new technical solution is needed to solve this problem. Summary of the Invention

[0005] Object of the Invention: In order to solve the problems of slow speed and poor accuracy in the segmentation completion and recognition of picking objects during unmanned picking, a method for object segmentation completion and recognition in an occlusion environment is provided, which provides an effective method for the segmentation completion and recognition of picking objects in an occlusion environment, and innovatively uses an improved interpolation algorithm to complete the contour of the picking object, and constructs a method for object segmentation completion and recognition in an occlusion environment, which has the advantages of strong practicability, high contour completion accuracy and strong anti-background environment interference ability.

[0006] Technical Solution: To achieve the above object, the present invention provides a method for object segmentation completion and recognition in an occlusion environment, including the following steps:

[0007] S1: Two consecutive video sequence images in an occlusion environment are collected through a camera, the two collected consecutive video sequence images are converted from the RGB space to the YCbCr color space, and the Y component is extracted for normalization processing to obtain two normalized images in the occlusion environment;

[0008] S2: The two normalized images in the occlusion environment are subjected to mean clustering segmentation to obtain two segmentation images in the occlusion environment;

[0009] S3: Edge extraction and edge tracking are performed on the two segmented images in the occluded environment. The contour reconstruction of the picking target is achieved by improving the interpolation algorithm. Then, a fitting rectangle is calculated and generated based on the segmented images with contour reconstruction in the two occluded environments, and an occluded repair image containing the fitting rectangle in the two occluded environments is obtained.

[0010] S4: Calculate the relative accuracy of the two occluded repair images containing the fitting rectangle. If the relative accuracy is qualified, use the trained multi-feature convolutional neural network model to identify the two occluded repair images containing the fitting rectangle, and obtain the segmented and completed recognition image of the picking target in the occluded environment, realizing the segmented and completed recognition of the target in the occluded environment.

[0011] Further, the step S1 specifically includes the following steps:

[0012] A1: Collect the real-time video in the occluded environment through the camera. Perform video sequence frame division on the obtained real-time video in the occluded environment, set to extract one frame every 2 seconds to obtain two consecutive video sequence images, and store the two consecutive video sequence images.

[0013] A2: According to the formula, convert the two consecutive video sequence images from the RGB color space to the YCbCr color space, extract the Y component, and obtain two luminance signal images in the occluded environment.

[0014] A3: Perform normalization processing on the two luminance signal images in the occluded environment to obtain two normalized images in the occluded environment.

[0015] Further, the formula for converting the image from the RGB color space to the YCbCr color space and extracting the Y component in the step A2 is:

[0016]

[0017] Y = (0.257 * R) + (0.504 * G) + (0.098 * B) + 16

[0018] Where: R is the red component of the two consecutive video sequence images; G is the green component of the two consecutive video sequence images; B is the blue component of the two consecutive video sequence images; Cb is the blue offset component image in the two occluded environments; Cr is the red offset component image in the two occluded environments; Y is the luminance signal image in the two occluded environments.

[0019] Further, the formula for the normalization processing in the step A3 is as follows:

[0020]

[0021] Where: X input(λ,μ)is the pixel value of the pixel at the λ-th row and μ-th column of the luminance signal images in two occluded environments; max(X input ) is the maximum pixel value of all pixels in the luminance signal images in two occluded environments; min(X input ) is the minimum pixel value of all pixels in the luminance signal images in two occluded environments; X output(λ,μ) is the normalized image pixel value of the λ-th row and μ-th column in two occluded environments output.

[0022] Further, the mean clustering segmentation in step S2 adopts K-means clustering segmentation, which specifically includes the following:

[0023] For K-means clustering segmentation, the number of clusters k needs to be specified manually, and k clustering centers are randomly generated. According to the distance from each data object in the normalized images in two occluded environments to the clustering centers, through repeated iterative operations, the positions of the clustering centers are calculated. Considering that there are picking targets and branches and leaves in the occluded environment, to reduce the calculation amount, the number of clusters k is specified as 2, and the K-means clustering segmentation formula is:

[0024]

[0025]

[0026] Where: D is the pixel value sample set of each pixel in the normalized images in two occluded environments; D1 is the pixel value of the first data object in the normalized images in two occluded environments; D2 is the pixel value of the second data object in the normalized images in two occluded environments; DN is the pixel value of the N-th data object in the normalized images in two occluded environments; U is the set of clustering centers in the normalized images in two occluded environments; U1 and U2 are the two specified clustering centers in the normalized images in two occluded environments respectively; C is the set of clusters in the normalized images in two occluded environments; C1 and C2 are the two specified clusters in the normalized images in two occluded environments respectively; CI is the I-th cluster in the set of clusters in the normalized images in two occluded environments; Z is the element of the set of clusters in the normalized images in two occluded environments; minE is the shortest distance from the sample point to the clustering center of the cluster.

[0027] K-means clustering segmentation can achieve dynamic clustering and has a certain self-adaptability. It can still effectively segment the normalized images in two occluded environments under different light intensities and obtain the segmented images in two occluded environments.

[0028] Further, step S3 specifically includes the following steps:

[0029] B1: Edge extraction and edge tracking are performed on the two segmented images in the occluded environment. Canny edge extraction is used for edge extraction, and eight-neighbor edge tracking is used for edge tracking to obtain the segmented images of edge extraction and edge tracking in the two occluded environments.

[0030] B2: The contour reconstruction of the picking target is realized by improving the interpolation algorithm, and the fitting rectangle frame is calculated and generated according to the picking target after contour reconstruction to obtain the occluded repair images with the fitting rectangle frames in the two occluded environments.

[0031] Further, the specific steps of step B1 are as follows:

[0032] C1: Gaussian filtering processing is performed on the two segmented images in the occluded environment according to the two-dimensional Gaussian filtering formula, and the high-frequency noise superimposed in the picking image is effectively filtered to obtain the segmented images after filtering processing in the two occluded environments.

[0033] The two-dimensional Gaussian filtering formula is as follows:

[0034]

[0035] Where: g(a, b) is the weighted average weight value of the pixel at the a-th row and b-th column of the pixel to be filtered; A is the weighted proportionality coefficient; u a represents that the pixel to be filtered is located in the u-th a row pixel of the segmented image in the two occluded environments; u b represents that the pixel to be filtered is located in the u-th b column pixel of the segmented image in the two occluded environments; σ a is the standard deviation calculated row by row for the segmented image in the two occluded environments; σ b is the standard deviation calculated column by column for the segmented image in the two occluded environments;

[0036] C2: In Canny edge detection, first, the image gradient of the two segmented images after filtering processing in the occluded environment is calculated, and then the non-maximum suppression operation is performed on the two segmented images after filtering processing in the occluded environment to obtain the segmented images after the non-maximum suppression operation in the two occluded environments.

[0037] Among them, the formula for image gradient calculation is:

[0038]

[0039]

[0040] Where: η is the gray value matrix of the segmented images in the occluded environment after two filtering processes; G1 is the gray value of the edge detection of the segmented images after filtering in the occluded environment for two horizontally adjacent images; G2 is the gray value of the edge detection of the segmented images after filtering in the occluded environment for two vertically adjacent images; G is the horizontal and vertical gray values of each pixel of the segmented images after filtering in the occluded environment for two adjacent images.

[0041] C3: Perform double-threshold detection and edge connection on the segmented images after non-maximum suppression operation in the occluded environment for two adjacent images to obtain the edge-extracted segmented images in the occluded environment for two adjacent images.

[0042] The formula for double-threshold detection and edge connection is as follows:

[0043]

[0044] Where: th1 and th2 are two set thresholds respectively; maxVal is used to determine whether to connect to the pixel points with values of th1 or th2. If so, assign 1, otherwise assign 0; src(α, β) is the pixel point value of the α-th row and β-th column of the segmented images after non-maximum suppression operation in the occluded environment for two adjacent images; dst(α, β) is the pixel point value of the α-th row and β-th column of the edge-extracted segmented images in the occluded environment for two adjacent images.

[0045] C4: Perform eight-neighborhood edge tracking on the edge-extracted segmented images in the occluded environment for two adjacent images. The eight-neighborhood edge tracking is to perform edge tracking clockwise, traverse the edge-extracted segmented images in the occluded environment for two adjacent images, find the non-zero pixel points, start from this point, and clockwise find the next non-zero pixel point encountered within the eight-neighborhood of this point, and iterate one by one until the last non-zero pixel point is found to complete the eight-neighborhood edge tracking, and obtain the segmented images of edge extraction and edge tracking in the occluded environment for two adjacent images.

[0046] Further, the specific step B2 is as follows:

[0047] D1: For the segmented images of edge extraction and edge tracking in the occluded environment for two adjacent images, extract the pixel point coordinates of the edges therein, use the following improved interpolation algorithm to reconstruct the contour of the picking target, and fill the contour to obtain the segmented images of contour reconstruction in the occluded environment for two adjacent images. The formula of the improved interpolation algorithm is as follows:

[0048] U = {(x o , y o ), (x1, y1)...(x i , y i )}

[0049] h i = x i - x i-1

[0050]

[0051] S i (x i ) = y i

[0052] S' i (x i ) = S' i-1 (x i )

[0053] m n = S'(x n )

[0054]

[0055] Among them, U is the set of edge pixel point coordinates of the segmented images for edge extraction and edge tracking in the two occluded environments; i is the number of edge pixel point coordinates of the segmented images for edge extraction and edge tracking in the two occluded environments; (x i , y i ) is the i-th edge pixel point coordinate of the segmented images for edge extraction and edge tracking in the two occluded environments; x i is the abscissa of the i-th edge pixel point of the segmented images for edge extraction and edge tracking in the two occluded environments; y i is the ordinate of the i-th edge pixel point of the segmented images for edge extraction and edge tracking in the two occluded environments; h i is the step size corresponding to the i-th edge pixel point of the segmented images for edge extraction and edge tracking in the two occluded environments; S(x) is the edge pixel point coordinate function of the segmented images for edge extraction and edge tracking in the two occluded environments; S i (x i ) is the function value of the segmented images for edge extraction and edge tracking in the two occluded environments at the i-th edge pixel point coordinate; S' i (x i ) is the first derivative value of the edge pixel point coordinate function of the segmented images for edge extraction and edge tracking in the two occluded environments at the abscissa of the i-th edge pixel point; n is the number of edge pixel point coordinates of the segmented images for contour reconstruction in the two occluded environments; S'(x n ) is the first derivative value of the edge pixel point coordinate function of the segmented images for contour reconstruction in the two occluded environments at the abscissa of the n-th edge pixel point; m n is the n-th first derivative value of the edge pixel point coordinate function of the segmented images for contour reconstruction in the two occluded environments; h n is the step size corresponding to the n-th edge pixel point of the segmented images for contour reconstruction in the two occluded environments; y n is the ordinate of the n-th edge pixel point of the segmented images for contour reconstruction in the two occluded environments;

[0056] D2: Calculate and generate a fitted rectangular box based on the segmented images reconstructed from the contours in two occluded environments, and obtain the occluded restoration images containing the fitted rectangular box in two occluded environments; the formula for the fitted rectangular box is:

[0057] u′ = {(x0, y0), (x1, y1)...(x p , y p )}

[0058] l = max{max(x p ) - min(x p ), max(y p ) - min(y p )}

[0059] w = min{max(x p ) - min(x p ), max(y p ) - min(y p )}

[0060] Among them, U′ is the set of edge pixel point coordinates of the segmented images reconstructed from the contours in two occluded environments; p is the number of edge pixel point coordinates of the segmented images reconstructed from the contours in two occluded environments; l is the length of the fitted rectangular box, and w is the width of the fitted rectangular box.

[0061] Further, the step S4 specifically includes the following steps:

[0062] E1: Calculate the relative accuracy of the occluded restoration images containing the fitted rectangular box in two occluded environments. The formula for the relative accuracy is:

[0063]

[0064] Among them: ε is the relative accuracy; TC is the number of pixel points in the intersection of the contours of the occluded restoration images containing the fitted rectangular box in two occluded environments; FC is the number of pixel points in the non - intersection of the contours of the occluded restoration images containing the fitted rectangular box in two occluded environments; TF is the number of pixel points in the intersection of the fitted rectangular boxes of the occluded restoration images containing the fitted rectangular box in two occluded environments; FF is the number of pixel points in the non - intersection of the fitted rectangular boxes of the occluded restoration images containing the fitted rectangular box in two occluded environments;

[0065] E2: If the relative accuracy is greater than 95%, use the trained multi - feature convolutional neural network model to identify the two occluded restoration images containing the fitted rectangular box. The structure composition and formula of the multi - feature convolutional neural network are as follows:

[0066] Extract the convolutional features of the data according to the following two formulas, and the function is to obtain the The convolutional feature map of a data unit, Taking values from 0 to m in sequence, input from the source data set, each data unit obtained is a two-dimensional array of n1*8. Take the first 6 columns as the input of the first part and send it to the input layer. The input layer selects 16 3*3 convolutional kernels, with a stride of 1 and a padding number of 1. The output size formula of the first convolutional layer is as follows:

[0067]

[0068] Among them, Z is the length of the convolutional output data; W is the length of the convolutional input data; P is the padding number; F is the length of the convolutional kernel; v represents the stride;

[0069] The output data size of the first convolutional layer calculated by the above formula is n1*6*16; use the rectified linear unit function to send the data to the second convolutional layer. The second convolutional layer uses 32 3*3 convolutional kernels, with a stride of 1 and a padding number of 1. Then, according to the output size formula of the convolutional layer, the output size of the second convolutional layer is n1*6*32; after the second convolutional layer, also use the rectified linear unit function to send it to the pooling layer. The pooling layer slides with a 2*2 rectangular window. The output size calculation of the pooling layer is as follows:

[0070]

[0071] Z′ is the length of the pooling layer output; W′ is the length of the pooling layer input; F′ is the length of the filter; v′ represents the stride in the horizontal direction;

[0072] The output size of the pooling layer obtained by the above formula is 50*3*32. Obtain the next convolutional layer and the next pooling layer in the same way. Then, according to the output size formula of the convolutional layer, the output size of the second convolutional layer is n1*96*6; according to the output size formula of the pooling layer, the output size of the convolutional layer in the second part is 50*48*6; perform vector transformation processing on the above output and the feature data set to obtain a three-layer fully connected network, and then obtain a multi-feature convolutional neural network;

[0073] E3: Use the trained multi-feature convolutional neural network model to identify two occluded repair images containing fitting rectangular frames, obtain the picking target segmentation completion and recognition image under the occluded environment, and realize the target segmentation completion and recognition under the occluded environment.

[0074] In the process of picking target segmentation, the clustering segmentation adopted by the method of the present invention can achieve dynamic clustering, and can still effectively segment the normalized images in two occluded environments under different light intensities; in the process of picking target completion, the improved interpolation algorithm adopted has universality, and can still achieve the contour reconstruction of the picking target under different picking target shapes and the presence of different occluders, and the accuracy is improved twice by relative accuracy; in the process of picking target recognition, the multi-feature convolutional neural network model adopted uses the data of adjacent moments to jointly analyze and recognize the picking target, and can still achieve the recognition of the picking target when the picking targets are blurred and overlapped. The method for target segmentation, completion and recognition in an occluded environment proposed by the present invention is of great significance for accurately and quickly recognizing picking target segmentation and improving the accuracy of picking target completion and recognition.

[0075] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0076] (1) The clustering segmentation adopted can achieve dynamic clustering, and can still effectively segment the normalized images in two occluded environments under different light intensities;

[0077] (2) In the process of picking target completion, the improved interpolation algorithm adopted has universality, and can still achieve the contour reconstruction of the picking target under different picking target shapes and the presence of different occluders, and the accuracy is improved twice by relative accuracy;

[0078] (3) In the process of picking target recognition, the multi-feature convolutional neural network model adopted uses the data of adjacent moments to jointly analyze and recognize the picking target, and can still achieve the recognition of the picking target when the picking targets are blurred and overlapped. Description of the drawings

[0079] Figure 1 is the flowchart of the method of the present invention;

[0080] Figure 2 is the continuous video sequence image obtained by the technology provided in the embodiment of the present invention in apple and orange orchards;

[0081] Figure 3 is the luminance signal image obtained by the technology provided in the embodiment of the present invention in apple and orange orchards;

[0082] Figure 4 is the normalized image obtained by the technology provided in the embodiment of the present invention in apple and orange orchards;

[0083] Figure 5 is the segmented image obtained by the technology provided in the embodiment of the present invention in apple and orange orchards;

[0084] Figure 6 are edge extraction and edge tracking images obtained by the technology provided in the embodiments of the present invention in apple and orange orchards;

[0085] Figure 7 are contour completion images obtained by the technology provided in the embodiments of the present invention in apple and orange orchards;

[0086] Figure 8 are segmentation completion and recognition images obtained by the technology provided in the embodiments of the present invention in apple and orange orchards. Detailed implementation manners

[0087] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.

[0088] The present invention provides a method for target segmentation completion and recognition in an occlusion environment, as Figure 1 shown, which includes the following steps:

[0089] S1: Collect two consecutive video sequence images in an occlusion environment through a camera, convert the two collected consecutive video sequence images from the RGB space to the YCbCr color space, extract the Y component for normalization processing, and obtain two normalized images in the occlusion environment;

[0090] S2: Perform mean clustering segmentation on the two normalized images in the occlusion environment to obtain two segmentation images in the occlusion environment;

[0091] S3: Perform edge extraction and edge tracking on the two segmentation images in the occlusion environment, implement contour reconstruction of the picking target through an improved interpolation algorithm, and then calculate and generate a fitting rectangle according to the segmentation images with contour reconstruction in the two occlusion environments to obtain two occlusion repair images with fitting rectangles in the occlusion environment;

[0092] S4: Calculate the relative accuracy of the two occlusion repair images with fitting rectangles in the occlusion environment. If the relative accuracy is qualified, use the trained multi-feature convolutional neural network model to identify the two occlusion repair images with fitting rectangles to obtain the picking target segmentation completion and recognition images in the occlusion environment, and realize the target segmentation completion and recognition in the occlusion environment.

[0093] To verify the actual effect of the method of the present invention, the method of the present invention is applied in this embodiment, specifically applied to apple and orange orchards, and the specific process is as follows:

[0094] A method for target segmentation completion and recognition in an occluded environment, comprising the following steps:

[0095] S1: Collect two consecutive video sequence images in an occluded environment through a camera, convert the two collected consecutive video sequence images from the RGB color space to the YCbCr color space, extract the Y component for normalization processing, and obtain two normalized images in the occluded environment;

[0096] S2: Perform mean clustering segmentation on the two normalized images in the occluded environment to obtain two segmented images in the occluded environment;

[0097] S3: Perform edge extraction and edge tracking on the two segmented images in the occluded environment, implement contour reconstruction of the picking target through an improved interpolation algorithm, and then calculate and generate a fitting rectangle according to the segmented images with contour reconstruction in the two occluded environments to obtain two occluded repair images with fitting rectangles in the occluded environment;

[0098] S4: Calculate the relative accuracy of the two occluded repair images with fitting rectangles in the occluded environment. If the relative accuracy is qualified, use the trained multi-feature convolutional neural network model to identify the two occluded repair images with fitting rectangles to obtain the picking target segmentation completion and recognition image in the occluded environment, and realize target segmentation completion and recognition in the occluded environment.

[0099] Step S1 specifically includes the following steps:

[0100] A1: Collect real-time video in an occluded environment through a camera. Perform video sequence frame division on the obtained real-time video in the occluded environment, set to extract one frame every 2 seconds to obtain two consecutive video sequence images, and store the two consecutive video sequence images;

[0101] A2: According to the formula, convert the two consecutive video sequence images from the RGB color space to the YCbCr color space, extract the Y component, and obtain two luminance signal images in the occluded environment;

[0102] A3: Perform normalization processing on the two luminance signal images in the occluded environment to obtain two normalized images in the occluded environment.

[0103] The formula for converting the image from the RGB color space to the YCbCr color space and extracting the Y component in step A2 is:

[0104]

[0105] Y = (0.257 * R) + (0.504 * G) + (0.098 * B) + 16

[0106] Where: R is the red component of two consecutive video sequence images; G is the green component of two consecutive video sequence images; B is the blue component of two consecutive video sequence images; Cb is the blue offset component image in two occluded environments; Cr is the red offset component image in two occluded environments; Y is the luminance signal image in two occluded environments.

[0107] The formula for the normalization process in step A3 is as follows:

[0108]

[0109] Where: X input(λ,μ) is the pixel value of the pixel at the λ-th row and μ-th column of the luminance signal image in two occluded environments; max(X input ) is the maximum pixel value of all pixel points of the luminance signal image in two occluded environments; min(X input ) is the minimum pixel value of all pixel points of the luminance signal image in two occluded environments; X output(λ,μ) is the pixel value of the normalized image at the λ-th row and μ-th column of the output two occluded environments.

[0110] The mean clustering segmentation in step S2 uses K-means clustering segmentation, which specifically includes the following:

[0111] For K-means clustering segmentation, the number of clusters k needs to be specified manually, and k cluster centers are randomly generated. According to the distance from each data object in the normalized image in two occluded environments to the cluster centers, after repeated iterative calculations, the positions of the cluster centers are calculated. Considering that there are picking targets and branches and leaves in the occluded environment, to reduce the computational amount, the number of clusters k is specified as 2. The K-means clustering segmentation formula is:

[0112]

[0113]

[0114] Where: D is the pixel value sample set of each pixel point in the normalized images under two occluded environments; D1 is the pixel value of the first data object in the normalized images under two occluded environments; D2 is the pixel value of the second data object in the normalized images under two occluded environments; DN is the pixel value of the Nth data object in the normalized images under two occluded environments; U is the set of cluster centers in the normalized images under two occluded environments; U1 and U2 are two specified cluster centers in the normalized images under two occluded environments respectively; C is the set of clusters in the normalized images under two occluded environments; C1 and C2 are two specified clusters in the normalized images under two occluded environments respectively; CI is the Ith cluster in the set of clusters in the normalized images under two occluded environments; Z is an element of the set of clusters in the normalized images under two occluded environments; minE is the shortest distance from the sample point to the cluster center of the cluster.

[0115] K-means clustering segmentation can achieve dynamic clustering and has a certain self-adaptability. It can still effectively segment the normalized images under two occluded environments under different light intensities and obtain the segmented images under two occluded environments.

[0116] Step S3 specifically includes the following steps:

[0117] B1: Perform edge extraction and edge tracking on the segmented images under two occluded environments. Canny edge extraction is used for edge extraction, and eight-neighborhood edge tracking is used for edge tracking to obtain the segmented images of edge extraction and edge tracking under two occluded environments;

[0118] B2: Implement the contour reconstruction of the picking target through an improved interpolation algorithm, calculate and generate a fitting rectangle according to the picking target after contour reconstruction, and obtain the occlusion repair images containing the fitting rectangle under two occluded environments.

[0119] Step B1 is specifically:

[0120] C1: Perform Gaussian filtering on the segmented images under two occluded environments according to the two-dimensional Gaussian filtering formula, effectively filtering out the superimposed high-frequency noise in the picking images to obtain the segmented images after filtering under two occluded environments;

[0121] The two-dimensional Gaussian filtering formula is as follows:

[0122]

[0123] Where: g(a, b) is the weighted average weight value of the pixel point at the a-th row and b-th column from the pixel point to be filtered; A is the weighted proportional coefficient; u a represents that the pixel point to be filtered is at the u-th a row pixel point of the segmented images under two occluded environments; ub Indicates that the pixel to be filtered is located in the \(u\)-th column of the segmented images in two occluded environments b column of pixels; \(\sigma\) a is the standard deviation calculated row by row for the segmented images in two occluded environments; \(\sigma\) b is the standard deviation calculated column by column for the segmented images in two occluded environments;

[0124] C2: For Canny edge detection, first calculate the image gradient for the segmented images after filtering in two occluded environments, and then perform non-maximum suppression operations on the segmented images after filtering in two occluded environments to obtain the segmented images after non-maximum suppression operations in two occluded environments;

[0125] Among them, the formula for image gradient calculation is:

[0126]

[0127]

[0128] Among them: \(\eta\) is the gray value matrix of the segmented images after filtering in two occluded environments; \(G1\) is the gray value of edge detection for the segmented images after filtering horizontally in two occluded environments; \(G2\) is the gray value of edge detection for the segmented images after filtering vertically in two occluded environments; \(G\) is the horizontal and vertical gray values of each pixel in the segmented images after filtering in two occluded environments;

[0129] C3: Perform double-threshold detection to connect edges on the segmented images after non-maximum suppression operations in two occluded environments to obtain the edge-extracted segmented images in two occluded environments;

[0130] Among them, the formula for double-threshold detection to connect edges is:

[0131]

[0132] Among them: \(th1\) and \(th2\) are two set thresholds respectively; \(maxVal\) is used to determine whether to connect to the pixel points with values of \(th1\) or \(th2\), if so, assign 1, otherwise assign 0; \(src(\alpha,\beta)\) is the pixel point value of the \(\alpha\)-th row and \(\beta\)-th column of the segmented images after non-maximum suppression operations in two occluded environments; \(dst(\alpha,\beta)\) is the pixel point value of the \(\alpha\)-th row and \(\beta\)-th column of the edge-extracted segmented images in two occluded environments;

[0133] C4: Perform eight-neighborhood edge tracking on the edge-extracted and segmented images in two occluded environments. The eight-neighborhood edge tracking is to perform edge tracking in a clockwise direction, traverse the edge-extracted and segmented images in two occluded environments, find the non-zero pixel points, use that point as the starting point, and clockwise search for the next non-zero pixel point encountered within the eight-neighborhood of that point. Iterate one by one until the last non-zero pixel point is found, complete the eight-neighborhood edge tracking, and obtain the edge-extracted and edge-tracked segmented images in two occluded environments.

[0134] Step B2 is specifically as follows:

[0135] D1: For the edge-extracted and edge-tracked segmented images in two occluded environments, extract the pixel coordinates of the edges therein, use the following improved interpolation algorithm to reconstruct the contour of the picking target, and fill the contour to obtain the contour-reconstructed segmented images in two occluded environments. The formula of the improved interpolation algorithm is as follows:

[0136] U = {(x0, y0), (x1, y1)...(x i , y i )}

[0137] h i = x i -x i-1

[0138]

[0139] S i (x i ) = y i

[0140] S′ i (x i ) = S′ i-1 (x i )

[0141] m n = S′(x n )

[0142]

[0143] Among them, U is the set of edge pixel coordinates of the edge-extracted and edge-tracked segmented images in two occluded environments; i is the number of edge pixel coordinates of the edge-extracted and edge-tracked segmented images in two occluded environments; (x i , y i ) is the i-th edge pixel coordinate of the edge-extracted and edge-tracked segmented images in two occluded environments; x i is the abscissa of the i-th edge pixel of the edge-extracted and edge-tracked segmented images in two occluded environments; y iis the ordinate of the i-th edge pixel of the segmented image for edge extraction and edge tracking in the two occluded environments; h i is the step size corresponding to the i-th edge pixel of the segmented image for edge extraction and edge tracking in the two occluded environments; S(x) is the edge pixel coordinate function of the segmented image for edge extraction and edge tracking in the two occluded environments; S i (x i ) is the function value of the segmented image for edge extraction and edge tracking in the two occluded environments at the coordinate of the i-th edge pixel; S′ i (x i ) is the first derivative value of the edge pixel coordinate function of the segmented image for edge extraction and edge tracking in the two occluded environments at the abscissa of the i-th edge pixel; n is the number of edge pixel coordinates of the segmented image for contour reconstruction in the two occluded environments; S′(x n ) is the first derivative value of the edge pixel coordinate function of the segmented image for contour reconstruction in the two occluded environments at the abscissa of the n-th edge pixel; m n is the n-th first derivative value of the edge pixel coordinate function of the segmented image for contour reconstruction in the two occluded environments; h n is the step size corresponding to the n-th edge pixel of the segmented image for contour reconstruction in the two occluded environments; y n is the ordinate of the n-th edge pixel of the segmented image for contour reconstruction in the two occluded environments;

[0144] D2: Calculate and generate a fitting rectangle according to the segmented image for contour reconstruction in the two occluded environments, and obtain an occlusion repair image containing the fitting rectangle in the two occluded environments; the fitting rectangle formula is:

[0145] U′ = {(x0, y0), (x1, y1)...(x p , y p )}

[0146] l = max{max(x p ) - min(x p ), max(y p ) - min(y p )}

[0147] w = min{max(x p ) - min(x p ), max(y p ) - min(y p )}

[0148] where U′ is the set of edge pixel coordinates of the segmented image for contour reconstruction in the two occluded environments; p is the number of edge pixel coordinates of the segmented image for contour reconstruction in the two occluded environments; l is the length of the fitting rectangle and w is the width of the fitting rectangle.

[0149] Step S4 specifically includes the following steps:

[0150] E1: Calculate the relative accuracy of two occlusion repair images with fitted rectangular frames in an occluded environment. The formula for relative accuracy is:

[0151]

[0152] Where: ε is the relative accuracy; TC is the number of pixel points in the intersection of the contours of two occlusion repair images with fitted rectangular frames in an occluded environment; FC is the number of pixel points in the non-intersection of the contours of two occlusion repair images with fitted rectangular frames in an occluded environment; TF is the number of pixel points in the intersection of the fitted rectangular frames of two occlusion repair images with fitted rectangular frames in an occluded environment; FF is the number of pixel points in the non-intersection of the fitted rectangular frames of two occlusion repair images with fitted rectangular frames in an occluded environment;

[0153] E2: If the relative accuracy is greater than 95%, use the trained multi-feature convolutional neural network model to identify two occlusion repair images with fitted rectangular frames. The structure composition and formula of the multi-feature convolutional neural network are as follows:

[0154] Extract convolutional features from the data according to the following two formulas, and the function is to obtain the convolutional feature map of the th data unit, Taking values from 0 to m in sequence, input from the source data set, and each data unit obtained is a two-dimensional array of n1*8. Take the first 6 columns as the input of the first part and send it to the input layer. The input layer selects 16 3*3 convolutional kernels, with a stride of 1 and a padding quantity of 1. The output size formula of the first convolutional layer is as follows:

[0155]

[0156] Where, Z is the length of the convolutional output data; W is the length of the convolutional input data; P is the padding quantity; F is the length of the convolutional kernel; v represents the stride;

[0157] The output data size of the first convolutional layer calculated by the above formula is n1*6*16; use the rectified linear unit function to send the data to the second convolutional layer. The second convolutional layer uses 32 3*3 convolutional kernels, with a stride of 1 and a padding quantity of 1. Then, according to the output size formula of the convolutional layer, the output size of the second convolutional layer is n1*6*32; after the second convolutional layer, also use the rectified linear unit function to send it to the pooling layer. The pooling layer slides with a rectangular window of 2*2 size, and the output size calculation of the pooling layer is as follows:

[0158]

[0159] Z′ is the length of the output of the pooling layer; W′ is the length of the input of the pooling layer; F′ is the length of the filter; v′ represents the stride in the horizontal direction;

[0160] From the above formula, the output size of the pooling layer is 50*3*32. In the same way, the next convolutional layer and the next pooling layer are obtained. Then, according to the output size formula of the convolutional layer, the output size of the second convolutional layer is n1*96*6; according to the output size formula of the pooling layer, the output size of the second part of the convolutional layer is 50*48*6. The above outputs are vector-transformed with the feature dataset to obtain a three-layer fully connected network, and a multi-feature convolutional neural network is obtained.

[0161] E3: Use the trained multi-feature convolutional neural network model to identify two occluded repair images containing fitted rectangular frames, obtain the segmented and completed picking target recognition images in the occluded environment, and realize the segmented and completed target recognition in the occluded environment.

[0162] After the above process, in this embodiment, the Figures 2 to 8 related images are obtained:

[0163] Referring to Figure 2 、 Figure 3 and Figure 4 , it can be reflected that in step S1 of this embodiment, two consecutive video sequence images in the occluded environment are collected by the camera, converted from the RGB space to the YCbCr color space, and the Y component is extracted for normalization processing to obtain two normalized images in the occluded environment. The normalization processing effectively improves the convergence speed of the model and the accuracy of the model.

[0164] Referring to Figure 5 , it can be reflected that in step S2 of this embodiment, the two normalized images in the occluded environment are subjected to mean clustering segmentation to obtain two segmented images in the occluded environment. The clustering segmentation adopted can achieve dynamic clustering and can still effectively segment the two normalized images in the occluded environment under different light intensities, and has the advantage of strong anti-background environment interference ability.

[0165] Referring to Figure 6 and Figure 7 , it can be reflected that in step S3 of this embodiment, edge extraction and edge tracking are performed on the two segmented images in the occluded environment, the contour reconstruction of the picking target is realized by improving the interpolation algorithm, and then the fitted rectangular frame is calculated and generated according to the segmented images with contour reconstruction in the two occluded environments to obtain two occluded repair images containing fitted rectangular frames in the occluded environment. The improved interpolation algorithm adopted is universal and can still realize the contour reconstruction of the picking target under different picking target shapes and the presence of different occluding objects, and has the advantage of high contour completion accuracy.

[0166] Reference Figure 8 , it can be seen that in step S4 of this embodiment, the relative accuracy of two occluded repair images containing fitted rectangular frames in an occluded environment is calculated. If the relative accuracy is qualified, the trained multi-feature convolutional neural network model is used to identify the two occluded repair images containing fitted rectangular frames, and the segmented and completed recognition image of the picking target in the occluded environment is obtained, realizing the segmentation, completion and recognition of the target in the occluded environment. The relative accuracy is used to improve the accuracy for the second time, and the multi-feature convolutional neural network model uses the data at adjacent times to jointly analyze and identify the picking target, and the recognition of the picking target can still be realized when the picking targets are blurred and overlapped.

Claims

1. A method for target segmentation, completion and recognition in an occluded environment, characterized in that, It includes the following steps: S1: Collect two consecutive video sequence images in the occlusion environment through a camera, convert the two collected consecutive video sequence images from the RGB space to the YCbCr color space, extract the Y component for normalization processing, and obtain two normalized images in the occlusion environment; S2: Perform mean clustering segmentation on the two normalized images in the occlusion environment to obtain two segmented images in the occlusion environment; S3: Perform edge extraction and edge tracking on the two segmented images in the occlusion environment, realize the contour reconstruction of the picking target through an improved interpolation algorithm, and then calculate and generate a fitting rectangle according to the segmented images with contour reconstruction in the two occlusion environments to obtain two occlusion repair images with fitting rectangles in the occlusion environment; S4: Calculate the relative accuracy of the two occlusion repair images with fitting rectangles in the occlusion environment. If the relative accuracy is qualified, use the trained multi-feature convolutional neural network model to identify the two occlusion repair images with fitting rectangles to obtain the picking target segmentation completion and recognition images in the occlusion environment, and realize the target segmentation completion and recognition in the occlusion environment; The formula for normalization processing is as follows: Where: X input(λ,μ) is the pixel value of the pixel points of the λ-th row and μ-th column of the luminance signal images in two occluded environments; max(X input ) is the maximum pixel value of all pixel points of the luminance signal images in two occluded environments; min(X input ) is the minimum pixel value of all pixel points of the luminance signal images in two occluded environments; X output(λ,μ) are the normalized image pixel values of λ rows and μ columns in the two output occlusion environments; The mean clustering segmentation in step S2 adopts K-means clustering segmentation, which specifically includes the following: For K-means clustering segmentation, the number of clusters k needs to be specified manually, and k clustering centers are randomly generated. According to the distance from each data object in the two normalized images in the occlusion environment to the clustering centers, after repeated iterative operations, the positions of the clustering centers are calculated. Considering that there are picking targets and branches and leaves in the occlusion environment, in order to reduce the calculation amount, the number of clusters k is specified as 2. The K-means clustering segmentation formula is: Where: D is the pixel value sample set of each pixel point in the two normalized images in the occlusion environment; D1 is the pixel value of the first data object in the two normalized images in the occlusion environment; D2 is the pixel value of the second data object in the two normalized images in the occlusion environment; DN is the pixel value of the Nth data object in the two normalized images in the occlusion environment; U is the set of clustering centers in the two normalized images in the occlusion environment; U1 and U2 are the two specified clustering centers in the two normalized images in the occlusion environment; C is the set of clusters in the two normalized images in the occlusion environment; C1 and C2 are the two specified clusters in the two normalized images in the occlusion environment; CI is the Ith cluster in the set of clusters in the two normalized images in the occlusion environment; Z is the element of the set of clusters in the two normalized images in the occlusion environment; minE is the shortest distance from the sample point to the clustering center of the cluster.

2. The method for target segmentation completion and recognition in an occlusion environment according to claim 1, wherein, The specific steps of step S1 are as follows: A1: Collect real-time video in the occlusion environment through a camera. Perform video sequence frame division on the obtained real-time video in the occlusion environment, set to extract one frame every 2 seconds to obtain two consecutive video sequence images, and store the two consecutive video sequence images; A2: Convert two consecutive video sequence images from the RGB color space to the YCbCr color space according to the formula, extract the Y component, and obtain two luminance signal images in the occluded environment. A3: Normalize the two luminance signal images in the occluded environment to obtain two normalized images in the occluded environment.

3. The method for target segmentation completion and recognition in an occlusion environment according to claim 2, characterized in that, The formula for converting the image from the RGB color space to the YCbCr color space and extracting the Y component in step A2 is as follows: Y = (0.257 * R) + (0.504 * G) + (0.098 * B) + 16 Where: R is the red component of the two consecutive video sequence images; G is the green component of the two consecutive video sequence images; B is the blue component of the two consecutive video sequence images; Cb is the blue offset component image in the two occluded environments; Cr is the red offset component image in the two occluded environments; Y is the luminance signal image in the two occluded environments.

4. The object segmentation completion and recognition method in an occluded environment according to claim 1, characterized in that, Step S3 specifically includes the following steps: B1: Perform edge extraction and edge tracking on the two segmented images in the occluded environment. Canny edge extraction is used for edge extraction, and eight-neighborhood edge tracking is used for edge tracking to obtain two segmented images with edge extraction and edge tracking in the occluded environment. B2: Implement the contour reconstruction of the picking target through an improved interpolation algorithm, calculate and generate a fitting rectangle according to the picking target after contour reconstruction, and obtain two occluded repair images with fitting rectangles in the occluded environment.

5. The target segmentation completion and recognition method in an occluded environment according to claim 4, characterized in that, Step B1 is specifically: C1: Perform Gaussian filtering on the two segmented images in the occluded environment according to the two-dimensional Gaussian filtering formula, effectively filtering out the superimposed high-frequency noise in the picking image to obtain two filtered segmented images in the occluded environment. The two-dimensional Gaussian filtering formula is as follows: Where: g(a, b) is the weighted average weight of the pixel at the a-th row and b-th column from the pixel to be filtered; A is the weighted proportionality coefficient; u a represents that the pixel to be filtered is located at the u a -th row of the pixel points in the segmented images under two occlusion environments; u b represents that the pixel to be filtered is located at the u b -th column of the pixel points in the segmented images under two occlusion environments; σ a is the standard deviation calculated row by row for the segmented images under two occlusion environments; σ b is the standard deviation calculated column by column for the segmented images under two occlusion environments; C2: Canny edge detection first calculates the image gradient of the two filtered segmented images in the occluded environment, and then performs non-maximum suppression on the two filtered segmented images in the occluded environment to obtain two segmented images after non-maximum suppression in the occluded environment. Among them, the formula for image gradient calculation is: Where: η is the gray value matrix of the two filtered segmented images in the occluded environment; G1 is the gray value of the edge detection of the two filtered segmented images in the occluded environment horizontally; G2 is the gray value of the edge detection of the two filtered segmented images in the occluded environment vertically; G is the horizontal and vertical gray values of each pixel of the two filtered segmented images in the occluded environment. C3: Perform double-threshold detection and connect edges on the two segmented images after non-maximum suppression in the occluded environment to obtain two edge-extracted segmented images in the occluded environment. The formula for double-threshold detection and connecting edges is as follows: Where: th1 and th2 are two set thresholds respectively; maxVal is used to judge whether to connect with the pixel points with values of th1 or th2. If so, assign 1, otherwise assign 0; src(α,β) is the pixel point value of the α-th row and β-th column of the two segmented images after non-maximum suppression in the occluded environment; dst(α,β) is the pixel point value of the α-th row and β-th column of the two edge-extracted segmented images in the occluded environment. C4: Perform eight-neighborhood edge tracking on the edge extraction and segmentation images in two occluded environments. The eight-neighborhood edge tracking is to perform edge tracking in a clockwise direction. Traverse the edge extraction and segmentation images in two occluded environments, find the non-zero pixel points, take a certain point as the starting point, and clockwise search for the next non-zero pixel point encountered within the eight neighborhoods of this point. Iterate one by one until the last non-zero pixel point is found, complete the eight-neighborhood edge tracking, and obtain the segmentation images of edge extraction and edge tracking in two occluded environments.

6. The method for target segmentation, completion and recognition in an occlusion environment according to claim 4, wherein The specific steps of step B2 are as follows: D1: For the segmentation images of edge extraction and edge tracking in two occluded environments, extract the pixel coordinates of the edges therein, use the following improved interpolation algorithm to reconstruct the contour of the picking target, and fill the contour to obtain the segmentation images of contour reconstruction in two occluded environments. The formula of the improved interpolation algorithm is as follows: U = {(x0, y0), (x1, y1)…(x i , y i )} h i = x i - x i-1 S i (x i ) = y i S′ i (x i ) = S′ i-1 (x i ) m n = S'(x n ) Among them, U is the set of edge pixel point coordinates of the segmented images for edge extraction and edge tracking in the two occluded environments; i is the number of edge pixel point coordinates of the segmented images for edge extraction and edge tracking in the two occluded environments; (x i , y i ) is the coordinate of the i-th edge pixel point of the segmented images for edge extraction and edge tracking in the two occluded environments; x i is the abscissa of the i-th edge pixel point of the segmented images for edge extraction and edge tracking in the two occluded environments; y i is the ordinate of the i-th edge pixel point of the segmented images for edge extraction and edge tracking in the two occluded environments; h i is the step size corresponding to the i-th edge pixel point of the segmented images for edge extraction and edge tracking in the two occluded environments; S(x) is the edge pixel point coordinate function of the segmented images for edge extraction and edge tracking in the two occluded environments; S i (x i ) is the function value of the segmented images for edge extraction and edge tracking in the two occluded environments at the coordinate of the i-th edge pixel point; S' i (x i ) is the first derivative value of the edge pixel point coordinate function of the segmented images for edge extraction and edge tracking in the two occluded environments at the abscissa of the i-th edge pixel point; n is the number of edge pixel point coordinates of the segmented images for contour reconstruction in the two occluded environments; S'(x n ) is the first derivative value of the edge pixel point coordinate function of the segmented images for contour reconstruction in the two occluded environments at the abscissa of the n-th edge pixel point; m n is the n-th first derivative value of the edge pixel point coordinate function of the segmented images for contour reconstruction in the two occluded environments; h n is the step size corresponding to the n-th edge pixel point of the segmented images for contour reconstruction in the two occluded environments; y n is the ordinate of the n-th edge pixel point of the segmented images for contour reconstruction in the two occluded environments; D2: Calculate and generate a fitting rectangle box based on the segmentation images of contour reconstruction in two occluded environments to obtain the occlusion repair images containing the fitting rectangle box in two occluded environments; the formula of the fitting rectangle box is: U′ = {(x0,y0),(x1,y1)…(x p ,y p )} l = max{max(x p ) - min(x p ), max(y p ) - min(y p )} w = min{max(x p ) - min(x p ), max(y p ) - min(y p )} where U' is the set of edge pixel coordinates of the segmentation images of contour reconstruction in two occluded environments; P is the number of edge pixel coordinates of the segmentation images of contour reconstruction in two occluded environments; l is the length of the fitting rectangle box, and w is the width of the fitting rectangle box.

7. A method for target segmentation completion and recognition in an occluded environment according to claim 1, characterized in that, The specific steps of step S4 include the following steps: E1: Calculate the relative accuracy of the occlusion repair images containing the fitting rectangle box in two occluded environments. The formula of the relative accuracy is: where: ε is the relative accuracy; TC is the number of pixel points in the intersection of the contours of the occlusion repair images containing the fitting rectangle box in two occluded environments; FC is the number of pixel points in the non-intersection of the contours of the occlusion repair images containing the fitting rectangle box in two occluded environments; TF is the number of pixel points in the intersection of the fitting rectangle boxes of the occlusion repair images containing the fitting rectangle box in two occluded environments; FF is the number of pixel points in the non-intersection of the fitting rectangle boxes of the occlusion repair images containing the fitting rectangle box in two occluded environments; E2: If the relative accuracy is greater than 95%, use the trained multi-feature convolutional neural network model to identify the two occlusion repair images containing the fitting rectangle box. The structure composition and formula of the multi-feature convolutional neural network are as follows: Convolutional feature extraction is performed on the data according to the following two formulas, with the function of obtaining the convolutional feature map of the th data unit. Taking values from 0 to m in sequence, input from the source data set, each data unit obtained is a two-dimensional array of n1*8. The first 6 columns are taken as the input of the first part and sent to the input layer. The input layer selects 16 3*3 convolutional kernels, with a stride of 1 and a padding number of 1. The output size formula of the first convolutional layer is as follows: where Z is the length of the convolution output data; W is the length of the convolution input data; P is the padding quantity; F is the length of the convolution kernel; v represents the stride; The output data size of the first convolutional layer calculated by the above formula is n1*6*16; use the rectified linear unit function to send the data to the second convolutional layer. The second convolutional layer uses 32 3*3 convolutional kernels, the stride is 1, and the padding quantity is 1. Then, according to the output size formula of the convolutional layer, the output size of the second convolutional layer is b1*6*32; after the second convolutional layer, also use the rectified linear unit function to send it to the pooling layer. The pooling layer slides with a 2*2 rectangular window. The output size of the pooling layer is calculated as follows: Z' is the length of the output of the pooling layer; W' is the length of the input of the pooling layer; F' is the length of the filter; ν' represents the stride in the horizontal direction; The output size of the pooling layer obtained from the above formula is 50*3*32. In the same way, the next convolutional layer and the next pooling layer are obtained. Then, according to the output size formula of the convolutional layer, the output size of the second convolutional layer is n1*96*6; according to the output size formula of the pooling layer, the output size of the second part of the convolutional layer is 50*48*6. The above outputs are subjected to vector transformation processing with the feature dataset to obtain a three-layer fully connected network, thus obtaining a multi-feature convolutional neural network. E3: Use the trained multi-feature convolutional neural network model to identify two occluded repair images containing fitted rectangular frames, obtain the segmented and completed picking target recognition image under the occluded environment, and realize the segmentation, completion and recognition of the target under the occluded environment.

Citation Information

Patent Citations

  • A road traffic sign recognition method based on possibility clustering and a convolutional neural network

    CN109886161A

  • Traffic sign recognition method based on capsule neural network

    CN111428556A