Contour detection method simulating complex cell characteristics
By simulating the connection method of complex cells in biological vision, the cluster-like connection network is designed, combined with inhibition and multi-scale convolution modules, the feature extraction capability of the coding network is enhanced, and the problem of low edge detection efficiency on devices with limited resources in the existing technology is solved, and better edge information extraction effect is achieved.
Patent Information
- Application Number
- CN202310907287.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-24
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-07-24
AI Technical Summary
The existing contour detection network is inefficient in operating on devices with limited computing resources and ignores the feature extraction capability of the encoded network, resulting in insufficient extraction of edge information.
A cluster-shaped connection network structure that simulates the complex cell connection mode of primary visual cortex in biological vision is designed, combining the SM inhibition module and the MSSM module to enhance the feature extraction capability of the coding network and perform feature fusion through the decoding network.
It improves edge detection performance, can efficiently extract more detailed features and edge information on limited resource devices, achieving better edge detection effect.
Smart Images

Figure CN117274286B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a contour detection method for simulating complex cell characteristics. Background Art
[0002] Most existing contour detection networks require pre-training on the large ImageNet dataset, which consumes a large amount of computing resources. Furthermore, due to the large number of parameters and computational complexity, they run inefficiently or even fail to run on devices with limited computing and storage capabilities. Therefore, researchers have begun to focus on how to compress the number of network parameters and computational complexity without significantly sacrificing network performance. Existing CNN-based edge detection methods typically employ an encoding + decoding network structure, where the encoding network generally uses an existing network model and is primarily used to extract image features and edge information. The decoding network, designed by researchers, integrates features and edge information, ultimately extracting image edges. Because researchers focus on the design of the decoding network and ignore the impact of the encoding network in the encoding-decoding network, existing edge detection methods often suffer from weaknesses such as weak feature extraction and insufficient edge information extraction.
[0003] Deep Learning Methods: Existing neural network contour detection methods typically use CNNs (such as VGG and ResNet) or Transformers (such as ViT and SwinT) as encoding networks to extract multi-scale features of the image. A decoding network is then designed to fuse the extracted features to produce a contour image. These methods often require large-scale pre-training on the ImageNet dataset to achieve effective detection results, placing high demands on hardware resources. They also ignore the inspiration provided by biological vision mechanisms for neural networks, resulting in low interpretability.
[0004] Biological methods: Most biological learning methods are based on traditional manual algorithms to simulate the spatial antagonism mechanism of mutual modulation between the classical and non-classical receptive fields of simple cells. There is very little research on modeling complex cell characteristics. Summary of the Invention
[0005] The present invention aims to provide a contour detection method that simulates the characteristics of complex cells. This method simulates the connection mode of complex cells in the primary visual cortex in biological vision and designs a clustered connection network structure, thereby enhancing the feature extraction capability of the coding network, enabling the coding network to obtain more detailed features and edge information, and achieving better edge detection performance.
[0006] The technical solutions of the present invention are as follows:
[0007] The contour detection method for simulating complex cell characteristics comprises the following steps:
[0008] A. Construct a neural network. The neural network structure is as follows:
[0009] Includes encoding network and decoding network;
[0010] The encoding network has five layers. The first layer is a 5*5 convolution. The second, third, fourth, and fifth layers are respectively equipped with ICCM modules that simulate complex cells. The specific settings are as follows:
[0011] The second layer is provided with a first ICCM module 2 and a second ICCM module 2; the third layer is provided with a first ICCM module 3, a second ICCM module 3, and a third ICCM module 3; the fourth layer is provided with a first ICCM module 4, a second ICCM module 4, a third ICCM module 4, and a fourth ICCM module 4; the fifth layer is provided with a first ICCM module 5, a second ICCM module 5, a third ICCM module 5, a fourth ICCM module 5, and a fifth ICCM module 5; there are one, two, and three AFM modules between the second and third, third and fourth, and fourth and fifth layers, respectively;
[0012] B. The original image is input into the 5*5 convolution of the first layer. After the 5*5 convolution processing, it is input into the first ICCM module 2 and the second ICCM module 2 of the second layer respectively. The processing results of the first ICCM module 2 and the second ICCM module 2 are obtained. On the one hand, these two processing results are output to the decoding network. On the other hand, they are input into the third layer respectively. The processing result of the first ICCM module 2 is input into the first ICCM module 3, and the processing result of the second ICCM module 2 is input into the third ICCM module 3. The processing results of the first ICCM module 2 and the second ICCM module 2 are processed by the AFM module and input into the second ICCM module 3.
[0013] C. The third layer processes to obtain the processing results of the first ICCM module 3, the processing results of the second ICCM module 3, and the processing results of the third ICCM module 3, and these three processing results are respectively input into the decoding network; in addition, the processing result of the first ICCM module 3 is input into the first ICCM module 4 in the fourth layer, and the processing result of the third ICCM module 3 is input into the fourth ICCM module 4 in the fourth layer. The processing results of the first ICCM module 3 and the processing results of the second ICCM module 3 are processed by the AFM module and input into the second ICCM module 4 in the fourth layer. The processing results of the second ICCM module 3 and the processing results of the third ICCM module 3 are processed by the AFM module and input into the third ICCM module 4 in the fourth layer;
[0014] D. The fourth layer processes to obtain the processing results of the first ICCM module 4, the processing results of the second ICCM module 4, the processing results of the third ICCM module 4, and the processing results of the fourth ICCM module 4, and these four processing results are respectively input into the decoding network; in addition, the processing result of the first ICCM module 4 is input into the first ICCM module 5 in the fifth layer, the processing result of the fourth ICCM module 4 is input into the fifth ICCM module 5 in the fifth layer, the processing result of the first ICCM module 4 and the processing result of the second ICCM module 4 are input into the second ICCM module 5 in the fifth layer, the processing result of the second ICCM module 4 and the processing result of the third ICCM module 4 are input into the third ICCM module 5 in the fifth layer, and the processing result of the third ICCM module 4 and the processing result of the fourth ICCM module 4 are input into the fourth ICCM module 5 in the fifth layer;
[0015] E. The fifth layer processes the first ICCM module 5 processing result, the second ICCM module 5 processing result, the third ICCM module 5 processing result, the fourth ICCM module 5 processing result, and the fifth ICCM module 5 processing result, and these five processing results are respectively input into the decoding network;
[0016] D. The decoding network fuses all the input results to obtain the final output contour.
[0017] The first ICCM module 2, the second ICCM module 2, the first ICCM module 5, the second ICCM module 5, the third ICCM module 5, the fourth ICCM module 5, and the fifth ICCM module 5 have the same basic structure, and respectively include an SM1 module and an MSSM module; after the input result is processed by the SM1 module and the MSSM module in sequence, the MSSM module processing result is added and fused with the input result to obtain the output result;
[0018] The first ICCM module 3, the second ICCM module 3, the third ICCM module 3, the first ICCM module 4, the second ICCM module 4, the third ICCM module 4, and the fourth ICCM module 4 have the same basic structure, which respectively include an SM2 module, an MSSM module, and a 1*1 convolution; after the input result is processed by the SM2 module and the MSSM module in sequence, the MSSM module processing result with the number of channels doubled is obtained, and the MSSM module processing result is added and fused with the input result after the number of channels is adjusted to be consistent through the 1*1 convolution to obtain the output result.
[0019] The processing process in the SM1 module is as follows
[0020] The input structure is divided into three paths. The first path is processed by 5*5 convolution and LeakyReLU function in sequence. The second path is processed by 1*1 convolution and ReLU function in sequence. The third path is processed by distributed spatial attention module. The results of the second and third paths are subtracted and then subtracted from the results of the first path to obtain the output result. During the three-path processing, the number of image channels remains unchanged.
[0021] The processing process in the SM2 module is as follows
[0022] The input results are divided into three paths. The first path is processed by 5*5 convolution and its channel number is doubled, and then processed by LeakyReLU function; the second path is processed by 1*1 convolution and its channel number is doubled, and then processed by ReLU function; the third path is processed by distributed spatial attention module and its channel number is doubled; the second and third processing results are subtracted, and then subtracted from the first processing result to obtain the output result.
[0023] The processing process in the distributed spatial attention module is as follows:
[0024] The input result is divided into two paths. The first path is processed by 1*1 convolution, and the second path is processed by ReLU function, 1*1 convolution, 3*3 convolution, and Softmax function in sequence. The output result is obtained by multiplying the results of the two paths.
[0025] When corresponding to the SM1 module, the number of channels of the input result is not changed during the first and second processing; when corresponding to the SM2 module, the 1*1 convolution processing of the first and second paths doubles the number of channels of the input result.
[0026] The input result is processed by shift operation. The shift operation process is as follows: the channel of the input result is divided into eight parts, and each part is moved in a different spatial direction. The specific moving directions are: upper left, upper, upper right, right, lower right, lower, lower left, and left;
[0027] The result of the Shift operation is three-way. The results of moving to the upper left, upper right, lower left, and lower right are input to the first path 3*3 convolution processing, the results of moving to the left and right are input to the second path 1*3 convolution processing, and the results of moving up and down are input to the third path 3*1 convolution processing; the three-way processing results are spliced by the Concat function and then processed by 3*3 convolution to obtain the output result.
[0028] The decoding network is divided into four levels, each level has 1, 3, 6, and 10 AFM modules respectively. The processing process is as follows:
[0029] In the first level, the processing results of the first ICCM module 2 and the processing results of the second ICCM module 2 are input into the AFM module for processing to obtain AFM processing result 1;
[0030] In the second level, the processing results of the first ICCM module 3 and the second ICCM module 3 are input to the AFM module for processing, obtaining AFM processing result 2; the processing results of the second ICCM module 3 and the third ICCM module 3 are input to the AFM module for processing, obtaining AFM processing result 3; AFM processing result 2 and AFM processing result 3 are input to the AFM module for processing, obtaining AFM processing result 4;
[0031] In the third level, the processing results of the first ICCM module 4 and the processing results of the second ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 5; the processing results of the second ICCM module 4 and the processing results of the third ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 6; the processing results of the third ICCM module 4 and the processing results of the fourth ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 7; AFM processing result 5 and AFM processing result 6 are input into the AFM module for processing, obtaining AFM processing result 8; AFM processing result 6 and AFM processing result 7 are input into the AFM module for processing, obtaining AFM processing result 9; AFM processing result 8 and AFM processing result 9 are input into the AFM module for processing, obtaining AFM processing result 10;
[0032] In the fourth layer, the processing results of the first ICCM module 5 and the second ICCM module 5 are input to the AFM module for processing, and an AFM processing result 11 is obtained; the processing results of the second ICCM module 5 and the third ICCM module 5 are input to the AFM module for processing, and an AFM processing result 12 is obtained; the processing results of the third ICCM module 5 and the fourth ICCM module 5 are input to the AFM module for processing, and an AFM processing result 13 is obtained; the processing results of the fourth ICCM module 5 and the fifth ICCM module 5 are input to the AFM module for processing, and an AFM processing result 14 is obtained; the AFM processing results 11 and 12 are input to the AFM module. The FM module processes to obtain AFM processing result 15; the AFM processing result 12 and the AFM processing result 13 are input into the AFM module for processing to obtain AFM processing result 16; the AFM processing result 13 and the AFM processing result 14 are input into the AFM module for processing to obtain AFM processing result 17; the AFM processing result 15 and the AFM processing result 16 are input into the AFM module for processing to obtain AFM processing result 18; the AFM processing result 16 and the AFM processing result 17 are input into the AFM module for processing to obtain AFM processing result 19; the AFM processing result 18 and the AFM processing result 19 are input into the AFM module for processing to obtain AFM processing result 20;
[0033] AFM processing results 1, AFM processing results 4, AFM processing results 10, and AFM processing results 20 are adjusted to the same number of channels after 1*1 convolution processing, and then feature union is performed through the cancatenate module. Then, they are processed by 1*1 convolution and Sigmoid function in sequence to obtain the final output contour.
[0034] The processing process in the AFM module is as follows: the two input results are added and fused, and the output result is obtained through 1*1 convolution.
[0035] The method of the present invention combines the characteristics of complex cells in the primary visual cortex in the biological vision mechanism, proposes a bionic contour detection neural network, and designs a clustered connection network structure by simulating the connection mode of complex cells in the primary visual cortex in biological vision by simulating complex cell modules as simulated complex cell modules, thereby enhancing the feature extraction capability of the coding network and enabling the coding network to obtain more detailed features and edge information. An SM suppression module is designed in the neural network, which is divided into SM1 module and SM2 module corresponding to different layer results, thereby achieving the suppression effect on non-target edges. The method of the present invention also designs an MSSM module as a multi-scale convolutional space modeling module, which enhances the ability to extract contextual feature information and achieves better edge detection performance.
[0036] This method utilizes a clustered connection network to enhance the feature extraction capabilities of the encoding network. It utilizes a suppression module to highlight important and meaningful features in the image. Furthermore, the MSSM module integrates adjacent feature information without losing any edge information, enriching the edge information extracted by the model. As shown in Example 2, the method achieved an ODS score of 0.802 on the BSDS500 dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a schematic diagram of the structure of a neural network according to Example 1 of the present invention. The left side of the figure shows a clustered encoding network for extracting feature information at different scales from an image; the right side of the figure shows a decoding network for fusing features at different scales.
[0038] Figure 2 This is a schematic structural diagram of the ICCM module of Example 1 of the present invention, wherein the corresponding modules in the upper figure are: first ICCM module 2, second ICCM module 2, first ICCM module 5, second ICCM module 5, third ICCM module 5, fourth ICCM module 5, and fifth ICCM module 5; the corresponding modules in the lower figure are: first ICCM module 3, second ICCM module 3, third ICCM module 3, first ICCM module 4, second ICCM module 4, third ICCM module 4, and fourth ICCM module 4;
[0039] Figure 3 Schematic diagram of the structure of the SM1 module and the SM2 module of Example 1 of the present invention;
[0040] Figure 4 This is a schematic diagram of the MSSM module structure of Example 1 of the present invention;
[0041] Figure 5 This is a schematic diagram of the structure of the distributed spatial attention module Spatial attention in Example 1 of the present invention;
[0042] Figure 6 Figure 1 shows an example of a decoding network and an AFM module structure according to Embodiment 1 of the present invention. (a) shows the AFM module, (b) shows the primary decoding method for the side output of the second layer, and (c) shows the primary decoding method for the side output of the fifth layer.
[0043] Figure 7 Comparison of the qualitative results of contour detection on the BSDS500 test set. From top to bottom: (a) original RGB image, (b) ground truth, (c) HED, (d) Example 1. DETAILED DESCRIPTION
[0044] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0045] Example 1
[0046] The contour detection method for simulating complex cell characteristics comprises the following steps:
[0047] A. Construct a neural network. The neural network structure is as follows:
[0048] Includes encoding network and decoding network;
[0049] The encoding network has five layers. The first layer is a 5*5 convolution. The second, third, fourth, and fifth layers are respectively equipped with ICCM modules that simulate complex cells. The specific settings are as follows:
[0050] The second layer is provided with a first ICCM module 2 and a second ICCM module 2; the third layer is provided with a first ICCM module 3, a second ICCM module 3, and a third ICCM module 3; the fourth layer is provided with a first ICCM module 4, a second ICCM module 4, a third ICCM module 4, and a fourth ICCM module 4; the fifth layer is provided with a first ICCM module 5, a second ICCM module 5, a third ICCM module 5, a fourth ICCM module 5, and a fifth ICCM module 5; there are one, two, and three AFM modules between the second and third, third and fourth, and fourth and fifth layers, respectively;
[0051] B. The original image is input into the 5*5 convolution of the first layer. After the 5*5 convolution processing, it is input into the first ICCM module 2 and the second ICCM module 2 of the second layer respectively. The processing results of the first ICCM module 2 and the second ICCM module 2 are obtained. On the one hand, these two processing results are output to the decoding network. On the other hand, they are input into the third layer respectively. The processing result of the first ICCM module 2 is input into the first ICCM module 3, and the processing result of the second ICCM module 2 is input into the third ICCM module 3. The processing results of the first ICCM module 2 and the second ICCM module 2 are processed by the AFM module and input into the second ICCM module 3.
[0052] C. The third layer processes to obtain the processing results of the first ICCM module 3, the processing results of the second ICCM module 3, and the processing results of the third ICCM module 3, and these three processing results are respectively input into the decoding network; in addition, the processing result of the first ICCM module 3 is input into the first ICCM module 4 in the fourth layer, and the processing result of the third ICCM module 3 is input into the fourth ICCM module 4 in the fourth layer. The processing results of the first ICCM module 3 and the processing results of the second ICCM module 3 are processed by the AFM module and input into the second ICCM module 4 in the fourth layer. The processing results of the second ICCM module 3 and the processing results of the third ICCM module 3 are processed by the AFM module and input into the third ICCM module 4 in the fourth layer;
[0053] D. The fourth layer processes to obtain the processing results of the first ICCM module 4, the processing results of the second ICCM module 4, the processing results of the third ICCM module 4, and the processing results of the fourth ICCM module 4, and these four processing results are respectively input into the decoding network; in addition, the processing result of the first ICCM module 4 is input into the first ICCM module 5 in the fifth layer, the processing result of the fourth ICCM module 4 is input into the fifth ICCM module 5 in the fifth layer, the processing result of the first ICCM module 4 and the processing result of the second ICCM module 4 are input into the second ICCM module 5 in the fifth layer, the processing result of the second ICCM module 4 and the processing result of the third ICCM module 4 are input into the third ICCM module 5 in the fifth layer, and the processing result of the third ICCM module 4 and the processing result of the fourth ICCM module 4 are input into the fourth ICCM module 5 in the fifth layer;
[0054] E. The fifth layer processes the first ICCM module 5 processing result, the second ICCM module 5 processing result, the third ICCM module 5 processing result, the fourth ICCM module 5 processing result, and the fifth ICCM module 5 processing result, and these five processing results are respectively input into the decoding network;
[0055] D. The decoding network fuses all the input results to obtain the final output contour.
[0056] The first ICCM module 2, the second ICCM module 2, the first ICCM module 5, the second ICCM module 5, the third ICCM module 5, the fourth ICCM module 5, and the fifth ICCM module 5 have the same basic structure, and respectively include an SM1 module and an MSSM module; after the input result is processed by the SM1 module and the MSSM module in sequence, the MSSM module processing result is added and fused with the input result to obtain the output result;
[0057] The first ICCM module 3, the second ICCM module 3, the third ICCM module 3, the first ICCM module 4, the second ICCM module 4, the third ICCM module 4, and the fourth ICCM module 4 have the same basic structure, which respectively include an SM2 module, an MSSM module, and a 1*1 convolution; after the input result is processed by the SM2 module and the MSSM module in sequence, the MSSM module processing result with the number of channels doubled is obtained, and the MSSM module processing result is added and fused with the input result after the number of channels is adjusted to be consistent through the 1*1 convolution to obtain the output result.
[0058] The processing process in the SM1 module is as follows
[0059] The input structure is divided into three paths. The first path is processed by 5*5 convolution and LeakyReLU function in sequence. The second path is processed by 1*1 convolution and ReLU function in sequence. The third path is processed by distributed spatial attention module. The results of the second and third paths are subtracted and then subtracted from the results of the first path to obtain the output result. During the three-path processing, the number of image channels remains unchanged.
[0060] The processing process in the SM2 module is as follows
[0061] The input results are divided into three paths. The first path is processed by 5*5 convolution and its channel number is doubled, and then processed by LeakyReLU function; the second path is processed by 1*1 convolution and its channel number is doubled, and then processed by ReLU function; the third path is processed by distributed spatial attention module and its channel number is doubled; the second and third processing results are subtracted, and then subtracted from the first processing result to obtain the output result.
[0062] The processing process in the distributed spatial attention module is as follows:
[0063] The input result is divided into two paths. The first path is processed by 1*1 convolution, and the second path is processed by ReLU function, 1*1 convolution, 3*3 convolution, and Softmax function in sequence. The output result is obtained by multiplying the results of the two paths.
[0064] When corresponding to the SM1 module, the number of channels of the input result is not changed during the first and second processing; when corresponding to the SM2 module, the 1*1 convolution processing of the first and second paths doubles the number of channels of the input result.
[0065] The input result is processed by shift operation. The shift operation process is as follows: the channel of the input result is divided into eight parts, and each part is moved in a different spatial direction. The specific moving directions are: upper left, upper, upper right, right, lower right, lower, lower left, and left;
[0066] The result of the Shift operation is three-way. The results of moving to the upper left, upper right, lower left, and lower right are input to the first path 3*3 convolution processing, the results of moving to the left and right are input to the second path 1*3 convolution processing, and the results of moving up and down are input to the third path 3*1 convolution processing; the three-way processing results are spliced by the Concat function and then processed by 3*3 convolution to obtain the output result.
[0067] The decoding network is divided into four levels, each level has 1, 3, 6, and 10 AFM modules respectively. The processing process is as follows:
[0068] In the first level, the processing results of the first ICCM module 2 and the processing results of the second ICCM module 2 are input into the AFM module for processing to obtain AFM processing result 1;
[0069] In the second level, the processing results of the first ICCM module 3 and the second ICCM module 3 are input to the AFM module for processing, obtaining AFM processing result 2; the processing results of the second ICCM module 3 and the third ICCM module 3 are input to the AFM module for processing, obtaining AFM processing result 3; AFM processing result 2 and AFM processing result 3 are input to the AFM module for processing, obtaining AFM processing result 4;
[0070] In the third level, the processing results of the first ICCM module 4 and the processing results of the second ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 5; the processing results of the second ICCM module 4 and the processing results of the third ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 6; the processing results of the third ICCM module 4 and the processing results of the fourth ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 7; AFM processing result 5 and AFM processing result 6 are input into the AFM module for processing, obtaining AFM processing result 8; AFM processing result 6 and AFM processing result 7 are input into the AFM module for processing, obtaining AFM processing result 9; AFM processing result 8 and AFM processing result 9 are input into the AFM module for processing, obtaining AFM processing result 10;
[0071] In the fourth layer, the processing results of the first ICCM module 5 and the second ICCM module 5 are input to the AFM module for processing, and an AFM processing result 11 is obtained; the processing results of the second ICCM module 5 and the third ICCM module 5 are input to the AFM module for processing, and an AFM processing result 12 is obtained; the processing results of the third ICCM module 5 and the fourth ICCM module 5 are input to the AFM module for processing, and an AFM processing result 13 is obtained; the processing results of the fourth ICCM module 5 and the fifth ICCM module 5 are input to the AFM module for processing, and an AFM processing result 14 is obtained; the AFM processing results 11 and 12 are input to the AFM module. The FM module processes to obtain AFM processing result 15; the AFM processing result 12 and the AFM processing result 13 are input into the AFM module for processing to obtain AFM processing result 16; the AFM processing result 13 and the AFM processing result 14 are input into the AFM module for processing to obtain AFM processing result 17; the AFM processing result 15 and the AFM processing result 16 are input into the AFM module for processing to obtain AFM processing result 18; the AFM processing result 16 and the AFM processing result 17 are input into the AFM module for processing to obtain AFM processing result 19; the AFM processing result 18 and the AFM processing result 19 are input into the AFM module for processing to obtain AFM processing result 20;
[0072] AFM processing results 1, AFM processing results 4, AFM processing results 10, and AFM processing results 20 are adjusted to the same number of channels after 1*1 convolution processing, and then feature union is performed through the cancatenate module. Then, they are processed by 1*1 convolution and Sigmoid function in sequence to obtain the final output contour.
[0073] The processing process in the AFM module is as follows: the two input results are added and fused, and the output result is obtained through 1*1 convolution.
[0074] Example 2
[0075] For the performance evaluation of the contour map output by the network, we use the same performance measurement standard as in Reference 1. The specific evaluation is shown in Formula (1).
[0076]
[0077] Where P represents precision and R represents recall. A larger value of F indicates better performance.
[0078] Document 1: Xie S, Tu Z. Holistically-Nested Edge Detection[J]. International Journal of Computer Vision, 2017, 125(1-3):3.
[0079] The parameters used in Reference 1 are the same as those in the original text, and are guaranteed to be the optimal parameters for the model.
[0080] Figure 5 The following table shows, from top to bottom, four natural images randomly selected from the Berkeley Segmentation Dataset (BSDS500), the corresponding true contour map, the optimal contour map detected by the method in Reference 1, and the optimal contour detected by the method of the present invention. The performance comparison of the edge detection method provided in Example 1 and the edge detection method in Reference 1 is shown in Table 1 below:
[0081] Table 1 Performance comparison of the edge detection method provided by Example 1 and the edge detection method provided by Reference 1
[0082]
Claims
1. A contour detection method simulating complex cell characteristics, characterized in that: The following steps are involved: A. Construct a neural network. The neural network structure is as follows: Includes encoding network and decoding network; The encoding network has five layers. The first layer is a 5*5 convolution. The second, third, fourth, and fifth layers are respectively equipped with ICCM modules that simulate complex cells. The specific settings are as follows: The second layer is provided with a first ICCM module 2 and a second ICCM module 2; the third layer is provided with a first ICCM module 3, a second ICCM module 3, and a third ICCM module 3; the fourth layer is provided with a first ICCM module 4, a second ICCM module 4, a third ICCM module 4, and a fourth ICCM module 4; the fifth layer is provided with a first ICCM module 5, a second ICCM module 5, a third ICCM module 5, a fourth ICCM module 5, and a fifth ICCM module 5; there are one, two, and three AFM modules between the second and third, third and fourth, and fourth and fifth layers, respectively; B. The original image is input into the 5*5 convolution of the first layer. After the 5*5 convolution processing, it is input into the first ICCM module 2 and the second ICCM module 2 of the second layer respectively. The processing results of the first ICCM module 2 and the second ICCM module 2 are obtained. On the one hand, these two processing results are output to the decoding network. On the other hand, they are input into the third layer respectively. The processing result of the first ICCM module 2 is input into the first ICCM module 3, and the processing result of the second ICCM module 2 is input into the third ICCM module 3. The processing results of the first ICCM module 2 and the second ICCM module 2 are processed by the AFM module and input into the second ICCM module 3. C. The third layer processes to obtain the processing results of the first ICCM module 3, the processing results of the second ICCM module 3, and the processing results of the third ICCM module 3, and these three processing results are respectively input into the decoding network; in addition, the processing result of the first ICCM module 3 is input into the first ICCM module 4 in the fourth layer, and the processing result of the third ICCM module 3 is input into the fourth ICCM module 4 in the fourth layer. The processing results of the first ICCM module 3 and the processing results of the second ICCM module 3 are processed by the AFM module and input into the second ICCM module 4 in the fourth layer. The processing results of the second ICCM module 3 and the processing results of the third ICCM module 3 are processed by the AFM module and input into the third ICCM module 4 in the fourth layer; D. The fourth layer processes to obtain the processing results of the first ICCM module 4, the processing results of the second ICCM module 4, the processing results of the third ICCM module 4, and the processing results of the fourth ICCM module 4, and these four processing results are respectively input into the decoding network; in addition, the processing result of the first ICCM module 4 is input into the first ICCM module 5 in the fifth layer, the processing result of the fourth ICCM module 4 is input into the fifth ICCM module 5 in the fifth layer, the processing result of the first ICCM module 4 and the processing result of the second ICCM module 4 are input into the second ICCM module 5 in the fifth layer, the processing result of the second ICCM module 4 and the processing result of the third ICCM module 4 are input into the third ICCM module 5 in the fifth layer, and the processing result of the third ICCM module 4 and the processing result of the fourth ICCM module 4 are input into the fourth ICCM module 5 in the fifth layer; E. The fifth layer processes the first ICCM module 5 processing result, the second ICCM module 5 processing result, the third ICCM module 5 processing result, the fourth ICCM module 5 processing result, and the fifth ICCM module 5 processing result, and these five processing results are respectively input into the decoding network; F. The decoding network fuses all the input results to obtain the final output contour; The decoding network is divided into four levels, each level has 1, 3, 6, and 10 AFM modules, respectively. The processing process is as follows: In the first level, the processing results of the first ICCM module 2 and the processing results of the second ICCM module 2 are input into the AFM module for processing to obtain AFM processing result 1; In the second level, the processing results of the first ICCM module 3 and the second ICCM module 3 are input to the AFM module for processing, obtaining AFM processing result 2; the processing results of the second ICCM module 3 and the third ICCM module 3 are input to the AFM module for processing, obtaining AFM processing result 3; AFM processing result 2 and AFM processing result 3 are input to the AFM module for processing, obtaining AFM processing result 4; In the third level, the processing results of the first ICCM module 4 and the processing results of the second ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 5; the processing results of the second ICCM module 4 and the processing results of the third ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 6; the processing results of the third ICCM module 4 and the processing results of the fourth ICCM module 4 are input into the AFM module for processing, obtaining AFM processing result 7; AFM processing result 5 and AFM processing result 6 are input into the AFM module for processing, obtaining AFM processing result 8; AFM processing result 6 and AFM processing result 7 are input into the AFM module for processing, obtaining AFM processing result 9; AFM processing result 8 and AFM processing result 9 are input into the AFM module for processing, obtaining AFM processing result 10; In the fourth layer, the processing results of the first ICCM module 5 and the second ICCM module 5 are input to the AFM module for processing, and an AFM processing result 11 is obtained; the processing results of the second ICCM module 5 and the third ICCM module 5 are input to the AFM module for processing, and an AFM processing result 12 is obtained; the processing results of the third ICCM module 5 and the fourth ICCM module 5 are input to the AFM module for processing, and an AFM processing result 13 is obtained; the processing results of the fourth ICCM module 5 and the fifth ICCM module 5 are input to the AFM module for processing, and an AFM processing result 14 is obtained; the AFM processing results 11 and 12 are input to the AFM module. The FM module processes to obtain AFM processing result 15; the AFM processing result 12 and the AFM processing result 13 are input into the AFM module for processing to obtain AFM processing result 16; the AFM processing result 13 and the AFM processing result 14 are input into the AFM module for processing to obtain AFM processing result 17; the AFM processing result 15 and the AFM processing result 16 are input into the AFM module for processing to obtain AFM processing result 18; the AFM processing result 16 and the AFM processing result 17 are input into the AFM module for processing to obtain AFM processing result 19; the AFM processing result 18 and the AFM processing result 19 are input into the AFM module for processing to obtain AFM processing result 20; AFM processing results 1, 4, 10, and 20 are adjusted to the same number of channels through 1*1 convolution, and then feature union is performed through the cancatenate module. They are then processed by 1*1 convolution and sigmoid function in sequence to obtain the final output contour. The processing process in the AFM module is as follows: the two input results are added and fused, and the output result is obtained through 1*1 convolution.
2. The contour detection method for simulating complex cell characteristics according to claim 1, wherein: The first ICCM module 2, the second ICCM module 2, the first ICCM module 5, the second ICCM module 5, the third ICCM module 5, the fourth ICCM module 5, and the fifth ICCM module 5 have the same basic structure, and respectively include an SM1 module and an MSSM module; The input result is processed by the SM1 module and the MSSM module in sequence, and the MSSM module processing result is added and fused with the input result to obtain the output result; The first ICCM module 3, the second ICCM module 3, the third ICCM module 3, the first ICCM module 4, the second ICCM module 4, the third ICCM module 4, and the fourth ICCM module 4 have the same basic structure, which respectively include an SM2 module, an MSSM module, and a 1*1 convolution; after the input result is processed by the SM2 module and the MSSM module in sequence, the MSSM module processing result with the number of channels doubled is obtained, and the MSSM module processing result is added and fused with the input result after the number of channels is adjusted to the same by the 1*1 convolution to obtain the output result; The processing process in the SM1 module is as follows The input structure is divided into three paths. The first path is processed by 5*5 convolution and LeakyReLU function in sequence. The second path is processed by 1*1 convolution and ReLU function in sequence. The third path is processed by distributed spatial attention module. The results of the second and third paths are subtracted and then subtracted from the results of the first path to obtain the output result. During the three-path processing, the number of image channels remains unchanged. The processing process in the SM2 module is as follows The input result is divided into three paths. The first path is processed by 5*5 convolution and its channel number is doubled, and then processed by LeakyReLU function; the second path is processed by 1*1 convolution and its channel number is doubled, and then processed by ReLU function; the third path is processed by distributed spatial attention module and its channel number is doubled; the second and third path processing results are subtracted, and then subtracted from the first path processing result to obtain the output result; The processing process in the MSSM module is as follows: The input result is processed by shift operation. The shift operation process is as follows: the channel of the input result is divided into eight parts, and each part is moved in a different spatial direction. The specific moving directions are: upper left, upper, upper right, right, lower right, lower, lower left, and left; The result of the Shift operation is three-way. The results of moving to the upper left, upper right, lower left, and lower right are input to the first path 3*3 convolution processing, the results of moving to the left and right are input to the second path 1*3 convolution processing, and the results of moving up and down are input to the third path 3*1 convolution processing; the three-way processing results are spliced by the Concat function and then processed by 3*3 convolution to obtain the output result.
3. The contour detection method for simulating complex cell characteristics according to claim 2, wherein: The processing process in the distributed spatial attention module is as follows: The input result is divided into two paths. The first path is processed by 1*1 convolution, and the second path is processed by ReLU function, 1*1 convolution, 3*3 convolution, and Softmax function in sequence. The output result is obtained by multiplying the results of the two paths. When corresponding to the SM1 module, the number of channels of the input result is not changed during the first and second processing; when corresponding to the SM2 module, the 1*1 convolution processing of the first and second paths doubles the number of channels of the input result.
Citation Information
Patent Citations
Retinal encoder for machine vision
CN103890781A
Color image edge detection method
CN114022505A