A Spatial Location Relationship Detection Method Based on a Homogeneous Residual Network
By designing a uniform residual network to keep the number of channels unchanged, the problems of insufficient utilization of deep information and overfitting in the prior art are solved, and a higher precision spatial position relationship detection is achieved.
Patent Information
- Application Number
- CN202210998966.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-19
AI Technical Summary
The prior art fails to effectively utilize depth information in spatial position relationship detection, resulting in insufficient detection accuracy, and the change in the number of channels during downsampling of conventional residual networks leads to an increase in the risk of overfitting.
A uniform residual network is adopted to keep the number of channels unchanged during deep information processing, a uniform residual network is designed to extract depth information features, and spatial position relationship detection is performed in combination with RGB image features to avoid overfitting.
It improves the accuracy of spatial position relationship detection, reduces the risk of overfitting, and improves the detection effect.
Smart Images

Figure CN115359122B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a method for detecting spatial position relationships based on a homogeneous residual network. Background Art
[0002] Computer vision can not only detect objects in an image, but also needs to further understand the relationships between various objects in the image, which is usually referred to as scene graph generation. For a given picture, scene graph generation requires inferring the relationship between any two arbitrarily selected objects in the picture. For example, in "a person rides a bicycle", "rides" is the relationship between "person" and "bicycle". After extracting visual features and spatial position features from an RGB image and statistical features from a dataset, reference [1] proposed a scene graph generation algorithm based on conditional random fields, and reference [2] used an environment-related attention mechanism to improve the accuracy of the model.
[0003] Spatial position relationship detection is a type of task in scene graph generation, which mainly focuses on the spatial position relationship between two objects. For example, judging the front-back position relationship between the goal line and the football in a football game. On the other hand, the objective world is three-dimensional. Due to the perspective, it is sometimes difficult to accurately judge whether the football has crossed the goal line through a single RGB image. This means that a complete understanding of the spatial position relationship needs to take into account depth information as well. Both reference [1] and [2] only considered RGB images and did not use the depth information of objects. To overcome this shortcoming, reference [3] used the depth information of objects when detecting spatial position relationships. The experimental results also proved that using depth information helps to improve the detection accuracy. However, reference [3] adopted a simple method of taking the average value when extracting depth information features, and the extracted depth features are greatly affected by the object placement angle.
[0004] In computer vision, a backbone network is generally used to extract features from an image. Commonly used backbone networks include VGG [9] and Residual Network (ResNet)
[10]
[11] , etc. VGG contains multiple convolutional layers and pooling layers. The convolutional layer is used to obtain local features of the image, and the pooling layer is used for feature aggregation and image downsampling. The residual network adds a residual structure on the basis of VGG, thus solving the problem that deep neural networks are difficult to train. A common design feature of VGG and the residual network is that the number of channels of the network is doubled each time downsampling is performed, so that the computational amount of each layer of the backbone network is approximately the same. This design also increases the number of parameters of the neural network, thus increasing the risk of overfitting.
[0005] References:
[0006] [1] B. Dai, Y. Zhang and D. Lin, "Detecting Visual Relationships with Deep Relational Networks," 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3298 - 3308, doi:10.1109 / CVPR.2017.352
[0007] [2] Bohan Zhuang, Lingqiao Liu, Chunhua Shen, and Ian Reid. Towards context - aware interaction recognition for visual relationship detection. In Proceedings of the IEEE International Conference on Computer Vision, pages 589–598, 2017
[0008] [3] Ding X, Li Y, Pan Y, et al. Exploring depth information for spatial relation recognition[C] / / 2020 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR). IEEE, 2020:279 - 284
[0009] [4] K. Yang, O. Russakovsky and J. Deng, "SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition," 2019 IEEE / CVF International Conference on Computer Vision (ICCV), 2019, pp. 2051 - 2060, doi:10.1109 / ICCV.2019.00214.
[0010] [5]Ankit Goyal, Kaiyu Yang, Dawei Yang, Jia Deng, “Rel3D: A Minimally Contrastive Benchmark for Grounding Spatial Relations in 3D”, Neural Information Processing Systems (NeurIPS), 2020
[0011] [6]Yikang Li, Wanli Ouyang, Xiaogang Wang, and Xiao’ou Tang. Vip-cnn: Visual phrase guided convolutional neural network. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017
[0012] [7]Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, and Tat-Seng Chua. Visual translation embedding network for visual relation detection. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017
[0013] [8]Bohan Zhuang, Lingqiao Liu, Chunhua Shen, and Ian Reid. Towards context-aware interaction recognition for visual relationship detection. In Proceedings of the IEEE International Conference on Computer Vision, pages 589–598, 2017
[0014] [9]K.Simonyan and A.Zisserman. Very deep convolutional networks for large-scale image recognition. In The International Conference on Learning Representations(ICLR), 2015.
[0015]
[10] K.He, X.Zhang, S.Ren, and J.Sun. Deep residual learning for image recognition. In The IEEE Conference on Computer Vision and Pattern Recognition(CVPR), 2016.
[0016]
[11] K.He, X.Zhang, S.Ren, and J.Sun. Identity mappings in deep residual networks. In European Conference on Computer Vision(ECCV), 2016. Summary of the Invention
[0017] The purpose of the present invention is to provide a method for detecting spatial position relationships based on a homogeneous residual network to improve the detection accuracy of spatial position relationships.
[0018] The method for detecting spatial position relationships based on a homogeneous residual network in this patent is based on the following principle: it uses a residual network to extract depth information features from the depth information of an image to overcome the disadvantages brought by directly averaging the object depth. In addition, a homogeneous residual network is designed so that the number of channels of features remains unchanged during downsampling when processing depth information, thereby effectively improving the detection accuracy of spatial position relationships and avoiding overfitting.
[0019] To achieve the above purpose, the present invention provides a method for detecting spatial position relationships based on a homogeneous residual network, including:
[0020] S1: Use a homogeneous residual network to extract depth information features from the depth information of an image;
[0021] S2: Extract the first type of spatial position features from object labels and object bounding boxes, and extract the second type of spatial position features from RGB images;
[0022] S3: Feed all the depth information features, the first type of spatial position features, and the second type of spatial position features into the spatial position relationship classification network to detect the spatial position relationship;
[0023] In the step S1, the homogeneous residual network is designed by the following method:
[0024] S11: Select one of the basic building block and the bottleneck building block as the building block of the homogeneous residual network;
[0025] S12: Set the number of building blocks in the residual network;
[0026] S13: Insert a 7×7 convolutional layer and a 3×3 max pooling layer before the first building block, and insert an average pooling layer after the last building block; the first building block refers to the building block that first processes the input image during the process of processing the input image;
[0027] S14: Select the neural network layer for downsampling;
[0028] S15: Set the number of input channels and output channels of each building block. Among them, all building blocks have the same number of convolutional layers, and the number of input channels and output channels of the same convolutional layer in different building blocks does not change with downsampling.
[0029] Preferably, the building block of the homogeneous residual network adopts the basic building block, and the basic building block is composed of two convolutional layers with a size of 3×3, as well as the corresponding normalization layer and activation function layer. The number of input channels and output channels of the two convolutional layers is the same and does not change with downsampling.
[0030] Preferably, in each basic building block, the number of input channels of the two 3×3 convolutional layers is 64, and the number of output channels is 64.
[0031] Preferably, the constituent unit of the homogeneous residual network adopts a bottleneck constituent unit. The bottleneck constituent unit is composed of three convolutional layers and corresponding normalization layers and activation function layers. The size of the first convolutional layer is 1×1, the size of the second convolutional layer is 3×3, and the size of the third convolutional layer is 1×1; for all bottleneck constituent units, the number of output channels of the second convolutional layer is the same as the number of output channels of the first convolutional layer and does not change with downsampling; for other bottleneck constituent units except the first bottleneck constituent unit, the number of input channels of the first convolutional layer is 4 times the number of input channels of the second convolutional layer and does not change with downsampling; for all bottleneck constituent units, the number of input channels of the third convolutional layer is the same as the number of input channels of the second convolutional layer and does not change with downsampling, and the number of output channels of the third convolutional layer is 4 times the number of output channels of the second convolutional layer and does not change with downsampling.
[0032] Preferably, the number of channels of the first bottleneck constituent unit is: the input channel of the first convolutional layer is 64, and the output channel is 64; the input channel of the second convolutional layer is 64, and the output channel is 64; the input channel of the third convolutional layer is 64, and the output channel is 256; the number of channels of other bottleneck constituent units except the first bottleneck constituent unit is: the input channel of the first convolutional layer is 256, and the output channel is 64; the input channel of the second convolutional layer is 64, and the output channel is 64; the input channel of the third convolutional layer is 64, and the output channel is 256; where the first bottleneck constituent unit refers to the bottleneck constituent unit that first processes data in the process of processing the input image.
[0033] Preferably, in the step S14, the total number of key layers is a predetermined positive integer, and the ordinal number of the neural network layer for downsampling is selected according to the total number of key layers of the residual network;
[0034] If the total number of key layers is 18, the neural network layers for downsampling are: the 7×7 convolutional layer, the 3×3 max-pooling layer, and the first 3×3 convolutional layer in the 3rd, 5th, and 7th constituent units;
[0035] If the total number of key layers is 34, the neural network layers for downsampling are: the 7×7 convolutional layer, the 3×3 max-pooling layer, and the first 3×3 convolutional layer in the 4th, 8th, and 14th constituent units;
[0036] If the total number of key layers is 50, the neural network layers for downsampling are: the 7×7 convolutional layer, the 3×3 max-pooling layer, and the 3×3 convolutional layer in the 4th, 8th, and 14th constituent units;
[0037] If the total number of key layers is 101, the downsampling neural network layers are: a 7×7 convolutional layer, a 3×3 max pooling layer, and 3×3 convolutional layers in the 4th, 8th, and 31st building blocks;
[0038] If the total number of key layers is 152, the downsampling neural network layers are: a 7×7 convolutional layer, a 3×3 max pooling layer, and 3×3 convolutional layers in the 4th, 12th, and 48th building blocks.
[0039] The spatial position relationship detection method based on the homogeneous residual network proposed in this patent uses a residual network to extract depth information features from the depth information of an image to overcome the disadvantages brought by directly averaging the object depth. In addition, the present invention designs a homogeneous residual network, in which the number of channels of the network remains unchanged during downsampling when processing depth information, so as to effectively improve the detection accuracy of the spatial position relationship and avoid overfitting. It should be noted that the number of channels remains unchanged only during downsampling when processing depth information. When extracting features from RGB images, if a traditional network is used, there will still be downsampling, but without this feature. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of a spatial position relationship detection method based on a homogeneous residual network according to an embodiment of the present invention.
[0041] Figure 2 is a design flowchart of the homogeneous residual network of a spatial position relationship detection method based on a homogeneous residual network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following further describes the present invention with reference to specific embodiments. It should be understood that the following embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0043] As Figure 1 shown is a spatial position relationship detection method based on a homogeneous residual network according to an embodiment of the present invention, and the method includes the following steps:
[0044] Step S1: Use a homogeneous residual network to extract depth information features from the depth information of the image; wherein, the number of input and output channels of the homogeneous residual network remains unchanged during downsampling;
[0045] Among them, the building blocks of the homogeneous residual network include one of a basic building block and a bottleneck building block. Generally, when the total number of key layers of the homogeneous residual network is less than 50, the basic building block is used, and when the total number of key layers of the homogeneous residual network is greater than or equal to 50, the bottleneck building block is used.
[0046] The basic form building block consists of two 3×3 convolutional layers, along with corresponding normalization layers and activation function layers. The number of output channels of the two convolutional layers is the same.
[0047] In this embodiment, the basic form building block is as follows:
[0048]
[0049] Among them, 3×3 represents a 3×3 convolutional layer, 64 indicates that the number of output channels of the feature is 64, and the large square brackets mean that the input of the first layer of the basic form building block is also directly connected to the output of the second layer. Here, the number of output channels of the feature is discussed first, and the number of input channels of the feature will be further discussed later.
[0050] The bottleneck form building block consists of three convolutional layers, along with corresponding normalization layers and activation function layers. The first convolutional layer is a 1×1 convolutional layer, the second convolutional layer is a 3×3 convolutional layer, and the third convolutional layer is a 1×1 convolutional layer. The number of output channels of the second convolutional layer is the same as that of the first convolutional layer, and the number of output channels of the third convolutional layer is 4 times that of the first convolutional layer.
[0051] In this embodiment, the bottleneck form building block is as follows:
[0052]
[0053] Among them, 1×1 represents a 1×1 convolutional layer, and the large square brackets mean that the input of the first layer of the bottleneck form building block is also directly connected to the output of the third layer.
[0054] In the present invention, a data-driven method is adopted to extract features from the depth information of an image, that is, the depth information features of the image are extracted through a neural network.
[0055] Existing neural networks:
[0056] In the prior art, for feature extraction, classic backbone networks such as conventional VGG and ResNet can be used. Traditional backbone networks such as VGG and ResNet are designed for RGB images. One of their design principles is that every time downsampling is performed, the number of output channels is doubled.
[0057] Specifically, the conventional ResNet will increase the number of channels of the network when encountering downsampling. For example, for a conventional ResNet composed of basic form building blocks, the number of output channels of its basic form building block will change from 64 to 128, then to 256, and finally to 512.
[0058]
[0059] Among them, each matrix represents a building block, and each row of the matrix represents a convolutional layer. The first column of the matrix is the size of the convolutional layer (such as 3×3), and the second column is the number of output channels of the convolutional layer (such as 64, 128, 256, 512).
[0060] For a conventional residual network composed of bottleneck building blocks, the bottleneck building blocks also undergo similar changes when encountering downsampling:
[0061]
[0062] Among them, each matrix represents a building block, and each row of the matrix represents a convolutional layer. The first column of the matrix is the size of the convolutional layer (such as 3×3, 1×1), and the second column is the number of output channels of the convolutional layer (such as 64, 128, 256, 512).
[0063] The neural network for depth information feature extraction of the present invention:
[0064] However, in the present invention, for spatial position relationship detection, depth information is used to infer the relative position between two objects and does not require so many channels to express. This patent designs a homogeneous residual network to replace the classical backbone network. The input and output channel numbers of the homogeneous residual network remain unchanged during downsampling, thereby reducing the number of parameters and the risk of overfitting.
[0065] In the homogeneous residual network designed in this patent, regardless of whether downsampling is encountered, the number of output channels of each convolutional layer corresponding to the depth information features always remains unchanged. For the basic building block, there is
[0066]
[0067] Among them, each matrix represents a building block, and each row of the matrix represents a convolutional layer. The first column of the matrix is the size of the convolutional layer (such as 3×3, 1×1), and the second column is the number of output channels of the convolutional layer.
[0068] For the bottleneck building block, there is:
[0069]
[0070] Among them, the meaning of each matrix is similar to that above.
[0071] For the number of input channels of the neural network, the following method is adopted to determine. If the basic building block is adopted, in each basic building block, the number of input channels of the two 3×3 convolutional layers is 64.
[0072] If the bottleneck building block is adopted, the number of input channels of the three convolutional layers of the first bottleneck building block is 64. The number of input channels of the other bottleneck building blocks except the first one is as follows: the number of input channels of the first convolutional layer is 256; the number of input channels of the second convolutional layer is 64; the number of input channels of the third convolutional layer is 64.
[0073] According to the above discussion, in the homogeneous residual network, if the basic building block is adopted, its structure is as follows:
[0074]
[0075] Among them, each matrix represents a building block. Each row of the matrix represents a convolutional layer. The first column of the matrix is the number of input channels of the convolutional layer, the second column is the size of the convolutional layer, and the third column is the number of output channels of the convolutional layer. That is to say, 64, 3×3, 64 refers to a convolutional layer with an input channel number of 64 and an output channel number of 64 and a size of 3×3.
[0076] If the bottleneck building block is adopted, then:
[0077] The structure of the first bottleneck building block is
[0078]
[0079] The structure of the remaining bottleneck building blocks is:
[0080]
[0081] Among them, the meaning of each matrix is similar to that above. That is to say, 256, 1×1, 64 refers to a 1×1 convolutional layer with an input channel number of 256 and an output channel number of 64.
[0082] As Figure 2 shown, in the step S1, the homogeneous residual network is designed by the following method:
[0083] Step S11: Select one of the two building blocks, namely the basic building block and the bottleneck building block.
[0084] As described above, the basic building block is composed of two 3×3 convolutional layers and the corresponding normalization layer and activation function layer. The output channel numbers of the two convolutional layers are the same and do not change with downsampling.
[0085] The bottleneck building block consists of three convolutional layers, along with corresponding normalization layers and activation function layers. The size of the first convolutional layer is 1×1, the size of the second convolutional layer is 3×3, and the size of the third convolutional layer is 1×1. For all bottleneck building blocks, the number of output channels of the second convolutional layer is the same as that of the first convolutional layer and does not change with downsampling. For all bottleneck building blocks except the first one, the number of input channels of the first convolutional layer is 4 times that of the second convolutional layer and does not change with downsampling. For all bottleneck building blocks, the number of input channels of the third convolutional layer is the same as that of the second convolutional layer and does not change with downsampling, and the number of output channels of the third convolutional layer is 4 times that of the second convolutional layer and does not change with downsampling.
[0086] Step S12: Set the number of building blocks in the residual network according to the total number of key layers in the residual network; where the total number of key layers is a pre-specified positive integer.
[0087] If the basic building block is selected, each basic building block occupies 2 key layers. Therefore, the number of building blocks is:
[0088]
[0089] If the bottleneck building block is selected, each bottleneck building block occupies 3 key layers. The number of building blocks is:
[0090]
[0091] Step S13: Insert a 7×7 convolutional layer and a 3×3 max pooling layer before the first building block, and insert an average pooling layer after the last building block; the first building block refers to the building block that performs data processing earliest in the process of processing the input image.
[0092] Specifically, if the depth information of the image is in HHA format, then the number of input channels of the 7×7 convolutional layer is 3, and the number of output channels is 64. If the depth information of the image is not in HHA format, then the number of input channels of the 7×7 convolutional layer is 1, and the number of output channels is 64. In addition, an average pooling layer is inserted after the last building block, and the size of the output image is 1×1.
[0093] Step S14: Select the ordinal number of the neural network layer for downsampling according to the total number of key layers in the residual network.
[0094] Specifically, if the total number of key layers is 18, then the neural network layers for downsampling are: the 7×7 convolutional layer, the 3×3 max pooling layer, and the first 3×3 convolutional layer in the 3rd, 5th, and 7th building blocks.
[0095] If the total number of key layers is 34, the downsampling neural network layers are: a 7×7 convolutional layer, a 3×3 max pooling layer, and the first 3×3 convolutional layer in the 4th, 8th, and 14th building blocks.
[0096] If the total number of key layers is 50, the downsampling neural network layers are: a 7×7 convolutional layer, a 3×3 max pooling layer, and the 3×3 convolutional layers in the 4th, 8th, and 14th building blocks.
[0097] If the total number of key layers is 101, the downsampling neural network layers are: a 7×7 convolutional layer, a 3×3 max pooling layer, and the 3×3 convolutional layers in the 4th, 8th, and 31st building blocks.
[0098] If the total number of key layers is 152, the downsampling neural network layers are: a 7×7 convolutional layer, a 3×3 max pooling layer, and the 3×3 convolutional layers in the 4th, 12th, and 48th building blocks.
[0099] Step S15: Set the input channels and output channels of each building block. Among them, all building blocks have the same number of convolutional layers, and in the homogeneous residual network, the input and output channels remain unchanged during downsampling. That is to say, the output channels of the same convolutional layer in different building blocks are the same.
[0100] Among them, if the basic building block is adopted, in each basic building block, the input channels of both 3×3 convolutional layers are 64, and the output channels are both 64.
[0101] If the bottleneck building block is adopted, the number of channels of the first bottleneck building block is: the input channel of the first convolutional layer is 64, and the output channel is 64; the input channel of the second convolutional layer is 64, and the output channel is 64; the input channel of the third convolutional layer is 64, and the output channel is 256. The number of channels of the other bottleneck building blocks except the first one is: the input channel of the first convolutional layer is 256, and the output channel is 64; the input channel of the second convolutional layer is 64, and the output channel is 64; the input channel of the third convolutional layer is 64, and the output channel is 256.
[0102] Step S2: Extract the first type of spatial position features from the object label and the object bounding box, and extract the second type of spatial position features from the RGB image; the documents DRNet[1], Vip-CNN[6], VTransE[7], PPR-FCN[8] provide several possible specific implementation manners of step S2.
[0103] Step S3: Feed all the extracted depth information features, first - type spatial position features, and second - type spatial position features into the spatial position relationship classification network for detecting spatial position relationships.
[0104] Among them, the existing literature DRNet[1], Vip - CNN[6], VTransE[7], PPR - FCN[8] provides possible specific implementation manners of the spatial position relationship classification network. The result of the detection of spatial position relationships is the spatial position relationship. Which specific types the spatial position relationship includes is related to the dataset used. In the SpatialSense dataset used in the experiment, there are a total of 9 spatial relationships, specifically: above (with contact), above (without contact), below, front, back, left, right, adjacent, inside.
[0105] In the detection of spatial position relationships, in addition to extracting features from the depth information of the image, features can also be extracted from object bounding boxes, object labels, RGB images, etc. Finally, all the extracted features are fed into the spatial position relationship classification network for detecting spatial position relationships.
[0106] The comparison between the conventional residual network and the homogeneous residual network of the present invention is as follows:
[0107] The following Table 1 compares the conventional residual network ResNet - 18 and the homogeneous residual network ResNet - 18, thus illustrating the differences between the homogeneous residual network and the conventional residual network.
[0108] Among them, the normalization layer and the activation function layer are omitted in the process of making the table.
[0109] The stage in Table 1 refers to the down - sampling stage. Each time down - sampling is performed, a new stage is entered.
[0110] Table 1 Comparison between Conventional ResNet - 18 and Homogeneous ResNet - 18
[0111]
[0112] Tables 2 - 6 illustrate the detailed structures of different types of homogeneous residual networks. Among them, Tables 2 - 3 show the detailed structures of the homogeneous residual network with the basic building block as the building unit, and Tables 4 - 6 show the detailed structures of the homogeneous residual network with the bottleneck building block as the building unit.
[0113] Table 2 Detailed Structure of Homogeneous Residual Network ResNet - 18
[0114]
[0115] In Table 2, the value of N_input is as follows: If the depth information of the image adopts the HHA format, then N_input = 3; otherwise, N_input = 1.
[0116] Table 3 Detailed Structure of the Homogeneous Residual Network ResNet-34
[0117]
[0118] Table 4 Detailed Structure of the Homogeneous Residual Network ResNet-50
[0119]
[0120] Table 5 Detailed Structure of the Homogeneous Residual Network ResNet-101
[0121]
[0122] Table 6 Detailed Structure of the Homogeneous Residual Network ResNet-152
[0123]
[0124] Experimental Results:
[0125] The following Table 7 applies the method proposed in this patent to a specific spatial position relationship detection task to verify its effectiveness. The test scenario is the SpatialSense dataset [4], and only the NYU Depth part is used because the Flickr part of SpatialSense does not provide depth information. The performance metric is the average recognition accuracy of 9 spatial position relationships.
[0126] Table 7 Comparison of the Average Recognition Accuracy between the Method of the Present Invention and Existing Algorithms in the Literature
[0127]
[0128] It can be seen that the performance of this patent has increased by 3.94%.
[0129] The above-mentioned are only the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. The above embodiments of the present invention can also be variously changed. All simple, equivalent changes and modifications made according to the claims and the content of the specification of the present invention application fall within the scope of the claims of the present invention patent. What is not described in detail in the present invention is conventional technical content.
Claims
1. A method for detecting spatial position relationships based on a homogeneous residual network, characterized in that Including: Step S1: Extract depth information features from the depth information of the image using a homogeneous residual network; Step S2: Extract the first type of spatial position features from object labels and object bounding boxes, and extract the second type of spatial position features from the RGB image; Step S3: Feed all depth information features, the first type of spatial position features, and the second type of spatial position features into a spatial position relationship classification network to detect spatial position relationships; In step S1, the homogeneous residual network is designed by the following method: Step S11: Select one of the basic building block and the bottleneck building block as the building block of the homogeneous residual network; Step S12: Set the number of building blocks in the residual network; Step S13: Insert a 7×7 convolutional layer and a 3×3 max pooling layer before the first building block, and insert an average pooling layer after the last building block; the first building block refers to the building block that first processes the input image during the process of processing the input image; Step S14: Select the neural network layer for downsampling; Step S15: Set the number of input channels and output channels of each building block, where all building blocks have the same number of convolutional layers, and the number of input channels and output channels of the same convolutional layer of different building blocks does not change with downsampling; The building block of the homogeneous residual network adopts the bottleneck building block. The bottleneck building block consists of three convolutional layers and corresponding normalization layers and activation function layers. The size of the first convolutional layer is 1×1, the size of the second convolutional layer is 3×3, and the size of the third convolutional layer is 1×1; for all bottleneck building blocks, the number of output channels of the second convolutional layer is the same as the number of output channels of the first convolutional layer and does not change with downsampling; for other bottleneck building blocks except the first bottleneck building block, the number of input channels of the first convolutional layer is 4 times the number of input channels of the second convolutional layer and does not change with downsampling; for all bottleneck building blocks, the number of input channels of the third convolutional layer is the same as the number of input channels of the second convolutional layer and does not change with downsampling, and the number of output channels of the third convolutional layer is 4 times the number of output channels of the second convolutional layer and does not change with downsampling.
2. The spatial position relationship detection method based on a homogeneous residual network according to claim 1, wherein The building block of the homogeneous residual network adopts the basic building block. The basic building block consists of two convolutional layers with a size of 3×3 and corresponding normalization layers and activation function layers. The number of input channels and output channels of the two convolutional layers is the same and does not change with downsampling.
3. The spatial position relationship detection method based on the homogeneous residual network according to claim 2, wherein In each basic building block, the number of input channels of the two 3×3 convolutional layers is 64, and the number of output channels is 64.
4. The method for detecting spatial position relationship based on homogeneous residual network according to claim 1, wherein The number of channels of the first bottleneck building block is as follows: the number of input channels of the first convolutional layer is 64, and the number of output channels is 64; the number of input channels of the second convolutional layer is 64, and the number of output channels is 64; the number of input channels of the third convolutional layer is 64, and the number of output channels is 256; the number of channels of the other bottleneck building blocks except the first bottleneck building block is: the number of input channels of the first convolutional layer is 256, and the number of output channels is 64; the number of input channels of the second convolutional layer is 64, and the number of output channels is 64; the number of input channels of the third convolutional layer is 64, and the number of output channels is 256; where the first bottleneck building block refers to the bottleneck building block that first processes the input image during the process of processing the input image.
5. The spatial position relationship detection method based on the homogeneous residual network according to claim 1, characterized in that, In the step S14, the total number of key layers is a pre-specified positive integer, and the ordinal number of the neural network layer for downsampling is selected according to the total number of key layers of the residual network; If the total number of key layers is 18, the neural network layers for downsampling are: a 7×7 convolutional layer, a 3×3 max-pooling layer, and the first 3×3 convolutional layer in the 3rd, 5th, and 7th building blocks; If the total number of key layers is 34, the neural network layers for downsampling are: a 7×7 convolutional layer, a 3×3 max-pooling layer, and the first 3×3 convolutional layer in the 4th, 8th, and 14th building blocks; If the total number of key layers is 50, the neural network layers for downsampling are: a 7×7 convolutional layer, a 3×3 max-pooling layer, and the 3×3 convolutional layers in the 4th, 8th, and 14th building blocks; If the total number of key layers is 101, the neural network layers for downsampling are: a 7×7 convolutional layer, a 3×3 max-pooling layer, and the 3×3 convolutional layers in the 4th, 8th, and 31st building blocks; If the total number of key layers is 152, the neural network layers for downsampling are: a 7×7 convolutional layer, a 3×3 max-pooling layer, and the 3×3 convolutional layers in the 4th, 12th, and 48th building blocks.
Citation Information
Patent Citations
Lung texture recognition method based on deep neural network extraction appearance and geometric features
CN108428229A
Image super-resolution reconstruction method based on cascade residual convolutional neural network
CN110276721A