A Multi-Granularity Human-Machine Co-Integration Environment Perception Method Based on Deep Learning
Through a multi-grained human-computer inclusive environment perception algorithm based on deep learning, multi-grained segmentation is used using RGB images and depth images, which solves the problem that a single particle size in the existing technology cannot meet the environment perception of a collaborative robot, and realizes fine environmental perception and motion control of the robot under different tasks and distances.
Patent Information
- Application Number
- CN202211008001.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-08-22
AI Technical Summary
The existing scene semantic segmentation methods mainly stay at a single granularity and cannot meet the environmental perception needs of collaborative robots under different distances and fineness requirements, resulting in the robot being unable to perform fine motion control and collaborative actions during human-machine collaboration assembly.
Using a multi-grained human-computer inclusive environment perception algorithm based on deep learning, multi-grained segmentation is performed by acquiring RGB images and depth images, encoded networks, pyramid pooling modules and decoding networks, providing scene segmentation of regions, entities and partial levels to realize environment perception switching of different granularities.
It provides collaborative robots with more complete environmental perception capabilities, and can adaptively switch the perceptual segmentation results of different granularities according to the environment and tasks, improving the accuracy and efficiency of collaborative behavior decision-making and motion planning.
Smart Images

Figure CN115482532B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-machine collaborative intelligent manufacturing assembly, and in particular to a multi-granularity human-machine co-integration environment perception method based on deep learning. Background Art
[0002] Industrial robots are an essential part of intelligent manufacturing systems. However, the deployment cost of traditional industrial robots is high, and the entire production line needs to be transformed and adapted around them, resulting in insufficient flexibility of the production line. In addition, small and medium-sized enterprises are limited by funds and scale and cannot deploy traditional industrial robots on a large scale. In this context, collaborative robots and the human-machine co-integration manufacturing mode have gradually received more and more attention. Humans are responsible for work links with high requirements for flexibility, touch, and flexibility, while collaborative robots use their advantages of speed and precision to be responsible for repetitive and programmed work links. This can not only meet the needs of production line flexibility, flexibility, and large-scale personalization but also improve production efficiency and reduce the labor burden of workers.
[0003] This close-range human-machine co-integration collaboration mode requires the robot to be able to perceive its human-machine co-integration environment in real time and accurately, so as to adaptively adjust its working posture and operation tasks according to the actions of the collaborative personnel and environmental changes. Early environmental perception technologies mainly relied only on depth cameras or tactile sensors to judge the distance between the robot and surrounding obstacles. In recent years, with the development of artificial intelligence technology, deep neural networks have been more widely used to perform semantic segmentation on scenes, that is, to divide an image into different category regions according to the semantic information of the visual image.
[0004] However, existing scene semantic segmentation methods mainly still stay in the single-granularity semantic segmentation mode (for example, regardless of the size of the scene, the human body is always segmented as a whole). They do not consider that collaborative robots may have different perception distances and tasks with different fineness levels during operation. For example, in the case of medium distance, the robot may only need to identify the worker's body as a whole to avoid collisions. However, in the case of close-range human-machine collaborative assembly, the robot needs to more finely segment and distinguish different parts of the human body such as the hand, arm, and body, so as to perform more precise motion control and collaborative actions. This variability of specific scenarios and tasks requires the scene perception algorithm of the collaborative robot to have semantic understanding capabilities with multiple different granularities. Summary of the Invention
[0005] The present invention provides a multi-granularity human-robot coexistence environment perception algorithm model based on deep learning. By simultaneously performing scene segmentation at three granularity levels from coarse to fine, namely region, entity, and part, it provides a more perfect environment perception ability for collaborative robots, enabling them to adaptively switch different granularity environment perception segmentation results according to different environments and tasks, so as to better make subsequent collaborative behavior decisions and motion planning.
[0006] To achieve the above object, the present invention is implemented through the following technical solutions:
[0007] A multi-granularity human-robot coexistence environment perception method based on deep learning, wherein the method includes the steps of:
[0008] Obtain the RGB image and depth image of the human-robot coexistence scene; wherein, the RGB image and the depth image are images taken of the same human-robot coexistence scene;
[0009] Input the RGB image and the depth image into the encoding network to obtain an encoded image;
[0010] Input the encoded image into the pyramid pooling module to obtain a pooled image;
[0011] Input the pooled image into the decoding network to obtain a decoded image;
[0012] Input the decoded image into the multi-granularity segmentation output module to obtain scene segmentation images at different granularity levels;
[0013] Wherein, the granularity levels include region level, entity level, and part level of the entity;
[0014] The decoding network includes: a first upsampling module, a second upsampling module, and a third upsampling module; the step of inputting the pooled image into the decoding network to obtain a decoded image includes:
[0015] Input the pooled image into the first upsampling module to obtain first refined segmentation images at different granularity levels and a first upsampled image;
[0016] Input the first upsampled image and the first refined segmentation image into the second upsampling module to obtain second refined segmentation images at different granularity levels and a second upsampled image;
[0017] Input the second upsampled image and the second refined segmentation image into the third upsampling module to obtain third refined segmentation images at different granularity levels and a third upsampled image, and use the third refined segmentation image and the third upsampled image as the decoded image.
[0018] The described multi-granularity human-machine coexistence environment perception method based on deep learning, wherein the encoding network includes: a first downsampling module, a second downsampling module, a third downsampling module, a fourth downsampling module, a first fusion module, a second fusion module, a third fusion module, and a fourth fusion module;
[0019] The step of inputting the RGB image and the depth image into the encoding network to obtain an encoded image includes:
[0020] Input the depth image and the RGB image into the first downsampling module respectively to obtain a first downsampled depth image and a first downsampled RGB image;
[0021] Input the first downsampled depth image and the first downsampled RGB image into the first fusion module to obtain a first fused image;
[0022] Input the first downsampled depth image and the first fused image into the second downsampling module respectively to obtain a second downsampled depth image and a second downsampled RGB image;
[0023] Input the second downsampled depth image and the second downsampled RGB image into the second fusion module to obtain a second fused image;
[0024] Input the second downsampled depth image and the second fused image into the third downsampling module respectively to obtain a third downsampled depth image and a third downsampled RGB image;
[0025] Input the third downsampled depth image and the third downsampled RGB image into the third fusion module to obtain a third fused image;
[0026] Input the third downsampled depth image and the third fused image into the fourth downsampling module respectively to obtain a fourth downsampled depth image and a fourth downsampled RGB image;
[0027] Input the fourth downsampled depth image and the fourth downsampled RGB image into the fourth fusion module to obtain a fourth fused image, and use the fourth fused image as the encoded image.
[0028] The described multi-granularity human-machine coexistence environment perception method based on deep learning, wherein the first fusion module is jump-connected to the third upsampling module; the second fusion module is jump-connected to the second upsampling module; the third fusion module is jump-connected to the first upsampling module;
[0029] The third upsampling module includes: a first convolutional module, a first upsampled depth convolutional module, and a refinement segmentation module with different granularity levels;
[0030] Inputting the second upsampled image and the second refined segmentation image into the third upsampling module to obtain third refined segmentation images and a third upsampled image with different granularity levels, includes:
[0031] Inputting the second upsampled image into the first convolutional module to obtain a first convolutional image;
[0032] Inputting the first convolutional image into the first upsampling depth convolutional module to obtain a first depth convolutional image;
[0033] Concatenating the first depth convolutional image and the first fusion image to obtain a third upsampled image;
[0034] Inputting the first convolutional image and the second refined segmentation image of the corresponding granularity level into the refined segmentation module of the corresponding granularity level to obtain a third refined segmentation image of the corresponding granularity level.
[0035] The method for multi-granularity human-machine co-integration environment perception based on deep learning, wherein, the refined segmentation module includes: a first upsampling convolutional module and a first convolutional layer; the step of inputting the first convolutional image and the second refined segmentation image of the corresponding granularity level into the refined segmentation module of the corresponding granularity level to obtain a third refined segmentation image of the corresponding granularity level, includes:
[0036] Inputting the second refined segmentation image of the granularity level into the first upsampling convolutional module to obtain a first upsampling convolutional image;
[0037] Concatenating the first convolutional image and the first upsampling convolutional image and then inputting them into the first convolutional layer to obtain a third refined segmentation image of the corresponding granularity level.
[0038] The method for multi-granularity human-machine co-integration environment perception based on deep learning, wherein, the first fusion module includes: a first pooling convolutional module and a second pooling convolutional module;
[0039] The step of inputting the first downsampled depth image and the first downsampled RGB image into the first fusion module to obtain a first fusion image, includes:
[0040] Inputting the first downsampled depth image into the first pooling convolutional module to obtain a first pooling convolutional image;
[0041] Inputting the first downsampled RGB image into the second pooling convolutional module to obtain a second pooling convolutional image;
[0042] Concatenating the first pooling convolutional image and the second pooling convolutional image to obtain a first fusion image.
[0043] The described multi-granularity human-machine coexistence environment perception method based on deep learning, wherein the multi-granularity segmentation output module includes: segmentation output modules of different granularity levels, and each granularity level's segmentation output module includes: a second upsampling convolutional module, a second convolutional layer, a second upsampling depth convolutional module, and a third upsampling depth convolutional module;
[0044] Inputting the decoded image into the multi-granularity segmentation output module to obtain scene segmentation images of different granularity levels includes:
[0045] Inputting the third upsampled image into the second convolutional layer to obtain a second convolutional image;
[0046] Inputting the third refined segmentation image of the granularity level into the second upsampling convolutional module in the segmentation output module of the corresponding granularity level to obtain a second upsampling convolutional image;
[0047] Inputting the concatenated second convolutional image and the first upsampling convolutional image into the second upsampling depth convolutional module to obtain a second depth convolutional image;
[0048] Inputting the second depth convolutional image into the third upsampling depth convolutional module to obtain a scene segmentation image of the corresponding granularity level.
[0049] The described multi-granularity human-machine coexistence environment perception method based on deep learning, wherein the pyramid pooling module includes: a pooling layer, a fourth convolutional layer, and an upsampling layer; inputting the encoded image into the pyramid pooling module to obtain a pooled image includes:
[0050] Inputting different regions of the encoded image into the pooling layer to obtain region-pooled images of different sizes;
[0051] Inputting the region-pooled images of different sizes into the fourth convolutional layer respectively to obtain region-convolutional images of different sizes;
[0052] Inputting the region-convolutional images of different sizes into the upsampling layer respectively to obtain corresponding region-upsampled images; the sizes of all region-upsampled images are the same as the size of the encoded image;
[0053] Concatenating the encoded image and all the region-upsampled images to obtain the pooled image.
[0054] The described multi-granularity human-machine coexistence environment perception method based on deep learning, wherein the decoding network is trained based on a weighted cross-entropy loss function;
[0055] The weighted cross-entropy loss function is:
[0056]
[0057] Among them, represents the category of the entity, represents the weighted cross-entropy loss function, w i represents the weight of the category of the i-th entity, y i represents the label image of the category of the i-th entity, p i represents the scene segmentation image of different granularity levels of the category of the i-th entity.
[0058] A computer device includes a memory and a processor. The memory stores a computer program. Among them, when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0059] A computer-readable storage medium stores a computer program. Among them, when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0060] Beneficial effects:
[0061] After obtaining the RGB image and the depth image, input the RGB image and the depth image into the encoding network, extract and fuse the features in the RGB image and the depth image to obtain an encoded image. After obtaining the encoded image, input the encoded image into the pyramid pooling module for global-local feature fusion extraction to obtain a pooled image. After obtaining the pooled image, input the pooled image into the decoded image to obtain a decoded image. After obtaining the decoded image, input it into the multi-granularity segmentation output module to obtain the scene segmentation images of different granularity levels, providing a more perfect environment perception ability for the collaborative robot, enabling it to adaptively switch the environmental perception segmentation results of different granularities according to the different environments and tasks, and thus better making subsequent collaborative behavior decisions and motion planning. Description of the Drawings
[0062] Figure 1 It is a schematic diagram of the multi-granularity scene segmentation criterion for the human-robot co-fusion assembly environment of the present invention;
[0063] Figure 2 It is a first structural schematic diagram of the encoding-decoding model of the present invention;
[0064] Figure 3 It is a second structural schematic diagram of the encoding-decoding model of the present invention;
[0065] Figure 4 It is a schematic diagram of the downsampling module of the present invention;
[0066] Figure 5 It is a schematic diagram of the fusion module of the present invention;
[0067] Figure 6Schematic diagram of the upsampling module in the present invention;
[0068] Figure 7 Schematic diagram of the multi-granularity segmentation output module in the present invention;
[0069] Figure 8 Schematic diagram of the pyramid pooling module in the present invention;
[0070] Figure 9 Schematic diagram for comparison of RGB images, depth images, scene segmentation images at different granularity levels, and label images at different granularity levels in the present invention;
[0071] Figure 10 Overall flowchart of the multi-granularity human-machine coexistence environment perception method based on deep learning in the present invention. Detailed implementation manners
[0072] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various implementation manners of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0073] Please refer to Figures 1 - 10 simultaneously. The present invention provides some embodiments of a multi-granularity human-machine coexistence environment perception method based on deep learning.
[0074] As Figure 2 and Figure 3 shown, the present invention provides an encoder-decoder model, and the encoder-decoder model includes: an encoding network, a pyramid pooling module, a decoding network, and a multi-granularity segmentation output module.
[0075] The encoding network is used for downsampling an image, the pyramid pooling module is used for global-local feature fusion, the decoding network is used for upsampling an image, and the multi-granularity segmentation output module is used for outputting scene segmentation images at different granularity levels. The granularity level refers to the size level of the segmentation image in the image. The encoder-decoder model of the present application is applied to the human-machine coexistence scenario, and the human-machine coexistence scenario refers to the scenario of collaborative robots and human-machine coexistence in manufacturing. That is to say, in the human-machine coexistence scenario, there are robots, workers, manufacturing tools, manufacturing raw materials and products, divided areas, etc. The sizes of the various targets in the human-machine coexistence scenario are large and small. In tasks with different fineness, the robot needs to perform image segmentation with different fineness. Therefore, the various targets in the human-machine coexistence scenario are divided according to the granularity level, and the number of levels of the specific granularity level can be set as needed, such as Figure 1As shown, in the present invention, the granularity levels are divided into three levels, namely the area level, the entity level, and the part level of the entity. The area level refers to the level of relatively large sizes in the image that are connected to form areas. For example, the area level can be specifically divided into human-machine collaboration areas, warehousing areas, office areas, free areas, inaccessible areas, etc. The entity level refers to the level of relatively small targets in the image. For example, the entity level can be specifically divided into humans, robots, automated guided vehicles, tools, batteries, computers, desks, chairs, etc. The part level of the entity refers to the levels of the various parts of the entity in the image. For example, the part level of a human can be divided into arms, palms, heads, feet, trunks, etc. The part level of a robot can be divided into a base, a main fixture, etc. The part level of an automated guided vehicle can be divided into a base, wheels, a load platform. The part level of a tool can be divided into a handle, a front end, etc. The part level of a battery can be divided into a casing, connection wires, short battery packs, long battery packs, etc. The part level of a computer can be divided into a chassis, a keyboard, a mouse, a monitor, etc.
[0076] The encoding network includes: a first downsampling module, a second downsampling module, a third downsampling module, a fourth downsampling module, a first fusion module, a second fusion module, a third fusion module, and a fourth fusion module. The encoding network includes four downsampling modules and four fusion modules. The downsampling modules are used to downsample the image, and the fusion modules are used to fuse the depth image and the RGB image. The downsampling module includes: one depth convolutional layer and two convolutional layers. The fusion module includes: two pooling convolutional modules, and the pooling convolutional module includes: one pooling layer and two convolutional layers.
[0077] The pyramid pooling module includes: one pooling layer, one convolutional layer, and one upsampling layer.
[0078] The decoding network includes: a first upsampling module, a second upsampling module, and a third upsampling module. The decoding network includes three upsampling modules, and the image is upsampled through the upsampling modules. The upsampling module includes: one convolutional module, one upsampling depth convolutional module, and three refinement segmentation modules; the convolutional module includes: five convolutional layers, the upsampling depth convolutional module includes: one upsampling layer and one depth convolutional layer, the refinement segmentation module includes: one upsampling convolutional module and one convolutional layer, and the upsampling convolutional module includes: one upsampling layer and two convolutional layers.
[0079] The multi-granularity segmentation output module includes: segmentation output modules of different granularity levels. Since there are 3 granularity levels, there are also 3 segmentation output modules. The segmentation output module includes: 1 upsampling convolution module, 1 convolution layer, and 2 upsampling depth convolution modules. The upsampling convolution module includes: 1 upsampling layer and 2 convolution layers. The upsampling depth convolution module includes: 1 upsampling layer and 1 depth convolution layer.
[0080] As Figures 1 - 3 and Figure 10 shown, a multi-granularity human-machine co-integration environment perception method based on deep learning of the present invention includes the following steps:
[0081] Step S100, obtaining an RGB image and a depth image of a human-machine co-integration scene; wherein, the RGB image and the depth image are images obtained by photographing the same human-machine co-integration scene.
[0082] Step S200, inputting the RGB image and the depth image into an encoding network to obtain an encoded image.
[0083] Step S300, inputting the encoded image into a pyramid pooling module to obtain a pooled image.
[0084] Step S400, inputting the pooled image into a decoding network to obtain a decoded image.
[0085] Step S500, inputting the decoded image into a multi-granularity segmentation output module to obtain scene segmentation images of different granularity levels.
[0086] Photograph a human-machine co-integration scene to obtain an RGB image and a depth image. The RGB image and the depth image can be obtained by using an RGB-D camera, or can be respectively obtained by an RGB camera and a depth camera. The RGB image and the depth image are images obtained by photographing the same human-machine co-integration scene. That is to say, the targets in the RGB image and the depth image are the same targets, and the positions of the targets are at the same positions.
[0087] After obtaining the RGB image and the depth image, input the RGB image and the depth image into the encoding network to extract and fuse the features in the RGB image and the depth image to obtain an encoded image. After obtaining the encoded image, input the encoded image into the pyramid pooling module for global-local feature fusion extraction to obtain a pooled image. After obtaining the pooled image, input the pooled image into the decoding image to obtain a decoded image. After obtaining the decoded image, input it into the multi-granularity segmentation output module to obtain scene segmentation images of different granularity levels, providing a more perfect environment perception ability for the collaborative robot, enabling it to adaptively switch the environmental perception segmentation results of different granularities according to different environments and tasks, so as to better make subsequent collaborative behavior decisions and motion planning.
[0088] In one implementation, the decoding network mainly consists of three consecutive upsampling modules, gradually decoding and amplifying the feature map. The decoding network includes: a first upsampling module, a second upsampling module, and a third upsampling module; step S400 includes:
[0089] Step S410: Input the pooled image into the first upsampling module to obtain a first refined segmentation image and a first upsampled image with different granularity levels.
[0090] Step S420: Input the first upsampled image and the first refined segmentation image into the second upsampling module to obtain a second refined segmentation image and a second upsampled image with different granularity levels.
[0091] Step S430: Input the second upsampled image and the second refined segmentation image into the third upsampling module to obtain a third refined segmentation image and a third upsampled image with different granularity levels, and use the third refined segmentation image and the third upsampled image as the decoded image.
[0092] Specifically, after obtaining the pooled image, input the pooled image into the first upsampling module to obtain a first upsampled image and a first refined segmentation image with different granularity levels. For example, three granularity levels of the first refined segmentation image can be obtained, specifically, the first refined segmentation image at the region level, the first refined segmentation image at the entity level, and the first refined segmentation image at the part level.
[0093] Then input the first upsampled image and the first refined segmentation image into the second upsampling module to obtain a second refined segmentation image and a second upsampled image with different granularity levels. It can be understood that if the first refined segmentation image at the region level is used as the input of the second upsampling module, the second refined segmentation image at the region level is obtained. If the first refined segmentation image at the entity level is used as the input of the second upsampling module, the second refined segmentation image at the entity level is obtained.
[0094] Then input the second upsampled image and the second refined segmentation image into the third upsampling module to obtain a third refined segmentation image and a third upsampled image with different granularity levels, and the third refined segmentation image and the third upsampled image can be used as the decoded image.
[0095] In one implementation, the encoding network includes: a first downsampling module, a second downsampling module, a third downsampling module, a fourth downsampling module, a first fusion module, a second fusion module, a third fusion module, and a fourth fusion module; step S200 includes:
[0096] Step S210: Input the depth image and the RGB image into the first downsampling module respectively to obtain a first downsampled depth image and a first downsampled RGB image.
[0097] Step S220: Input the first downsampled depth image and the first downsampled RGB image into the first fusion module to obtain a first fused image.
[0098] Step S230: Input the first downsampled depth image and the first fused image into the second downsampling module respectively to obtain a second downsampled depth image and a second downsampled RGB image.
[0099] Step S240: Input the second downsampled depth image and the second downsampled RGB image into the second fusion module to obtain a second fused image.
[0100] Step S250: Input the second downsampled depth image and the second fused image into the third downsampling module respectively to obtain a third downsampled depth image and a third downsampled RGB image.
[0101] Step S260: Input the third downsampled depth image and the third downsampled RGB image into the third fusion module to obtain a third fused image.
[0102] Step S270: Input the third downsampled depth image and the third fused image into the fourth downsampling module respectively to obtain a fourth downsampled depth image and a fourth downsampled RGB image.
[0103] Step S280: Input the fourth downsampled depth image and the fourth downsampled RGB image into the fourth fusion module to obtain a fourth fused image, and use the fourth fused image as the encoded image.
[0104] Specifically, the first downsampling module, the second downsampling module, the third downsampling module, and the fourth downsampling module adopt downsampling modules with the same structure. As Figure 4 shown, this downsampling module includes: a 7×7 depth convolution layer, a 1×1 convolution layer, and a 1×1 convolution layer connected in sequence.
[0105] The purpose of the encoder network is to extract RGB and depth features and aggregate them at different stages to better utilize the complementary information in the RGB and depth maps. Both the encoding and decoding processes are carried out in a step-by-step manner. During the encoding process, 4 downsampling modules are used. In a step-by-step fusion manner, the RGB image and the depth image are fused, and specifically 4 fusion modules are used for fusion. At each stage of the backbone network of the encoder network, the depth features are fused into the RGB branch through the fusion module.
[0106] First, use the first downsampling module to downsample the depth image and the RGB image respectively, and obtain the first downsampled depth image and the first downsampled RGB image. Then, use the first fusion module to fuse the first downsampled depth image and the first downsampled RGB image to obtain the first fused image. Next, perform the next step of downsampling. After 4 times of downsampling, the fourth fused image is obtained, and the fourth fused image is used as the encoded image.
[0107] To accelerate the training, skip connections are adopted between the fusion module and the upsampling module. Specifically, the first fusion module is connected to the third upsampling module by a skip connection; the second fusion module is connected to the second upsampling module by a skip connection; the third fusion module is connected to the first upsampling module by a skip connection. The first upsampling module includes: a first convolution module, a first upsampled depth convolution module, and refinement segmentation modules with different granularity levels. In one implementation, step S410 includes:
[0108] Step S411: Input the pooled image into the first convolution module of the first upsampling module to obtain the first convolution image.
[0109] Step S412: Input the first convolution image into the first upsampled depth convolution module of the first upsampling module to obtain the first depth convolution image.
[0110] Step S413: Concatenate the first depth convolution image and the third fused image to obtain the first upsampled image.
[0111] Step S414: Input the first convolution image into the refinement segmentation module corresponding to the granularity level to obtain the first refined segmentation image corresponding to the granularity level.
[0112] It should be noted that due to the small size of the pooled image, the refinement segmentation modules with different granularity levels in the first upsampling module do not receive the input of the refined segmentation image. Only the first convolution image is used as the input of the refinement segmentation module corresponding to the granularity level to obtain the first refined segmentation image corresponding to the granularity level. In one implementation, the refinement segmentation module includes: a first upsampled convolution module and a first convolution layer; step S414 includes:
[0113] Step S4141: Input the first convolution image into the first upsampled convolution module of the refinement segmentation module corresponding to the granularity level to obtain the first upsampled convolution image.
[0114] Step S4142: Concatenate the first convolution image and the first upsampled convolution image and then input them into the first convolution layer to obtain the first refined segmentation image corresponding to the granularity level.
[0115] In one implementation, the second upsampling module includes: a first convolutional module, a first upsampling depth convolutional module, and refinement segmentation modules of different granularity levels; step S420 includes:
[0116] Step S421: Input the first upsampled image into the first convolutional module of the second upsampling module to obtain a first convolutional image.
[0117] Step S422: Input the first convolutional image into the first upsampling depth convolutional module of the second upsampling module to obtain a first depth convolutional image.
[0118] Step S423: Concatenate the first depth convolutional image and the second fused image to obtain a second upsampled image.
[0119] Step S424: Input the first convolutional image and the second refinement segmentation image of the granularity level into the refinement segmentation module of the corresponding granularity level to obtain the second refinement segmentation image of the corresponding granularity level.
[0120] In one implementation, the refinement segmentation module includes: a first upsampling convolutional module and a first convolutional layer; step S424 includes:
[0121] Step S4241: Input the second refinement segmentation image of the granularity level into the first upsampling convolutional module of the refinement segmentation module of the corresponding granularity level to obtain a first upsampling convolutional image.
[0122] Step S4242: Concatenate the first convolutional image and the first upsampling convolutional image and then input them into the first convolutional layer to obtain the first refinement segmentation image of the corresponding granularity level.
[0123] In one implementation, the third upsampling module includes: a first convolutional module, a first upsampling depth convolutional module, and refinement segmentation modules of different granularity levels; step S430 includes:
[0124] Step S431: Input the second upsampled image into the first convolutional module of the third upsampling module to obtain a first convolutional image.
[0125] Step S432: Input the first convolutional image into the first upsampling depth convolutional module of the third upsampling module to obtain a first depth convolutional image.
[0126] Step S433: Concatenate the first depth convolutional image and the first fused image to obtain a third upsampled image.
[0127] Step S434: Input the first convolutional image and the second refined segmentation image of the granularity level into the refined segmentation module corresponding to the granularity level to obtain the third refined segmentation image corresponding to the granularity level.
[0128] In one implementation, the refined segmentation module includes: a first upsampling convolutional module and a first convolutional layer; Step S434 includes:
[0129] Step S4341: Input the second refined segmentation image of the granularity level into the first upsampling convolutional module to obtain a first upsampling convolutional image.
[0130] Step S4342: Concatenate the first convolutional image and the first upsampling convolutional image and then input them into the first convolutional layer to obtain the third refined segmentation image corresponding to the granularity level.
[0131] Specifically, as Figure 6 shown, the first convolutional module includes: a 3×3 convolutional layer, a 3×1 convolutional layer, a 1×3 convolutional layer, a 3×1 convolutional layer, and a 1×3 convolutional layer connected in sequence. The first upsampling depth convolutional module includes: an upsampling layer and a 3×3 depth convolutional layer connected in sequence. The first upsampling convolutional module includes: 1 upsampling layer and 2 3×3 convolutional layers. The first convolutional layer uses a 1×1 convolutional layer.
[0132] In one implementation, the first fusion module includes: a first pooling convolutional module and a second pooling convolutional module; Step S220 includes:
[0133] Step S221: Input the first downsampled depth image into the first pooling convolutional module of the first fusion module to obtain a first pooling convolutional image.
[0134] Step S222: Input the first downsampled RGB image into the second pooling convolutional module of the first fusion module to obtain a second pooling convolutional image.
[0135] Step S223: Concatenate the first pooling convolutional image and the second pooling convolutional image to obtain a first fusion image.
[0136] In one implementation, the second fusion module includes: a first pooling convolutional module and a second pooling convolutional module; Step S240 includes:
[0137] Step S241: Input the second downsampled depth image into the first pooling convolutional module of the second fusion module to obtain a first pooling convolutional image.
[0138] Step S242: Input the second downsampled RGB image into the second pooling convolutional module of the second fusion module to obtain a second pooling convolutional image.
[0139] Step S243: Concatenate the first pooled convolutional image and the second pooled convolutional image to obtain a second fused image.
[0140] In one implementation, the third fusion module includes: a first pooled convolutional module and a second pooled convolutional module; Step S260 includes:
[0141] Step S261: Input the third downsampled depth image into the first pooled convolutional module of the third fusion module to obtain a first pooled convolutional image.
[0142] Step S262: Input the third downsampled RGB image into the second pooled convolutional module of the third fusion module to obtain a second pooled convolutional image.
[0143] Step S263: Concatenate the first pooled convolutional image and the second pooled convolutional image to obtain a third fused image.
[0144] In one implementation, the fourth fusion module includes: a first pooled convolutional module and a second pooled convolutional module; Step S280 includes:
[0145] Step S281: Input the fourth downsampled depth image into the first pooled convolutional module of the fourth fusion module to obtain a first pooled convolutional image.
[0146] Step S282: Input the fourth downsampled RGB image into the second pooled convolutional module of the fourth fusion module to obtain a second pooled convolutional image.
[0147] Step S283: Concatenate the first pooled convolutional image and the second pooled convolutional image to obtain a fourth fused image.
[0148] Specifically, as Figure 5 shown, the first pooled convolutional module includes: a pooling layer, a 1×1 convolutional layer, and a 1×1 convolutional layer connected in sequence. The second pooled convolutional module includes: a pooling layer, a 1×1 convolutional layer, and a 1×1 convolutional layer connected in sequence.
[0149] In one implementation, the multi-granularity segmentation output module includes: segmentation output modules of different granularity levels, and each segmentation output module of a granularity level includes: a second upsampling convolutional module, a second convolutional layer, a second upsampling depth convolutional module, and a third upsampling depth convolutional module. Step S500 specifically includes:
[0150] Step S510: Input the third upsampled image into the second convolutional layer to obtain a second convolutional image.
[0151] Step S520: Input the third refined segmentation image of the grain size level into the second upsampling convolution module in the segmentation output module corresponding to the grain size level to obtain a second upsampling convolution image.
[0152] Step S530: Concatenate the second convolution image and the first upsampling convolution image and input them into the second upsampling depth convolution module to obtain a second depth convolution image.
[0153] Step S540: Input the second depth convolution image into the third upsampling depth convolution module to obtain a scene segmentation image corresponding to the grain size level.
[0154] Specifically, as Figure 7 shown, the second upsampling convolution module includes: 1 upsampling layer and 2 3×3 convolution layers. The second convolution layer uses a 3×3 convolution layer. The second upsampling depth convolution module includes: an upsampling layer and a 3×3 depth convolution layer connected in sequence. The third upsampling depth convolution module includes: an upsampling layer and a 3×3 depth convolution layer connected in sequence. The multi-grain segmentation output module restores the feature map to the input size, so as to generate a pixel-level segmentation result.
[0155] After the encoded image is obtained by feature extraction in the encoding network, the pyramid pooling module is used to process features with different pooling scales, which can aggregate global and local context information. As Figure 8 shown, the pyramid pooling module includes: a pooling layer, a fourth convolution layer, and an upsampling layer. Step S300 includes:
[0156] Step S310: Input different regions of the encoded image into the pooling layer to obtain region pooling images of different sizes.
[0157] Step S320: Input the region pooling images of different sizes into the fourth convolution layer respectively to obtain region convolution images of different sizes.
[0158] Step S330: Input the region convolution images of different sizes into the upsampling layer respectively to obtain corresponding region upsampling images; the sizes of the region upsampling images are all the same as the size of the encoded image.
[0159] Step S340: Concatenate the encoded image and all the region upsampling images to obtain the pooled image.
[0160] For segmentation supervision, cross-entropy loss and an additional weight term are adopted. Based on the number of pixels presented in the ground truth segmentation map, the decoding network is trained based on the weighted cross-entropy loss function; the weighted cross-entropy loss function is:
[0161]
[0162] Among them, represents the category of the entity, represents the weighted cross-entropy loss function, w i represents the weight of the category of the i-th entity, y i represents the label image of the category of the i-th entity, p i represents the scene segmentation image of different granularity levels of the category of the i-th entity.
[0163] As Figure 9 shown, the two images under the vertical bar of RGB-B are the RGB image and the depth image respectively, the two images under the vertical bar of Area are the label image (i.e., the ground truth, GT) at the region level and the scene segmentation image (i.e., the prediction image, Pred) at the region level respectively, the two images under the vertical bar of Entity are the label image at the entity level and the scene segmentation image at the entity level respectively, and the two images under the vertical bar of Part are the label image at the part level and the scene segmentation image at the part level respectively.
[0164] The advantages of this invention are as follows:
[0165] 1) It integrates the scene perception division of multiple granularity levels into a unified encoding-decoding model, efficiently reusing the multi-scale features extracted by the deep convolutional neural network to perform multi-granularity scene semantic segmentation simultaneously, improving the computational efficiency, reducing the model running time consumption, and being able to better meet the requirements of human-machine collaboration applications with high real-time requirements.
[0166] 2) Compared with the encoding part of the existing deep learning models for scene semantic segmentation, the model proposed in this invention adopts a more advanced image-depth dual encoding network to extract scene visual information, and proposes an image-depth feature fusion module based on the channel attention mechanism to fuse the extracted multi-scale visual information, being able to better obtain more robust and stronger representative features.
[0167] 3) Compared with the decoding part of the existing deep learning models for scene semantic segmentation, the decoding part of the model proposed in this invention proposes network component modules for multi-granularity scene segmentation such as multi-scale additional segmentation task supervision, bottom-up layer-by-layer refinement connection, and multi-branch segmentation output end, being able to better optimize the multi-scale decoding process and provide the required final multi-granularity segmentation output.
[0168] The encoding-decoding model of this invention is trained by the following steps:
[0169] Obtain the RGB image and the depth image of the human-machine coexistence scene;
[0170] Input the RGB image and the depth image into an encoding-decoding model to obtain scene segmentation images at different granularity levels;
[0171] Update the parameters of the encoding-decoding model according to the scene segmentation images and label images at different granularity levels;
[0172] When the preset training conditions are met, obtain a trained encoding-decoding model.
[0173] Specifically, during the training process, based on the weighted cross-entropy loss function, perform backpropagation to calculate the gradients of the network parameters of the encoding-decoding model and update the parameters according to the gradients.
[0174] Based on the multi-granularity human-machine coexistence environment perception method based on deep learning according to any one of the above embodiments, the present invention also provides an embodiment of a computer device. The computer device of the present invention includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0175] Obtain the RGB image and the depth image of the human-machine coexistence scene; wherein, the RGB image and the depth image are images taken of the same human-machine coexistence scene;
[0176] Input the RGB image and the depth image into an encoding network to obtain an encoded image;
[0177] Input the encoded image into a pyramid pooling module to obtain a pooled image;
[0178] Input the pooled image into a decoding network to obtain a decoded image;
[0179] Input the decoded image into a multi-granularity segmentation output module to obtain scene segmentation images at different granularity levels;
[0180] When the processor executes the computer program, the following steps are also implemented:
[0181] Input the pooled image into a first upsampling module to obtain first refined segmentation images and a first upsampled image at different granularity levels;
[0182] Input the first upsampled image and the first refined segmentation image into the second upsampling module to obtain second refined segmentation images and a second upsampled image at different granularity levels;
[0183] Input the second upsampled image and the second refined segmentation image into the third upsampling module to obtain third refined segmentation images and a third upsampled image at different granularity levels, and use the third refined segmentation image and the third upsampled image as the decoded image.
[0184] Based on the multi-granularity human-machine coexistence environment perception method based on deep learning according to any of the above embodiments, the present invention also provides an embodiment of a computer-readable storage medium. The computer-readable storage medium of the present invention stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0185] Obtain the RGB image and the depth image of the human-machine coexistence scene; wherein, the RGB image and the depth image are images obtained by photographing the same human-machine coexistence scene;
[0186] Input the RGB image and the depth image into an encoding network to obtain an encoded image;
[0187] Input the encoded image into a pyramid pooling module to obtain a pooled image;
[0188] Input the pooled image into a decoding network to obtain a decoded image;
[0189] Input the decoded image into a multi-granularity segmentation output module to obtain scene segmentation images of different granularity levels;
[0190] When the computer program is executed by a processor, the following steps are also implemented:
[0191] Input the pooled image into a first upsampling module to obtain first refined segmentation images and a first upsampled image of different granularity levels;
[0192] Input the first upsampled image and the first refined segmentation image into the second upsampling module to obtain second refined segmentation images and a second upsampled image of different granularity levels;
[0193] Input the second upsampled image and the second refined segmentation image into the third upsampling module to obtain third refined segmentation images and a third upsampled image of different granularity levels, and use the third refined segmentation image and the third upsampled image as the decoded image.
[0194] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent replacements or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A multi-granularity human-machine coexistence environment perception method based on deep learning, characterized in that The method includes the steps of: Obtaining an RGB image and a depth image of a human-machine coexistence scenario; wherein, the RGB image and the depth image are images obtained by photographing the same human-machine coexistence scenario; Inputting the RGB image and the depth image into an encoding network to obtain an encoded image; Inputting the encoded image into a pyramid pooling module to obtain a pooled image; Inputting the pooled image into a decoding network to obtain a decoded image; Inputting the decoded image into a multi-granularity segmentation output module to obtain scene segmentation images of different granularity levels; Wherein, the granularity levels include a region level, an entity level, and a part level of the entity; the encoding network includes: a first downsampling module, a second downsampling module, a third downsampling module, a fourth downsampling module, a first fusion module, a second fusion module, a third fusion module, and a fourth fusion module; the pyramid pooling module includes: a pooling layer, a fourth convolutional layer, and an upsampling layer; The decoding network includes: a first upsampling module, a second upsampling module, and a third upsampling module; the first fusion module is skip-connected to the third upsampling module; the second fusion module is skip-connected to the second upsampling module; the third fusion module is skip-connected to the first upsampling module; the third upsampling module includes: a first convolutional module, a first upsampling depth convolutional module, and refinement segmentation modules of different granularity levels; the refinement segmentation module includes: a first upsampling convolutional module and a first convolutional layer; the first fusion module includes: a first pooling convolutional module and a second pooling convolutional module; the multi-granularity segmentation output module includes: segmentation output modules of different granularity levels, and each segmentation output module of a granularity level includes: a second upsampling convolutional module, a second convolutional layer, a second upsampling depth convolutional module, and a third upsampling depth convolutional module; The step of inputting the pooled image into the decoding network to obtain a decoded image includes: Inputting the pooled image into the first upsampling module to obtain first refinement segmentation images of different granularity levels and a first upsampled image; Inputting the first upsampled image and the first refinement segmentation images into the second upsampling module to obtain second refinement segmentation images of different granularity levels and a second upsampled image; Inputting the second upsampled image and the second refinement segmentation images into the third upsampling module to obtain third refinement segmentation images of different granularity levels and a third upsampled image, and taking the third refinement segmentation images and the third upsampled image as the decoded image.
2. The multi-granularity human-machine coexistence environment perception method based on deep learning according to claim 1, characterized in that The step of inputting the RGB image and the depth image into the encoding network to obtain an encoded image includes: Respectively inputting the depth image and the RGB image into the first downsampling module to obtain a first downsampled depth image and a first downsampled RGB image; Inputting the first downsampled depth image and the first downsampled RGB image into the first fusion module to obtain a first fused image; Respectively inputting the first downsampled depth image and the first fused image into the second downsampling module to obtain a second downsampled depth image and a second downsampled RGB image; Input the second downsampled depth image and the second downsampled RGB image into the second fusion module to obtain a second fusion image; Input the second downsampled depth image and the second fusion image into the third downsampling module respectively to obtain a third downsampled depth image and a third downsampled RGB image; Input the third downsampled depth image and the third downsampled RGB image into the third fusion module to obtain a third fusion image; Input the third downsampled depth image and the third fusion image into the fourth downsampling module respectively to obtain a fourth downsampled depth image and a fourth downsampled RGB image; Input the fourth downsampled depth image and the fourth downsampled RGB image into the fourth fusion module to obtain a fourth fusion image, and use the fourth fusion image as the encoded image.
3. The multi-granularity human-machine coexistence environment perception method based on deep learning according to claim 2, wherein, The step of inputting the second upsampled image and the second refined segmentation image into the third upsampling module to obtain third refined segmentation images and a third upsampled image at different granularity levels includes: Input the second upsampled image into the first convolutional module to obtain a first convolutional image; Input the first convolutional image into the first upsampled depth convolutional module to obtain a first depth convolutional image; Concatenate the first depth convolutional image and the first fusion image to obtain a third upsampled image; Input the first convolutional image and the second refined segmentation image at the granularity level into the refined segmentation module at the corresponding granularity level to obtain a third refined segmentation image at the corresponding granularity level.
4. The multi-granularity human-machine coexistence environment perception method based on deep learning according to claim 3, wherein, The step of inputting the first convolutional image and the second refined segmentation image at the granularity level into the refined segmentation module at the corresponding granularity level to obtain a third refined segmentation image at the corresponding granularity level includes: Input the second refined segmentation image at the granularity level into the first upsampled convolutional module to obtain a first upsampled convolutional image; Concatenate the first convolutional image and the first upsampled convolutional image and then input them into the first convolutional layer to obtain a third refined segmentation image at the corresponding granularity level.
5. The multi-granularity human-machine coexistence environment perception method based on deep learning according to claim 2, characterized in that, The step of inputting the first downsampled depth image and the first downsampled RGB image into the first fusion module to obtain a first fusion image includes: Input the first downsampled depth image into the first pooling convolutional module to obtain a first pooling convolutional image; Input the first downsampled RGB image into the second pooling convolutional module to obtain a second pooling convolutional image; Concatenate the first pooling convolutional image and the second pooling convolutional image to obtain a first fusion image.
6. The multi-granularity human-machine coexistence environment perception method based on deep learning according to claim 2, wherein, The step of inputting the decoded image into the multi-granularity segmentation output module to obtain scene segmentation images at different granularity levels includes: Input the third upsampled image into the second convolutional layer to obtain a second convolutional image; Input the third refined segmentation image at the granularity level into the second upsampled convolutional module in the segmentation output module at the corresponding granularity level to obtain a second upsampled convolutional image; Concatenate the second convolutional image and the second upsampled convolutional image and then input them into the second upsampled depth convolutional module to obtain a second depth convolutional image; Input the second depth convolution image into the third upsampling depth convolution module to obtain a scene segmentation image corresponding to the granularity level.
7. The multi-granularity human-machine coexistence environment perception method based on deep learning according to claim 1, characterized in that, The step of inputting the encoded image into the pyramid pooling module to obtain a pooled image includes: Input different regions of the encoded image into the pooling layer to obtain region pooled images of different sizes; Input the region pooled images of different sizes into the fourth convolution layer respectively to obtain region convolution images of different sizes; Input the region convolution images of different sizes into the upsampling layer respectively to obtain corresponding region upsampled images; the sizes of all the region upsampled images are the same as the size of the encoded image; Concatenate the encoded image and all the region upsampled images to obtain the pooled image.
8. The method for multi-granularity human-machine co-existing environment perception based on deep learning according to claim 1, characterized in that Train the decoding network based on the weighted cross-entropy loss function; The weighted cross-entropy loss function is: Among them, represents the category of the entity, represents the weighted cross-entropy loss function, w i represents the weight of the category of the i-th entity, y i represents the label image of the category of the i-th entity, p i represents the scene segmentation image of different granularity levels of the category of the i-th entity.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.