Method for detecting a target, apparatus for detecting a target, electronic device, and computer storage medium
Patent Information
- Application Number
- JP2024532488
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-30
- Filing Date
- 2022-11-17
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2042-11-17
AI Technical Summary
Existing target detection methods have low detection efficiency, which hinders their effectiveness in applications such as automated assembly by intelligent robots.
A method involving instance segmentation to obtain a segmentation mask, matching positional relationship features between target and standard pixels, and using a pre-trained instance segmentation model with feature extraction and fusion networks to enhance detection accuracy and efficiency.
Improves target detection efficiency by reducing data matching volume and enhancing accuracy, particularly in noisy and occluded scenarios, while maintaining robustness.
Smart Images

Figure 00000017_0000 
Figure 00000018_0000 
Figure 00000019_0000
Abstract
Description
[Technical field]
[0001] TECHNICAL FIELD Embodiments of the present application relate to the field of image processing technology, and in particular to a method for detecting a target, an apparatus for detecting a target, an electronic device, and a computer storage medium.
[0002] background Target detection technology can be applied to various scenarios as the technology matures. For example, in fields such as industrial production, workpieces can be automatically picked and assembled by using an intelligent robot through target detection technology. Specifically, an image containing a workpiece may be first acquired, and then target detection is performed on the image to obtain pose information (position information and attitude information) of the target workpiece, and then the intelligent robot acquires the target workpiece and assembles the target workpiece according to the pose information.
[0003] In the existing target detection methods, the detection efficiency in use is relatively low. Therefore, how to improve the target detection efficiency is an urgent problem to be solved.
[0004] overview In view of this, to solve this technical problem, the embodiments of the present application provide a method for detecting a target, an apparatus for detecting a target, an electronic device, and a computer storage medium to solve the deficiencies of relatively low detection efficiency in the related art.
[0005] According to a first aspect, an embodiment of the present application provides a method for detecting a target, the method for detecting a target comprising: obtaining a target image including the target object; performing instance segmentation on the target image to obtain a segmentation mask corresponding to the target object; and obtaining, based on the segmentation mask, position relationship features between target pixels in a target region where the target object is located in the target image; Obtaining standard pixel-to-pixel positional relationship characteristics within a preset region of interest of a target object in a standard image; and The method includes matching positional relationship features between the target pixels and positional relationship features between the standard pixels to obtain a correspondence relationship between the target pixels and the standard pixels, and obtaining pose information of the target object based on the correspondence relationship.
[0006] Optionally, obtaining, based on the segmentation mask, position relationship features between target pixels in a target region where a target object is located in the target image, includes: The method includes: pairing target pixels in a target region in which a target object is located in the target image based on a segmentation mask to obtain a plurality of target pixel pairs; and obtaining, for each target pixel pair, a position relationship feature between two target pixels in the target pixel pair; To obtain standard pixel-to-pixel relationship features within a preset region of interest in a standard image: acquiring a standard image and a preset region of interest of a target object within the standard image; and The method includes pairwise combining standard pixels within a preset region of interest to obtain a plurality of standard pixel pairs, and obtaining, for each standard pixel pair, a positional relationship feature between two standard pixels in the standard pixel pair.
[0007] Optionally, for each target pixel pair, a position relationship feature between two target pixels in the target pixel pair is obtained based on a distance between the two target pixels, an angle between normal vectors respectively corresponding to the two target pixels, and an angle between the normal vectors corresponding to the two target pixels and a connecting line between the two target pixels; For each standard pixel pair, a positional relationship feature between the two standard pixels in the standard pixel pair is obtained based on the distance between the two standard pixels, the angle between the normal vectors corresponding to the two standard pixels respectively, and the angle between the normal vectors corresponding to the two standard pixels and the connecting line between the two standard pixels.
[0008] Optionally, performing instance segmentation on the target image to obtain a segmentation mask corresponding to the target object includes: The method includes inputting the target image into a pre-trained instance segmentation model and performing instance segmentation on the target image by using the instance segmentation model to obtain a segmentation mask corresponding to the target object.
[0009] Optionally, the instance segmentation model includes: a feature extraction network, a feature fusion network, a region generation network, a feature alignment layer, a classification and regression network, and a segmentation mask network; Performing instance segmentation on a target image by inputting the target image into a pre-trained instance segmentation model and using the instance segmentation model to obtain a segmentation mask corresponding to the target object includes: inputting the target image into a feature extraction network in the pre-trained instance segmentation model, and performing multi-scale feature extraction on the target image by using the feature extraction network to obtain multiple levels of initial feature maps corresponding to the target image; performing feature fusion on the multiple levels of initial feature maps by using a feature fusion network to obtain a fused feature map; Obtaining information about the initial region of the target object by using a region generation network based on the fused feature map; performing feature extraction on the initial feature map by using a feature alignment layer based on information about the initial region to obtain a region feature map in the initial feature map corresponding to the initial region; The method includes obtaining category information and location information of the target object by using a classification and regression network based on the region feature map; and obtaining a segmentation mask corresponding to the target object by using a segmentation mask network based on the region feature map.
[0010] Optionally, performing feature fusion on multiple levels of initial feature maps by using a feature fusion network to obtain a fused feature map includes: performing a convolution operation on each initial feature map by using a feature fusion network to obtain multiple levels of initial dimensionality reduced feature maps; Sequentially performing a fusion process for every two adjacent levels of the initial dimensionality reduced feature maps according to descending order of levels to obtain an initial fused feature map, and updating the initial dimensionality reduced feature map of the lower level of the adjacent level by using the initial fused feature map, where the size of the initial dimensionality reduced feature map of the higher level is smaller than the size of the initial dimensionality reduced feature map of the lower level; performing a convolution operation on each initial fused feature map to obtain multiple levels of dimensionality reduced feature maps; and The method includes sequentially performing a fusion process on every two adjacent levels of the dimensionality reduced feature maps according to an ascending order of the levels to obtain a transition feature map, performing a fusion process on the transition feature map and the initial feature map to obtain a fused feature map, and using the fused feature map to update the upper-level dimensionality reduced feature map of the adjacent level, where the size of the upper-level dimensionality reduced feature map is smaller than the size of the lower-level dimensionality reduced feature map.
[0011] Optionally, the feature extraction network includes two concatenated convolutional layers, in which the size of a convolutional kernel of the previous convolutional layer is 1*1, the convolutional stride of the previous convolutional layer is 1, and the convolutional stride of the subsequent convolutional layer is less than or equal to the size of the convolutional kernel of the subsequent convolutional layer.
[0012] According to a second aspect, an embodiment of the present application provides an apparatus for detecting a target, the apparatus comprising: a target image capture module configured to capture a target image including the target object; a segmentation mask acquisition module configured to perform instance segmentation on the target image to obtain a segmentation mask corresponding to the target object; a first positional relationship feature acquisition module configured to acquire positional relationship features between target pixels in a target region where a target object is located in the target image based on the segmentation mask; a second positional relationship feature acquisition module configured to acquire positional relationship features between standard pixels in a preset region of interest in a standard image including a standard object corresponding to the target object; and a pose information acquisition module configured to match positional relationship features between the target pixels and positional relationship features between the standard pixels to acquire a correspondence relationship between the target pixels and the standard pixels, and acquire pose information of the target object based on the correspondence relationship.
[0013] According to a third aspect, an embodiment of the present application provides an electronic device including a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface completing communication with each other by using the communication bus, the memory being configured to store at least one executable instruction, the executable instruction causing the processor to implement the method of detecting a target according to the first aspect.
[0014] According to a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program which, when executed by a processor, implements a method of detecting a target according to the first aspect or any embodiment thereof.
[0015] The embodiments of the present application provide a method for detecting a target, an apparatus for detecting a target, an electronic device, and a computer storage medium. In the embodiments of the present application, after obtaining the positional relationship feature between target pixels in a target region in a target image, the positional relationship feature is matched with the positional relationship feature between standard pixels in a preset region of interest of a target object in a standard image, and pose information of the target object can be obtained based on the correspondence between the target pixels and the standard pixels obtained through matching. Since the preset region of interest is only a part of the entire target object, the amount of data of the positional relationship feature between standard pixels in the preset region of interest is relatively small compared with the positional relationship feature between all pixels of the target object in the entire standard image, so when the positional relationship feature between target pixels matches the positional relationship feature between standard pixels, the amount of data of the feature to be matched is also relatively small. Therefore, the matching speed is increased, thereby improving the overall efficiency of target detection.
[0016] Certain particular embodiments of the present application will now be described by way of example, and not by way of limitation, with reference to the accompanying drawings, in which like reference numbers indicate the same or similar components or parts. Those skilled in the art will appreciate that the accompanying drawings are not necessarily drawn to scale. [Brief description of the drawings]
[0017] [Figure 1] 1 is a schematic flow chart of a method for detecting a target according to an embodiment of the present application; [Diagram 2] FIG. 2 is a schematic diagram of a positional relationship feature according to an embodiment of the present application. [Diagram 3] 1 is a schematic flow chart of obtaining a fused feature map according to an embodiment of the present application; [Figure 4] FIG. 1 is a schematic diagram of the convolution process of an existing feature extraction network. [Diagram 5]FIG. 2 is a schematic diagram of a convolution process of a feature extraction network according to an embodiment of the present application; [Figure 6] 1 is a schematic diagram of an apparatus for detecting a target according to an embodiment of the present application;
[0018] Detailed Description In order to allow those skilled in the art to better understand the technical solution of the present application, the following will clearly and completely describe the technical solution of the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. The described embodiments are obviously not all the embodiments of the present application, but only some embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application shall fall within the protection scope of the embodiments of the present application.
[0019] It should be noted that the terms "first" and "second" in this application are only for distinguishing names, do not represent an order relationship, and cannot be understood as indicating or implying relative importance or implicitly specifying the number of technical features shown. For example, the first user, the second user and the third user are only for distinguishing different users.
[0020] Hereinafter, specific implementations of the embodiments of the present application will be further described with reference to the accompanying drawings of the embodiments of the present application.
[0021] An embodiment of the present application provides a method for detecting a target. Figure 1 is a schematic flow chart of a method for detecting a target according to an embodiment of the present application. As shown in Figure 1, the method includes the following steps:
[0022] Step 101. Obtain a target image containing the target object.
[0023] For example, in a scenario where an intelligent robot is performing a picking task, an image including a target workpiece may be acquired as a target image. This embodiment of the present application is not limited to a specific manner of acquiring a target image.
[0024] Step 102. Perform instance segmentation on the target image to obtain a segmentation mask corresponding to the target object.
[0025] Step 103. Based on the segmentation mask, obtain position relationship features between target pixels in the target region where the target object is located in the target image.
[0026] In this step, the positional relationship feature between the target pixels in the target region may be a feature representing a relative positional relationship between the pixels in the target region. The positional relationship feature may be obtained based on the coordinates of the pixels in the target region.
[0027] Step 104: Obtain standard inter-pixel positional relationship features within a preset region of interest of the target object in the standard image.
[0028] The standard image is an image acquired for a target object having a standard posture at a standard position. Both the standard position and the standard posture may be preset according to actual requirements. The scenario of an intelligent robot picking up a task is still used as an example. It is assumed that the target object is a rectangular workpiece to be assembled, the standard position may be a preset position that is convenient for the robot to assemble the workpiece, for example, a preset position on an assembly table, and the standard posture may be a preset posture that is convenient for the robot to assemble the workpiece, for example, the long side of the workpiece is parallel to the edge of the assembly table. Correspondingly, an image acquired for a target object whose long side is parallel to the edge of the assembly table and is at a preset position of the assembly table may be the standard image.
[0029] In this step, the preset region of interest may be a region that can represent a specific attribute of the target object, or the preset region of interest may be a specific region that distinguishes the target object from other non-target objects. For example, if the target object is a workpiece, the preset region of interest may be a specific texture region, a specific structure region, a specific text region, etc. that distinguishes the workpiece from other workpieces.
[0030] In practical applications, the preset regions of interest may be dynamically set and adjusted according to different practical scenarios, and in this embodiment of the present application, the setting manner of the preset regions of interest is not limited.
[0031] Corresponding to the positional relationship feature between the target pixels in the target region, the positional relationship feature between the standard pixels may be a feature representing a relative positional relationship between pixels in the preset region of interest. The positional relationship feature may be obtained based on the coordinates of the pixels in the preset region of interest.
[0032] Step 105. Match the positional relationship features between the target pixels and the standard pixels to obtain the corresponding relationship between the target pixels and the standard pixels, and obtain the pose information of the target object based on the corresponding relationship.
[0033] The pose information of the target object includes position information and orientation information of the target object.
[0034] Specifically, after the correspondence between the target pixel and the standard pixel is obtained, initial pose information of the target object in the target image may be obtained based on the correspondence. Specifically, the initial pose information may include initial position information (initial target area) of the target object in the target image and initial attitude information (initial angle information) of the target object in the target image relative to the target object in the standard image. Then, the initial target area is rotated by an initial attitude angle to obtain a rotated initial target area; the initial target area and the initial rotation angle are iteratively adjusted based on the correspondence between the target pixel and the standard pixel to obtain more accurate position information and attitude information of the target object as the pose information of the target object in the target image.
[0035] In the embodiment of the present application, after obtaining the positional relationship feature between the target pixels in the target region in the target image, the positional relationship feature is matched with the positional relationship feature between the standard pixels in the preset region of interest of the target object in the standard image, and pose information of the target object can be obtained based on the correspondence between the target pixels and the standard pixels obtained through matching. Since the preset region of interest is only a part of the entire target object, the amount of data of the positional relationship feature between the standard pixels in the preset region of interest is relatively small compared with the positional relationship feature between all pixels of the target object in the entire standard image, so when the positional relationship feature between the target pixels matches the positional relationship feature between the standard pixels, the amount of data of the feature to be matched is also relatively small. Therefore, the matching speed is increased, thereby improving the overall efficiency of target detection.
[0036] Optionally, in an embodiment of the present application, step 103 includes the following steps: To obtain multiple target pixel pairs, the method may be implemented as a step of combining target pixels in a target region where a target object is located in a target image in pairs based on a segmentation mask, and obtaining, for each target pixel pair, a positional relationship feature between the two target pixels in the target pixel pair.
[0037] Correspondingly, step 104 may be divided into the following steps: This may be implemented as the steps of obtaining a standard image and a preset region of interest of a target object in the standard image; combining standard pixels of the preset region of interest in the standard image in pairs to obtain a plurality of standard pixel pairs; and obtaining, for each standard pixel pair, a positional relationship feature between two standard pixels in the standard pixel pair.
[0038] Specifically, the positional relationship feature between two target pixels may be a feature representing the relative positional relationship between two target pixels; correspondingly, the positional relationship feature between two standard pixels may be a feature representing the relative positional relationship between two standard pixels.
[0039] Optionally, in an embodiment of the present application, for each target pixel pair, a position relationship feature between two target pixels in the target pixel pair is obtained based on a distance between the two target pixels, an angle between normal vectors respectively corresponding to the two target pixels, and an angle between the normal vectors corresponding to the two target pixels and a connecting line between the two target pixels; For each standard pixel pair, a positional relationship feature between two standard pixels in the standard pixel pair is obtained based on the distance between the two standard pixels, the angle between the normal vectors corresponding to the two standard pixels respectively, and the angle between the normal vectors corresponding to the two standard pixels and the connecting line between the two standard pixels.
[0040] Specifically, for each target pixel pair, by using the distance F1 between the two target pixels, the angle F2 between the normal vectors corresponding to the two target pixels respectively, and the angle (F3 and F4) between the normal vectors corresponding to the two target pixels and the connecting line between the two target pixels, a four-dimensional vector (F1, F2, F3 and F4) may be constructed and used as the positional relationship feature between the two target pixels in the target pixel pair. FIG. 2 is a schematic diagram of the positional relationship feature according to an embodiment of the present application. In FIG. 2, the target pixels are m1 and m2 respectively, and F1 is the distance between m1 and m2. N1 is the normal vector corresponding to m1, N2 is the normal vector corresponding to m2, and F2 is the angle (may be expressed in radians) between N1 and N2. F3 is the angle (may be expressed in radians) between N1 and F1. F4 is the angle (may be expressed in radians) between N2 and F1. Correspondingly, the position relationship feature F between the target pixel m1 and the target pixel m2 is (F1, F2, F3 and F4).
[0041] Correspondingly, the positional relationship feature between the two standard pixels in each standard pixel pair may be constructed in this manner, the details of which will not be described again here.
[0042] In this embodiment of the present application, the positional relationship feature between two pixels in a pixel pair is obtained based on the distance between the two pixels, the angle between the normal vectors corresponding to the two pixels, and the angle between the normal vectors corresponding to the two pixels and the connecting line between the two pixels. That is, in this embodiment of the present application, the positional relationship feature between the two pixels is expressed in four different dimensions jointly. Therefore, the obtained positional relationship feature can more accurately represent the relative positional relationship between the two pixels, and the target pixel is matched to the standard pixel based on the positional relationship feature to obtain a more accurate correspondence between the target pixel and the standard pixel, thereby improving the accuracy of target detection.
[0043] Optionally, in an embodiment of the present application, step 102 includes the following steps: It may be implemented as a step of performing instance segmentation on the target image by inputting the target image into a pre-trained instance segmentation model and using the instance segmentation model to obtain a segmentation mask corresponding to the target object.
[0044] Furthermore, in an embodiment of the present application, the instance segmentation model may include a feature extraction network, a feature fusion network, a region generation network, a feature alignment layer, a classification and regression network, and a segmentation mask network, and performing instance segmentation on a target image by inputting the target image into the pre-trained instance segmentation model and using the instance segmentation model to obtain a segmentation mask corresponding to a target object includes: inputting the target image into a feature extraction network in the pre-trained instance segmentation model, and performing multi-scale feature extraction on the target image by using the feature extraction network to obtain multiple levels of initial feature maps corresponding to the target image; performing feature fusion on the multiple levels of initial feature maps by using a feature fusion network to obtain a fused feature map; Obtaining information about the initial region of the target object by using a region generation network based on the fused feature map; performing feature extraction on the initial feature map by using a feature alignment layer based on information about the initial region to obtain a region feature map in the initial feature map corresponding to the initial region; The method may include obtaining category information and location information of the target object by using a classification and regression network based on the region feature map, and obtaining a segmentation mask corresponding to the target object by using a segmentation mask network based on the region feature map.
[0045] Optionally, in an embodiment of the present application, performing feature fusion on multiple levels of initial feature maps by using a feature fusion network to obtain a fused feature map, as a subsequent step, namely: performing a convolution operation on each initial feature map by using a feature fusion network to obtain multiple levels of initial dimensionality-reduced feature maps; sequentially performing a fusion process for every two adjacent levels of the initial dimensionality reduced feature maps according to descending order of levels to obtain an initial fused feature map, and updating the initial dimensionality reduced feature map of the lower level of the adjacent level by using the initial fused feature map, where the size of the initial dimensionality reduced feature map of the higher level is smaller than the size of the initial dimensionality reduced feature map of the lower level; performing a convolution operation on each initial fused feature map to obtain multiple levels of dimensionality reduced feature maps; It may be implemented as the steps of: sequentially performing a fusion process for every two adjacent levels of the dimensionality reduced feature maps according to ascending order of levels to obtain a transition feature map; performing a fusion process on the transition feature map and the initial feature map to obtain a fused feature map; and using the fused feature map to update the dimensionality reduced feature map of the higher level of the adjacent level, where the size of the dimensionality reduced feature map of the higher level is smaller than the size of the dimensionality reduced feature map of the lower level.
[0046] Specifically, Figure 3 is a schematic flow chart of obtaining a fused feature map according to an embodiment of the present application. C1, C2, C3, C4 and C5 are five levels of initial feature maps corresponding to a target image obtained by using a feature extraction network in an instance segmentation model, respectively. The process of performing a fusion process on the five initial feature maps by using a feature fusion network to obtain a fused feature map is as follows:
[0047] Since the C1 level of the initial feature map contains relatively little semantic information, feature fusion may be performed based on C2, C3, C4 and C5. Specifically, to obtain multiple levels of initial dimensionality-reduced feature maps, convolution operations (the size of the convolution kernel may be 1*1) are performed on C2, C3, C4 and C5 respectively, where a convolution operation is performed on C5 to obtain an initial dimensionality-reduced feature map P5. Then, to obtain an initial fused feature map P4, P5 and the initial dimensionality-reduced feature map corresponding to C4 (the feature map obtained after the convolution operation is performed on C4) are first fused according to descending order of levels (in order to obtain a feature map with the same size as the initial dimensionality-reduced feature map, before fusion is performed, upsampling needs to be performed on P5, and then the element corresponding to P5 obtained after upsampling and the element corresponding to the initial dimensionality-reduced feature map corresponding to C4 may be added to obtain the initial fused feature map P4). Furthermore, the initial dimensionality-reduced feature map corresponding to C4 is updated by using P4. Then, P4 is fused with the initial dimension-reduced feature map corresponding to C3 (the feature map obtained after performing a convolution operation on C3) to obtain an initial fused feature map P3 (in order to obtain a feature map of the same size as the initial dimension-reduced feature map, upsampling needs to be performed on P4 before fusion is performed, and then the elements corresponding to P4 obtained after upsampling and the elements corresponding to the initial dimension-reduced feature map corresponding to C3 may be added to obtain the initial fused feature map P3). Furthermore, downsampling may be further performed on P5 to obtain P6. So far, the initial fused feature maps P6, P5, P4 and P3 have been obtained. Convolution operations are respectively performed on P6, P5, P4 and P3 to obtain multiple levels of dimension-reduced feature maps, where a convolution operation is performed on P3 to obtain a dimension-reduced feature map N2.Then, N2 and the dimension-reduced feature map corresponding to P4 are first fused according to the ascending order of levels to obtain a transition feature map, and the transition feature map is fused with C3 to obtain a fused feature map N3. Furthermore, the dimension-reduced feature map corresponding to P4 is updated by using N3. N3 and the dimension-reduced feature map corresponding to P5 are fused to obtain a transition feature map, and the transition feature map is fused with C4 to obtain a fused feature map N4. Furthermore, the dimension-reduced feature map corresponding to P5 is updated by using N4. N4 and the dimension-reduced feature map corresponding to P6 are fused to obtain a transition feature map, and the transition feature map is fused with C5 to obtain a fused feature map N5. So far, the final fused feature maps N2, N3, N4 and N5 have been obtained.
[0048] In this process, when two different feature maps are to be fused, first, upsampling and enlargement operations may be performed on a feature map with a relatively small size, so that the two feature maps to be fused have the same size, and then elements may be added at corresponding positions in the feature maps of the same size to obtain a fused feature map.
[0049] In the feature fusion manner provided in the embodiment of the present application, the obtained fused feature map is obtained after performing multiple levels of fusion processing on the initial feature map. For example, the fused feature map N3 integrates the features in P3, P4 and C3 simultaneously, the fused feature map N4 integrates the features in P4, P5 and C4 simultaneously, and the fused feature map N5 integrates the features in P5, P6 and C5 simultaneously. Therefore, in this embodiment of the present application, the obtained fused feature map can include more features, especially smaller features, so that the method for detecting a target provided in this embodiment of the present application has higher detection accuracy. Moreover, the method for detecting a target provided in this embodiment of the present application also has better robustness in the case of noise, clutter and partial occlusion.
[0050] Optionally, in one embodiment of the present application, the feature extraction network includes two concatenated convolutional layers, in which the size of a convolutional kernel of the previous convolutional layer is 1*1, the convolutional stride of the previous convolutional layer is 1, and the convolutional stride of the subsequent convolutional layer is less than or equal to the size of a convolutional kernel of the subsequent convolutional layer.
[0051] By using the feature extraction network provided in the embodiments of the present application, more and fewer features may be extracted, which can improve the accuracy of target detection.
[0052] Hereinafter, the advantageous effects of the embodiments of the present application will be described in detail by using specific examples.
[0053] 4 is a schematic diagram of a convolution processing process of an existing feature extraction network, and FIG. 5 is a schematic diagram of a convolution processing process of a feature extraction network according to an embodiment of the present application. Specifically,
[0054] 4, the existing feature extraction network includes two concatenated convolution layers, where the size of the convolution kernel H1 of the previous convolution layer is 1*1, the convolution stride of the previous convolution layer is 2, the size of the convolution kernel H2 of the subsequent convolution layer is 3*3, and the convolution stride of the subsequent convolution is 1. By using the feature extraction network, feature extraction is performed on the image.
[0055] 5, the feature extraction network in this embodiment of the present application includes two convolution layers connected in series, where the convolution kernel of the previous convolution layer is also H1, but the convolution stride of the previous convolution layer is 1, and the convolution kernel of the subsequent convolution layer is also H2, but the convolution stride of the subsequent convolution layer is 2. By using the feature extraction network, feature extraction is performed on the image.
[0056] By comparing Fig. 4 and Fig. 5, it can be seen that in Fig. 4, when passing through the preceding convolution layer, all pixels in the second column, the fourth column, the second row and the fourth row in the image are skipped during convolution, i.e., these pixels do not participate in the convolution operation. However, in Fig. 5, when passing through the preceding convolution layer, all pixels in the image participate in the convolution operation. Therefore, the feature extraction network provided in this embodiment of the present application may provide more and less features than existing feature extraction networks. Thus, the accuracy of target detection can be improved.
[0057] According to the method for detecting a target provided in any one of the embodiments, an embodiment of the present application provides an apparatus for detecting a target. As shown in FIG. 6, FIG. 6 is a schematic diagram of an apparatus for detecting a target according to an embodiment of the present application. The apparatus 60 for detecting a target includes a target image acquisition module 601, a segmentation mask acquisition module 602, a first positional relationship feature acquisition module 603, a second positional relationship feature acquisition module 604, and a pose information acquisition module 605.
[0058] The first target image acquisition module 601 is configured to acquire a target image including a target object.
[0059] The segmentation mask obtainment module 602 is configured to perform instance segmentation on the target image to obtain a segmentation mask corresponding to the target object.
[0060] The first positional relationship feature acquiring module 603 is configured to acquire, based on the segmentation mask, positional relationship features between target pixels in a target region where a target object is located in the target image.
[0061] The second positional relationship feature acquisition module 604 is configured to acquire positional relationship features between standard pixels in a preset region of interest in a standard image containing a standard object corresponding to the target object.
[0062] The pose information acquisition module 605 is configured to acquire a correspondence relationship between the target pixels and the standard pixels, and match the positional relationship features between the target pixels and the positional relationship features between the standard pixels to acquire pose information of the target object based on the correspondence relationship.
[0063] Optionally, in one embodiment, the first location relationship characteristic acquisition module 602 further comprises: To obtain a plurality of target pixel pairs, the method is configured to pair-combine target pixels in a target region in which a target object is located in the target image based on a segmentation mask, and obtain, for each target pixel pair, a positional relationship feature between two target pixels in the target pixel pair.
[0064] The second positional relationship feature acquisition module 603 is further configured to acquire a standard image and a preset region of interest in the standard image, pair-wise combine standard pixels in the preset region of interest in the standard image to acquire a plurality of standard pixel pairs, and acquire, for each standard pixel pair, a positional relationship feature between two standard pixels in the standard pixel pair.
[0065] Optionally, in an embodiment of the present application, for each target pixel pair, a position relationship feature between two target pixels in the target pixel pair is obtained based on a distance between the two target pixels, an angle between normal vectors respectively corresponding to the two target pixels, and an angle between the normal vectors corresponding to the two target pixels and a connecting line between the two target pixels; For each standard pixel pair, a positional relationship feature between the two standard pixels in the standard pixel pair is obtained based on the distance between the two standard pixels, the angle between the normal vectors corresponding to the two standard pixels respectively, and the angle between the normal vectors corresponding to the two standard pixels and the connecting line between the two standard pixels.
[0066] Optionally, in an embodiment of the present application, the segmentation mask acquisition module 602 further comprises: The system is configured to perform instance segmentation on the target image by inputting the target image into a pre-trained instance segmentation model and using the instance segmentation model to obtain a segmentation mask corresponding to the target object.
[0067] Optionally, in an embodiment of the present application, the instance segmentation model includes a feature extraction network, a feature fusion network, a feature alignment layer, a classification and regression network, and a segmentation mask network, and the segmentation mask acquisition module 602 further comprises: inputting the target image into a feature extraction network in the pre-trained instance segmentation model, and performing multi-scale feature extraction on the target image by using the feature extraction network to obtain a multi-level initial feature map corresponding to the target image; performing feature fusion on the multiple-level initial feature maps by using a feature fusion network to obtain a fused feature map; Obtain information about the initial region of the target object based on the fused feature map by using a region generation network; performing feature extraction on the initial feature map by using a feature alignment layer based on information about the initial region to obtain a region feature map in the initial feature map corresponding to the initial region; The classification and regression network is configured to obtain category information and location information of the target object based on the region feature map, and the segmentation mask network is configured to obtain a segmentation mask corresponding to the target object based on the region feature map.
[0068] Optionally, in an embodiment of the present application, when performing feature fusion on the multiple-level initial feature maps by using a feature fusion network to obtain a fused feature map, the segmentation mask acquisition module 602 further comprises: Perform a convolution operation on each initial feature map by using a feature fusion network to obtain multiple levels of initial dimensionality reduced feature maps; Sequentially perform a fusion process for each two adjacent levels of the initial dimensionality reduced feature maps according to descending order of levels to obtain an initial fused feature map, and use the initial fused feature map to update the initial dimensionality reduced feature map of the lower level of the adjacent level, where the size of the initial dimensionality reduced feature map of the higher level is smaller than the size of the initial dimensionality reduced feature map of the lower level; Perform a convolution operation on each initial fused feature map to obtain multiple levels of dimensionality reduced feature maps; The method is configured to sequentially perform a fusion process for every two adjacent levels of the dimensionality reduced feature maps according to an ascending order of the levels to obtain a transition feature map, perform a fusion process on the transition feature map and the initial feature map to obtain a fused feature map, and use the fused feature map to update the dimensionality reduced feature map of the upper level of the adjacent level, where the size of the dimensionality reduced feature map of the upper level is smaller than the size of the dimensionality reduced feature map of the lower level.
[0069] Optionally, in one embodiment of the present application, the feature extraction network includes two concatenated convolutional layers, in which the size of a convolutional kernel of the previous convolutional layer is 1*1, the convolutional stride of the previous convolutional layer is 1, and the convolutional stride of the subsequent convolutional layer is less than or equal to the size of a convolutional kernel of the subsequent convolutional layer.
[0070] The target detection device in this embodiment of the present application is configured to implement the target detection method according to the above-mentioned method embodiment, and has the advantageous effects of the corresponding method embodiment. Details will not be described again here. Furthermore, for the functional implementation of each module in the target detection device in this embodiment of the present application, reference can be made to the description of the corresponding part in the method embodiment. Details will not be described again here.
[0071] Based on the method for detecting a target according to any one of the multiple embodiments, one embodiment of the present application provides an electronic device including a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface complete communication with each other by using the communication bus, and the memory is configured to store at least one executable instruction, and the executable instruction causes the processor to perform the method for detecting a target according to any one of the multiple embodiments.
[0072] Based on the method for detecting a target according to any one of the embodiments, one embodiment of the present application provides a computer storage medium, which stores a computer program, which, when executed by a processor, implements the method according to any one of the embodiments.
[0073] It should be noted that according to implementation needs, in order to achieve the objectives of the embodiments of the present application, each component / step described in the embodiments of the present application can be divided into more components / steps, or two or more components / steps or some operations of the components / steps can be combined into a new component / step.
[0074] The method according to the embodiment of the present application can be implemented in hardware and firmware, or can be implemented as software or computer code that can be stored in a recording medium (e.g., CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or can be implemented by computer code downloaded from a network and stored in a remote recording medium or a non-transitory machine-readable medium, and computer code stored in a local recording medium. Thus, the method described herein can be processed by such software stored in a recording medium by using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (e.g., ASIC or FPGA). It should be understood that a computer, a processor, a microprocessor controller, or a programmable hardware includes a storage component (e.g., RAM, ROM, or flash memory) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method of detecting a target described herein is implemented. Furthermore, when a general-purpose computer accesses the code for implementing the method of detecting a target described herein, the execution of the code turns the general-purpose computer into a dedicated computer for performing the method of detecting a target described herein.
[0075] Furthermore, it should be noted that the terms "include," "comprise," or any variation of these terms, are intended to cover a non-exclusive inclusion. Thus, a process, method, article, or device that includes a set of elements may not only include such elements, but also include other elements not expressly identified, or may include inherent elements of the process, method, article, or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude other identical elements in the process, method, article, or device that includes the element.
[0076] Those skilled in the art may realize that the exemplary units and method steps described with reference to the embodiments disclosed herein may be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in the manner of hardware or software is related to the design constraints of specific applications and technical solutions. Those skilled in the art may use different methods to implement the described functions for each specific application, but this implementation should not be considered to go beyond the scope of the embodiments of the present application.
[0077] All the embodiments in this specification are described in a progressive manner, and the same or similar parts in the embodiments are referred to such embodiments, and the description of each embodiment focuses on the difference from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, and therefore will be briefly described, and the relevant parts may be referred to the partial description in the method embodiments.
[0078] The above implementations are merely used to describe the embodiments of the present application, and are not intended to limit the embodiments of the present application. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application, and the scope of patent protection of the embodiments of the present application should be defined by the claims. [Explanation of symbols]
[0079] 101 Acquiring a target image including a target object 102 performing instance segmentation on the target image to obtain a segmentation mask corresponding to the target object. 103. obtaining positional relationship features between target pixels in a target region where a target object is located in a target image based on a segmentation mask; 104. Obtaining positional relationship characteristics between standard pixels in a preset region of interest in a standard image containing a standard object corresponding to a target object. 105. To obtain a correspondence between the target pixels and the standard pixels, the positional relationship features between the target pixels and the positional relationship features between the standard pixels are matched, and pose information of the target object is obtained based on the correspondence. m1 and m2 target pixels N1 Normal vector corresponding to target pixel m1 N2 Normal vector corresponding to target pixel m2 F. Positional relationship feature between target pixel m1 and target pixel m2 F1 Distance between two target pixels F2 The angle between the normal vectors corresponding to the two target pixels F3 and F4 The angle between the normal vectors corresponding to the two target pixels and the connecting line between the two target pixels Convolution kernel of the previous convolutional layer in the H1 feature extraction network H2 Convolution kernels for later convolutional layers in feature extraction networks 60 Target detection device 601 Target Image Acquisition Module 602 Segmentation Mask Acquisition Module 603 First Positional Relationship Feature Acquisition Module 604 Second Positional Relationship Feature Acquisition Module 605 Pose Information Acquisition Module
Claims
1. 1. A method for detecting a target, comprising: acquiring a target image including a target object (101); performing instance segmentation on the target image to obtain a segmentation mask corresponding to the target object (102); obtaining (103) a positional relationship characteristic between target pixels in a target region where the target object is located in the target image based on the segmentation mask; Obtaining (104) standard pixel-to-pixel relationship features within a preset region of interest of a target object in a standard image; and and matching the positional relationship features between the target pixels and the standard pixels to obtain a correspondence between the target pixels and the standard pixels, and obtaining pose information of the target object based on the correspondence (105). method.
2. Obtaining (103) a positional relationship feature between target pixels in a target region where the target object is located in the target image based on the segmentation mask includes: pairwise combining the target pixels in the target region where the target object is located in the target image based on the segmentation mask to obtain a plurality of target pixel pairs, and for each target pixel pair, obtaining a positional relationship feature between two target pixels in the target pixel pair; Obtaining (104) standard pixel-to-pixel relationship features within a preset region of interest of a target object in a standard image includes: acquiring the standard image and the preset region of interest of a target object within the standard image; and 2. The method of claim 1, further comprising: combining the standard pixels in the preset region of interest in pairs to obtain a plurality of standard pixel pairs; and for each standard pixel pair, obtaining a positional relationship characteristic between two standard pixels in the standard pixel pair.
3. For each target pixel pair, the position relationship feature between the two target pixels in the target pixel pair is obtained based on a distance between the two target pixels, an angle between normal vectors corresponding to the two target pixels respectively, and an angle between the normal vectors corresponding to the two target pixels and a connecting line between the two target pixels; 3. The method of claim 2, wherein for each standard pixel pair, the positional relationship feature between the two standard pixels in the standard pixel pair is obtained based on a distance between the two standard pixels, an angle between normal vectors corresponding to the two standard pixels, respectively, and an angle between the normal vectors corresponding to the two standard pixels and a connecting line between the two standard pixels.
4. Performing instance segmentation on the target image to obtain a segmentation mask corresponding to the target object (102) includes:
2. The method of claim 1 , further comprising: performing instance segmentation on the target image by inputting the target image into a pre-trained instance segmentation model and using the instance segmentation model to obtain the segmentation mask corresponding to the target object.
5. The instance segmentation model includes: a feature extraction network, a feature fusion network, a region generation network, a feature alignment layer, a classification and regression network, and a segmentation mask network; Performing instance segmentation on the target image by inputting the target image into a pre-trained instance segmentation model and using the instance segmentation model to obtain the segmentation mask corresponding to the target object, includes: performing multi-scale feature extraction on the target image by inputting the target image into the feature extraction network in the pre-trained instance segmentation model and using the feature extraction network to obtain a multi-level initial feature map corresponding to the target image; performing feature fusion on the multiple-level initial feature maps by using the feature fusion network to obtain a fused feature map; obtaining information about an initial region of the target object by using the region generation network based on the fused feature map; performing feature extraction on the initial feature map by using the feature alignment layer based on the information about the initial region to obtain a region feature map in the initial feature map corresponding to the initial region; 5. The method of claim 4, further comprising: obtaining category information and location information of the target object by using the classification and regression network based on the region feature map; and obtaining the segmentation mask corresponding to the target object by using the segmentation mask network based on the region feature map.
6. Performing feature fusion on the multiple-level initial feature maps by using the feature fusion network to obtain a fused feature map includes: performing a convolution operation on each of the initial feature maps by using the feature fusion network to obtain multiple levels of initial dimensionality-reduced feature maps; Sequentially performing a fusion process for every two adjacent levels of the initial dimensionality reduced feature maps according to descending order of levels to obtain an initial fused feature map, and using the initial fused feature map to update the initial dimensionality reduced feature map of the lower level of the adjacent level, where the size of the initial dimensionality reduced feature map of the higher level is smaller than the size of the initial dimensionality reduced feature map of the lower level; performing the convolution operation on each of the initial fused feature maps to obtain a multi-level dimensionality reduced feature map; and 6. The method of claim 5, further comprising: sequentially performing a fusion process for every two adjacent levels of the dimensionality reduced feature maps according to an ascending order of levels to obtain a transition feature map; performing a fusion process on the transition feature map and the initial feature map to obtain a fused feature map; and using the fused feature map to update the dimensionality reduced feature map of the upper level of the adjacent level, in which the size of the dimensionality reduced feature map of the upper level is smaller than the size of the dimensionality reduced feature map of the lower level.
7. 6. The method of claim 5, wherein the feature extraction network includes two concatenated convolutional layers, a size of a convolutional kernel of a previous convolutional layer is 1*1, a convolutional stride of the previous convolutional layer is 1, and a convolutional stride of a subsequent convolutional layer is less than or equal to a size of a convolutional kernel of the subsequent convolutional layer.
8. 1. An apparatus for detecting a target, comprising: A target image acquisition module (601) configured to acquire a target image including a target object; a segmentation mask acquisition module (602) configured to perform instance segmentation on the target image to obtain a segmentation mask corresponding to the target object; a first positional relationship feature acquisition module (603) configured to acquire positional relationship features between target pixels in a target region in which the target object is located in the target image based on the segmentation mask; a second positional relationship feature acquisition module (604) configured to acquire positional relationship features between standard pixels within a preset region of interest of the target object in the standard image; a pose information acquisition module (605) configured to match the positional relationship features between the target pixels and the positional relationship features between the standard pixels to acquire a correspondence relationship between the target pixels and the standard pixels, and acquire pose information of the target object based on the correspondence relationship. Device.
9. An electronic device including a processor, a memory, a communication interface, and a communication bus, The processor, the memory and the communication interface complete communication with each other by using the communication bus; The memory is configured to store at least one executable instruction, the executable instruction causing the processor to implement a method for detecting a target according to any one of claims 1 to 7. Electronic devices.
10. A computer storage medium for storing a computer program, The computer program, when executed by a processor, implements a method for detecting a target according to any one of claims 1 to 7. Computer storage media.