Affordability Detection Method and Related Device
By using multiple feature extraction branches and semantic coding layers in affordability detection, the weight of the feature map is determined and fused, the problem of poor distinction between information loss and irrelevant information in the prior art is solved, and the effect of image segmentation is improved.
Patent Information
- Application Number
- CN202210486904.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-05-06
AI Technical Summary
The prior art is prone to lose information of fine structures and information of small objects in affordability detection, and fails to effectively distinguish irrelevant information, resulting in unsatisfactory identification results.
Multiple feature extraction branches obtain multiple sets of feature maps of the image to be identified, use deep feature maps to determine the weight of the feature map, and fuse multiple sets of feature maps with reference feature maps into enhanced feature maps, and finally decode them to obtain affordability detection results.
By carrying the weight of deep feature information, the information conducive to image segmentation is enhanced, and irrelevant information is suppressed, thereby improving the segmentation effect of the recognized image.
Smart Images

Figure CN114863105B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image segmentation. Specifically, it relates to an affordance detection method and related devices. Background Art
[0002] Affordance reflects the functional possibilities presented by an object in an environment. Specifically, it not only requires identifying the object from the image to be recognized but also determining the function of each part of the object. Therefore, affordance detection requires pixel-level segmentation and recognition of different functional regions of the object.
[0003] In related technologies, there is an inventive concept of downsampling the high-resolution feature map of the image to be recognized to a low resolution and then restoring it to a high resolution. However, this concept connects features of different resolutions in a cascaded manner, resulting in a relatively small resolution of the deep feature map, which is prone to losing information about fine structures and small objects.
[0004] Therefore, in other related technologies, it is proposed to fuse the high-resolution feature map of the shallow layer of the image to be recognized with the low-resolution feature map of the deep layer and then perform image segmentation based on the fused feature map. It is found that this method does not distinguish irrelevant information for affordance detection, resulting in sometimes less than ideal recognition results for the image to be recognized. Summary of the Invention
[0005] To overcome at least one deficiency in the prior art, this application provides an affordance detection method that can obtain better affordance detection results when performing affordance detection on an image, including:
[0006] In a first aspect, an embodiment of this application provides an affordance detection method applied to an affordance detection device. The affordance detection device is configured with a pre-trained affordance detection model, and the affordance detection model includes multiple feature extraction branches and a semantic encoding layer. The method includes:
[0007] Obtaining multiple groups of feature maps of the image to be recognized through the multiple feature extraction branches, where the multiple groups of feature maps carry feature information of the image to be recognized from shallow to deep;
[0008] Inputting the deep feature maps in the multiple groups of feature maps into the semantic encoding layer to obtain the weights of each of the multiple groups of feature maps;
[0009] Fusing the multiple groups of feature maps with a reference feature map of the image to be recognized into an enhanced feature map according to the weights of each of the multiple groups of feature maps; where the reference feature map carries spatial structure information of the image to be recognized;
[0010] Decode the enhanced feature map to obtain the affordance detection result of the image to be recognized.
[0011] In a second aspect, an embodiment of the present application provides an affordance detection device, which is applied to an affordance detection device configured with a pre-trained affordance detection model. The affordance detection model includes multiple feature extraction branches and a semantic encoding layer. The affordance detection device includes:
[0012] An image encoding module, configured to obtain multiple groups of feature maps of the image to be recognized through the multiple feature extraction branches, where the multiple groups of feature maps carry the feature information of the image to be recognized from shallow to deep;
[0013] The image encoding module is further configured to input the deep feature maps in the multiple groups of feature maps into the semantic encoding layer to obtain the weights of the multiple groups of feature maps respectively;
[0014] An image decoding module, configured to fuse the multiple groups of feature maps with the reference feature map of the image to be recognized into an enhanced feature map according to the weights of the multiple groups of feature maps respectively; where the reference feature map carries the spatial structure information of the image to be recognized;
[0015] The image decoding module is further configured to decode the enhanced feature map to obtain the affordance detection result of the image to be recognized.
[0016] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the above-mentioned affordance detection method.
[0017] In a fourth aspect, an embodiment of the present application provides an affordance detection device, which includes a processor and a memory. The memory stores a computer program, which when executed by the processor, implements the above-mentioned affordance detection method.
[0018] Compared with the prior art, the present application has the following beneficial effects:
[0019] In the affordance detection method and related devices provided in this embodiment, the affordance detection device obtains multiple groups of feature maps carrying the feature information of the image to be recognized from shallow to deep; then, it uses the target feature map carrying the deep feature information to determine the weights of each group of feature maps, and fuses each group of feature maps with the reference feature map into an enhanced feature map according to the weights of each group of feature maps; finally, it decodes the enhanced feature map to obtain the affordance detection result of the image to be recognized. Since the deep feature information has richer semantic information and is suitable for global semantic encoding, the weights determined by the target feature map carrying the deep feature information can enhance the information beneficial to image segmentation and suppress the information irrelevant to image segmentation, so as to achieve the purpose of improving the segmentation effect of the image to be recognized. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is an example diagram of the image to be recognized provided in the embodiment of the present application;
[0022] Figure 2 It is a schematic diagram of the traditional image segmentation effect provided in the embodiment of the present application;
[0023] Figure 3 It is a schematic diagram of the image segmentation effect of the affordance detection provided in the embodiment of the present application;
[0024] Figure 4 It is a schematic diagram of the structure of the affordance detection device provided in the embodiment of the present application;
[0025] Figure 5 It is a schematic flowchart of the affordance detection method provided in the embodiment of the present application;
[0026] Figure 6 It is one of the schematic diagrams of the structure of the affordance detection model provided in the embodiment of the present application;
[0027] Figure 7 It is another schematic diagram of the structure of the affordance detection model provided in the embodiment of the present application;
[0028] Figure 8 It is a schematic diagram of the structure of the affordance detection device provided in the embodiment of the present application.
[0029] Icons: 10 - Hand hammer; 101 - Hammer head; 102 - Hammer handle; 120 - Memory; 130 - Processor; 140 - Communication unit; 201 - Image encoding module; 202 - Image decoding module. Detailed implementation manners
[0030] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. Components of the embodiments of the present application generally described and illustrated in the figures herein may be arranged and designed in a variety of different configurations.
[0031] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but is merely representative of selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0032] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it need not be further defined and explained in subsequent figures.
[0033] In the description of the present application, it should be noted that the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0034] This embodiment relates to affordance detection. To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, first, in combination with Figure 1 the meaning of affordance detection will be described.
[0035] As Figure 1 shown in the image to be recognized, the target object included is a hand hammer 10; in traditional object detection methods, only semantic segmentation of the target object in the image to be recognized needs to be performed as a whole, and the segmentation result can be as Figure 2 shown.
[0036] In some application scenarios, it is not only necessary to segment the target object from the image to be recognized, but also necessary to further segment the target object to determine each part of the target object. For example, Figure 3 In the example shown, when it is necessary to control the end of the robotic arm to hold the hammer 10 to perform some tasks, it is necessary to perform affordance detection on the hammer 10, distinguish the hammer head 101 and the hammer handle 102 of the hammer 10 from the image to be recognized, and the segmentation result can be as Figure 3 shown.
[0037] Therefore, in order to perform affordance detection on the image to be recognized, in the related art, the high-resolution feature map of the image to be recognized is downsampled to a low resolution, and then restored from the low-resolution feature map to a high resolution to perform affordance detection on the image to be recognized, which is likely to lose the information of fine structures and small objects. Therefore, in some other related technologies, it is proposed to fuse the high-resolution feature image of the shallow layer of the image to be recognized with the low-resolution feature map of the deep layer, and then perform image segmentation based on the fused feature map to determine each part of the target object in the image to be recognized.
[0038] However, although this method overcomes the problem that when the high-resolution feature map of the image to be recognized is downsampled to a low resolution and then restored from the low-resolution feature map to a high resolution, it is likely to lose the information of fine structures and small objects; however, this method does not distinguish the irrelevant information of affordance detection, resulting in sometimes less than ideal recognition results for the image to be recognized.
[0039] It should be noted that based on the discovery of the above technical problems, the inventor has proposed the following technical solutions through creative labor to solve or improve the above problems. It should be noted that the defects existing in the above solutions in the prior art are all the results obtained by the inventor after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the embodiments of the present application below for the above problems should all be the contributions made by the inventor to the present application during the invention creation process, and should not be understood as the technical content known to those skilled in the art.
[0040] In view of the above technical problems, this embodiment provides an affordance detection method applied to an affordance detection device. In this method, the affordance detection device acquires multiple groups of feature maps carrying the feature information of the image to be recognized from shallow to deep; then, uses the target feature map carrying the deep feature information to determine the weights of each group of feature maps, and fuses each group of feature maps with a reference feature map into an enhanced feature map according to the weights of each group of feature maps; finally, decodes the enhanced feature map to obtain the affordance detection result of the image to be recognized. Since the weights of each group of feature maps can enhance the information beneficial to image segmentation and suppress the information irrelevant to image segmentation; therefore, the segmentation effect of the image to be recognized is improved.
[0041] In some embodiments, the affordance detection device may be a server. Among them, the server may be a single server or a server group. The server group may be centralized or distributed (for example, the server may be a distributed system). In some embodiments, the server may be local or remote relative to the user terminal. In some embodiments, the server may be implemented on a cloud platform; by way of example only, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, multi-cloud, etc., or any combination thereof. In some embodiments, the server may be implemented on an electronic device having one or more components.
[0042] Of course, the affordance detection device may also be a user terminal. Among them, the user terminal may include a mobile terminal, a tablet computer, a laptop computer, or a built-in device in a motor vehicle, etc., or any combination thereof. In some embodiments, the mobile terminal may include a smart home device, a wearable device, a smart mobile device, a virtual reality device, or an augmented reality device, etc., or any combination thereof. In some embodiments, the smart home device may include a smart lighting device, a control device for smart electrical appliances, a smart monitoring device, a smart TV, a smart camera, or an intercom, etc., or any combination thereof. In some embodiments, the wearable device may include a smart bracelet, smart shoelaces, smart glasses, a smart helmet, a smart watch, smart clothing, a smart backpack, smart accessories, etc., or any combination thereof. In some embodiments, the smart mobile device may include a smart phone, a personal digital assistant (PDA), a gaming device, a navigation device, or a point of sale (POS) device, etc., or any combination thereof.
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, this embodiment also provides a structural schematic diagram of the affordance detection device. AsFigure 4 As shown, the affordance detection device may include a memory 120, a processor 130, and a communication unit 140. Each of the memory 120, the processor 130, and the communication unit 140 is electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more communication buses or signal lines.
[0044] Among them, the memory 120 may be an information recording device based on any electronic, magnetic, optical, or other physical principles for recording execution instructions, data, etc. In some embodiments, the memory 120 may be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, etc.
[0045] Among them, by way of example only, the volatile memory may be a random access memory (RAM). The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, etc.; the storage drive may be a disk drive, a solid-state drive, any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof, etc.
[0046] The communication unit 140 is configured to send and receive data via a network. In some embodiments, the network may include a wired network, a wireless network, an optical fiber network, a telecommunication network, an intranet, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Network (WLAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a Public Switched Telephone Network (PSTN), a Bluetooth network, a ZigBee network, or a Near Field Communication (NFC) network, etc., or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include a wired or wireless network access point, such as a base station and / or a network switching node, and one or more components of the service request processing system may be connected to the network via the access point to exchange data and / or information.
[0047] The processor 130 may be an integrated circuit chip with signal processing capabilities, and the processor may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the above-mentioned processor may include a Central Processing Unit (CPU), an Application-Specific Integrated Circuit (ASIC), an Application-Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC), or a microprocessor, etc., or any combination thereof.
[0048] Based on the relevant descriptions in the above embodiments, the following combinesFigure 5 The flowchart shown is a detailed description of the availability detection method provided by the present embodiment. However, it should be understood that the operations of the flowchart may not be implemented in order, and steps without logical contextual relationships may be reversed in order or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart, or may remove one or more operations from the flowchart under the guidance of the content of this application.
[0049] The availability detection device in this embodiment is configured with a pre-trained availability detection model, which includes multiple feature extraction branches and a semantic encoding layer. Based on the availability detection model, Figure 5 As shown, the method includes:
[0050] S101, obtaining multiple groups of feature maps of the image to be identified through multiple feature extraction branches.
[0051] Among them, multiple sets of feature maps carry feature information of the image to be identified from shallow to deep layers. It should be understood that for the same convolutional neural network, the shallow feature information in this embodiment represents the feature information output by the network layer close to the input layer; the deep feature information represents the feature information output by the network layer far from the input layer.
[0052] It should also be understood that the existing end-to-end availability detection model usually includes an encoder and a decoder, wherein the encoder includes a plurality of convolutional layers connected in sequence, for extracting feature information from the image to be identified and converting it into a feature map; and the decoder includes a plurality of deconvolutional layers connected in sequence, for segmenting the target object in the image to be identified according to the feature information in the feature map.
[0053] Similar to the existing end-to-end availability detection model, the present embodiment also includes an encoder, but different from the existing availability detection model, in the availability detection model of the present embodiment, the encoder is composed of multiple parallel feature extraction branches, and each feature extraction branch includes multiple convolutional layers connected in sequence.
[0054] In addition, this embodiment takes into account that the residual network has many excellent characteristics when performing image recognition processing. Therefore, this embodiment introduces the residual network into the traditional end-to-end availability detection model to improve the performance of the traditional availability detection model. Therefore, the availability detection model in this embodiment also includes a residual network layer, and the residual network layer includes a plurality of residual units corresponding to a plurality of feature extraction branches. Based on the multiple feature extraction branches and the residual network layer in the availability detection model, step S101 can obtain multiple sets of feature maps of the image to be identified by the following implementation method:
[0055] S101-1, obtaining an image to be recognized.
[0056] S101-2. Input the image to be recognized into the residual network layer to obtain multiple groups of initial feature maps output by multiple residual units.
[0057] Exemplarily, as Figure 6 shown, the affordance detection model is functionally divided into an encoder and a decoder. The encoder in this model includes 4 branches, respectively denoted as b1, b2, b3, and b4. Among them, branch b1 represents a direct connection branch for transmitting reference feature maps, and b2, b3, and b4 represent multiple feature extraction branches; after the input initial feature maps are subjected to feature extraction, 3 groups of feature maps of the image to be recognized are output. Among them, after the affordance detection model is trained, the 3 feature extraction branches can respectively extract different feature information with emphasis.
[0058] As introduced in the above embodiments, the residual network layer includes multiple residual units corresponding one by one to multiple feature extraction branches. Therefore, as Figure 7 shown, the input end of each feature extraction branch in the figure is connected to the output end of the corresponding residual unit. The affordance detection device inputs the image to be recognized into the residual network layer. The output result of the first residual unit serves as the reference feature map of the direct connection channel b1, and the output result of the second residual unit serves as the initial feature map of the feature extraction branch b2; the output result of the third residual unit serves as the initial feature map of the feature extraction branch b3; the output result of the fourth residual unit serves as the initial feature map of the feature extraction branch b4.
[0059] In order to enable those skilled in the art to use the content of this application, a possible implementation manner of the residual network layer is given below. For those skilled in the art, without departing from the spirit and scope of this application, on the basis of the basic embodiment, those skilled in the art can adaptively adjust the structural parameters of the residual network layer to adapt to other embodiments and application scenarios.
[0060] The residual network layer in this embodiment can be directly modified based on the Resnet101 network. Specifically, it is residual learning among three layers, and the three convolutional kernels are 1×1, 3×3, and 1×1 respectively; and there are a total of 4 residual units, corresponding to 1 direct connection channel and 3 feature extraction branches. Among them, the first residual unit contains 3 residual blocks and has 3 3×3 convolutional layers; the second residual unit contains 4 residual blocks and has 4 3×3 convolutional layers; the third residual unit contains 6 residual blocks and has 6 3×3 convolutional layers; the fourth residual unit contains 3 residual blocks and has 3 3×3 convolutional layers.
[0061] Continue to refer to Figure 7, Based on this residual network, the affordance detection device first uses a 3×3 convolution to extract features from the image to be recognized, obtaining a reference feature map with 64 channels. Compared with the image to be recognized, the size of the feature map is halved.
[0062] Each feature extraction branch is used to extract features from the initial feature map output by the corresponding residual unit, so that the three feature extraction branches respectively output feature maps with different dimensions. It should be noted that in this embodiment, the feature maps output by each network layer in the same feature extraction branch maintain the same feature size.
[0063] Continue to refer to Figure 7 , In this affordance detection model, from the previous feature extraction branch to the next feature extraction branch, a downsampling with a stride of 2 is used to achieve transitional connection. Among them, downsampling is used to reduce the spatial size of the initial feature map of the next feature extraction branch and improve the computational efficiency of the network. Specifically, compared with b2, the size of the initial feature map of b3 is halved, and the number of channels is twice that of the initial feature map of b2; compared with b3, the size of the initial feature map of b4 is halved, and the number of channels is twice that of the initial feature map of b3.
[0064] In this way, in this embodiment, the residual network layer is introduced into the affordance detection model, and by virtue of the characteristics of the residual network, the segmentation performance of the entire affordance detection model is improved.
[0065] S101-3, Input multiple groups of initial feature maps into multiple feature extraction branches according to the corresponding relationship to obtain multiple groups of feature maps of the image to be recognized.
[0066] Furthermore, in order to further deepen the depth of feature fusion between feature extraction branches in this embodiment, during the feature extraction process of each feature extraction branch, the features between branches are fused.
[0067] Continue to refer to Figure 6 , The affordance detection device inputs three groups of initial feature maps into feature extraction branches b2, b3, and b4 respectively; for each feature extraction branch, the following method is used to fuse features with the remaining feature extraction branches to obtain the corresponding feature maps:
[0068]
[0069]
[0070] In the formula, y i represents a group of feature maps output by feature extraction branch b i , and f ij (x j ) represents a group of feature maps x to be fused output by feature extraction branch b j to be fused. jBefore fusing with a set of feature maps to be fused output by the feature extraction branch b i The sampling process that needs to be performed before fusing with a set of feature maps to be fused output by the feature extraction branch b. 2×(i - j)times↓ indicates that x j needs to be downsampled by a factor of 2×(i - j); 2×(j - i)times↑ indicates that x j needs to be upsampled by a factor of 2×(j - i). That is to say, the symbol ↓ represents the downsampling operation, and the symbol ↑ represents the upsampling operation.
[0071] Continue to refer to Figure 6 , for the feature extraction branch b2, after the availability detection device upsamples the feature maps to be fused output by the feature extraction branch b3 and the feature maps to be fused output by the feature extraction branch b4, and then adds them pixel by pixel to the feature maps to be fused output by the feature extraction branch b2, a set of feature maps corresponding to the extraction branch b2 is obtained;
[0072] For the feature extraction branch b3, after the availability detection device downsamples the feature maps to be fused output by the feature extraction branch b2 and upsamples the feature maps to be fused output by the feature extraction branch b4, and then adds them pixel by pixel to the feature maps to be fused output by the feature extraction branch b3, a set of feature maps corresponding to the extraction branch b3 is obtained;
[0073] For the feature extraction branch b4, after the availability detection device downsamples the feature maps to be fused output by the feature extraction branch b2 and the feature extraction branch b3, and then adds them pixel by pixel to the feature maps to be fused output by the feature extraction branch b4, a set of feature maps corresponding to the extraction branch b4 is obtained.
[0074] Exemplarily, assume that the size of the image to be recognized is [W, H, 3], where the sizes of the feature maps to be fused corresponding to the three feature extraction branches b2, b3, and b4 are respectively Therefore, when performing feature fusion between the three feature extraction branches, 3×3 convolution can be used to downsample the high - resolution features, and the nearest - neighbor method can be used to upsample the low - resolution features. After transforming the features of different branches into the same resolution and number of channels, pixel - by - pixel addition is used for fusion, so that the features of each branch are fused with the features of other branches.
[0075] Refer to again Figure 5 , after step S101, the availability detection method further includes:
[0076] S102, inputting the deep - layer feature maps in the multiple sets of feature maps into the semantic encoding layer to obtain the weights of each set of feature maps.
[0077] It should be understood that the deep feature information has richer semantic information and is suitable for encoding the global semantics, so as to explore the relationships between various types of features in the image to be recognized at the global level. Therefore, in this embodiment, the deep feature map carrying the deep feature information is selected to determine the weights of each group of feature maps.
[0078] Exemplarily, continue to refer to Figure 6 the three feature extraction branches shown. In the order of the features extracted from the shallow layer to the deep layer, they are b2, b3, and b4 in sequence. Therefore, the features extracted by b4 are the deepest and carry the most semantic information. Then, the feature map output by branch b4 can be used as the deep feature map and input into the semantic encoding layer to obtain the weights of each group of feature maps.
[0079] In an optional implementation manner, the semantic encoding layer in this embodiment can be implemented based on NetVLAD. Therefore, step S102 may include the following implementation manners:
[0080] S102-1, encoding the deep feature map by using NetVLAD through the semantic encoding layer to obtain the global feature.
[0081] S102-2, converting the global feature into the weight coefficients of each group of feature maps in the following manner:
[0082]
[0083] In the formula, e includes the weight coefficients of each group of feature maps, represents the normalization result of V, V represents the global feature, represents converting into a c×1×1 vector through the fully connected layer. c corresponds to the number of groups of feature maps, and σ represents the sigmoid function.
[0084] It should be noted that NetVLAD can effectively capture the overall features of pictures in image retrieval. Therefore, it is usually used in the field of image retrieval. The inventor's research finds that the deep feature information output by multiple feature extraction branches in this embodiment is similar to the overall features of pictures captured during image retrieval and also has rich semantic information. Therefore, it is used as the semantic encoding layer in this embodiment to explore the relationships between various types of features in the image to be recognized at the global level, so as to determine the weights of each group of feature maps.
[0085] The principle of NetVLAD is that it is assumed that the dictionary of an image is represented as D = {d1, d2…, d n}, which contains a total of n coding words, representing n cluster centers. The learning process of the dictionary can be obtained through training in the network, where the number n of coding words can be set artificially. According to the set n, the corresponding number of coding words can be learned through backpropagation in an end-to-end manner during training. Therefore, the value of n can be adjusted in different datasets.
[0086] In the specific implementation, it is assumed that the input feature is represented as a three-dimensional matrix of c×h×w, where c represents the number of channels, and h and w represent the height and width of the feature map respectively; then, it is regarded as m local descriptors x of dimension c i , and x i and the coding words are used for residual calculation, and the residual result is expressed as:
[0087] r i,j = x i - d j
[0088] Based on this residual expression, a soft assignment method can be used to determine the weight of each channel in the deep feature, and the corresponding calculation formula is:
[0089]
[0090] In this way, the weight is obtained according to the distance from each local feature to the cluster center. Specifically, the closer x i is to d j , the closer the weight is to 1, and vice versa, the closer it is to 0. The global feature corresponding to the local feature is expressed as:
[0091]
[0092] where V(i,j) is an n×c vector, reflecting the residual distribution of the local feature in n coding classifications.
[0093] After normalizing the obtained V(i,j), it is input to the fully connected layer for processing to convert it into a c×1×1 vector; and it is mapped by the sigmoid function to a weight coefficient value e between 0 and 1 as the weight of each channel of the deep feature. For example, for Figure 6 the three groups of feature maps output by the three feature extraction branches in, the value of c in c×1×1 needs to be 3.
[0094] Continue to refer to Figure 5 , after step S102, the affordance detection method further includes:
[0095] S103, according to the weights of each group of feature maps, fuse each group of feature maps with the reference feature map of the image to be recognized into an enhanced feature map.
[0096] Among them, the reference feature map carries the spatial structure information of the image to be recognized. In addition, through research, it is found that when extracting features from the image to be recognized by a convolutional neural network, the shallow feature maps carry rich spatial structure information of the image to be recognized (for example, the edge contour of the target object in the image to be recognized and the size information of the image to be recognized, etc.); therefore, in order to further improve the segmentation effect of the image to be recognized, in this embodiment, the feature map obtained by performing feature extraction on the image to be recognized once is used as the reference feature map to correct the edge contour information in the fused feature map. Therefore, step S103 may include the following implementation manners:
[0097] S103-1, according to the respective weights of multiple groups of feature maps, fuse the multiple groups of feature maps into a fused feature map in the following manner:
[0098]
[0099] In the formula, represents the fused feature map, F includes multiple groups of feature maps, represents performing weighted summation on multiple groups of feature maps and their respective weight coefficients;
[0100] Assume that the multiple groups of feature maps in F are respectively represented as f1, f2, f3…, f n , then the relationship between the fused feature map and the multiple groups of feature maps can be expressed as:
[0101]
[0102] Among them, a n f n in a n represents the weight of the nth group of feature maps f n . In some implementation manners, a group of feature maps may include multiple channels. Therefore, when the feature map includes multiple channels, the feature maps of multiple channels can be fused into one feature map and then weighted summation is performed.
[0103] In addition, since the feature dimensions of shallow features and deep features are different, before fusing multiple groups of feature maps, it is necessary to adjust multiple groups of feature maps to the same scale. Considering that the feature maps with deep feature information lose a lot of spatial structure information, in this embodiment, the feature map with shallow feature information is selected as the size standard to perform upsampling on other feature maps. In this way, the fused feature map not only carries shallow spatial structure information but also carries deep semantic information.
[0104] For example, it can be Figure 6The feature map output by the feature extraction branch b2 in [ ] is used as the size standard, and the feature maps output by b3 and b4 are upsampled so that the feature maps output by b3 and b4 have the same size as the feature map output by b2, and then they are fused.
[0105] S103-2. Fuse the fused feature with the reference feature map to obtain an enhanced feature map.
[0106] Regarding this reference feature map, you can continue to refer to Figure 6 , in Figure 6 In addition to the 3 feature extraction branches shown, the affordance detection model also includes a direct connection channel. Based on this direct connection channel, in the decoder, the reference feature map extracted from the image to be recognized through one-time feature extraction is fused with the feature maps of the 3 feature extraction branches in the encoder; an enhanced feature map corrected by the reference feature map is obtained; finally, the decoder performs a decoding operation on the enhanced feature map to convert it into the segmentation result of the image to be recognized.
[0107] Among them, the reference feature map and the fused feature map can be fused in various ways, and those skilled in the art can make an adaptive selection according to the needs of the implementation scenario. In some embodiments, the two can be added at the feature level. In other embodiments, the two can be added at the channel level.
[0108] Taking the addition at the channel level as an example, assuming that the fused feature map has 3 channels and the reference feature map has 1 channel, if added at the channel level, the obtained enhanced feature map has 4 channels.
[0109] Refer to Figure 5 again. After step S103, the affordance detection method further includes:
[0110] S104. Decode the enhanced feature map to obtain the affordance detection result of the image to be recognized.
[0111] In this way, based on the above embodiments, the affordance detection device obtains multiple groups of feature maps carrying the feature information of the image to be recognized from shallow to deep; then, uses the target feature map carrying the deep feature information to determine the weights of each group of feature maps, and fuses each group of feature maps with the reference feature map into an enhanced feature map according to the weights of each group of feature maps; finally, decodes the enhanced feature map to obtain the affordance detection result of the image to be recognized. Since the deep feature information has richer semantic information and is suitable for global semantic encoding, the weights determined by the target feature map carrying the deep feature information can enhance the information beneficial to image segmentation and suppress the information irrelevant to image segmentation, so as to achieve the purpose of improving the segmentation effect of the image to be recognized.
[0112] In addition, since the availability detection model in this embodiment is implemented based on the principle of artificial neural network, it is necessary to train the availability detection model with the same structure through sample images to become the availability detection model. Therefore, in some embodiments, an image in RGB format can be collected, and the availability category of each pixel can be annotated using an annotation tool to obtain a sample image; then, the sample image is divided into a training set and a test set for training and verifying the availability detection model. In other embodiments, sample images can also be obtained from an existing public data set for training. For example, the public data set can be IIT-AFF. Regardless of the means used, the sample image can also be preprocessed with some images, such as scaling, rotation, flipping, scaling and random center cropping, as well as mean normalization.
[0113] Based on the same inventive concept as the availability detection method provided in this embodiment, this embodiment also provides an availability detection device, which includes at least one software function module that can be stored in the memory 120 in software form or fixed in the operating system (OS) of the availability detection device. The availability detection device is configured with a pre-trained availability detection model, and the availability detection model includes multiple feature extraction branches and a semantic encoding layer. Based on the availability detection model, please refer to Figure 8 , from the functional point of view, the affordance detection device can include:
[0114] The image encoding module 201 is used to obtain multiple groups of feature maps of the image to be identified through multiple feature extraction branches, wherein the multiple groups of feature maps carry feature information of the image to be identified from shallow to deep layers.
[0115] The image encoding module 201 is further used to input the deep feature maps in the multiple groups of feature maps into the semantic encoding layer to obtain the weights of the multiple groups of feature maps.
[0116] In this embodiment, the image encoding module 201 is used to implement Figure 5 For the detailed description of the image encoding module 201, please refer to the detailed description of steps S101-S102.
[0117] The image decoding module 202 is used to fuse the multiple groups of feature maps with the reference feature map of the image to be identified into an enhanced feature map according to the respective weights of the multiple groups of feature maps; wherein the reference feature map carries the spatial structure information of the image to be identified.
[0118] The image decoding module 202 is also used to decode the enhanced feature map to obtain the availability detection result of the image to be identified.
[0119] In this embodiment, the image decoding module 202 is used to implement Figure 5 Steps S103 - S104 in
[0120] In addition, in each embodiment of the present application, each functional module can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part. Moreover, in some embodiments, the availability detection device may further include other software functional modules for implementing other steps or sub - steps of the availability detection method provided in this embodiment.
[0121] It should also be understood that if the above - mentioned embodiments are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer - readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.
[0122] Therefore, this embodiment also provides a computer - readable storage medium. This computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the availability detection method provided in this embodiment. Among them, the computer - readable storage medium can be various media such as a USB flash drive, a mobile hard disk, a read - only memory (ROM, Read - Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0123] An availability detection device provided in this embodiment. The availability detection device may include a processor 130 and a memory 120. The processor 130 and the memory 120 can communicate via a system bus. Moreover, the memory 120 stores a computer program, and the processor realizes the availability detection method provided in this embodiment by reading and executing the computer program corresponding to the above - mentioned embodiment in the memory 120.
[0124] It should be understood that the devices and methods disclosed in the above embodiments can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0125] As described above, these are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An affordance detection method, characterized in that, Applied to an affordance detection device, the affordance detection device being configured with a pre-trained affordance detection model, the affordance detection model including multiple feature extraction branches and a semantic encoding layer, the method includes: Obtaining multiple groups of feature maps of the image to be recognized through the multiple feature extraction branches, where the multiple groups of feature maps carry the feature information of the image to be recognized from shallow to deep; Inputting the deep feature maps in the multiple groups of feature maps into the semantic encoding layer to obtain the weights of the multiple groups of feature maps respectively, including: Encoding the deep feature maps by the semantic encoding layer in the way of NetVLAD to obtain global features; Converting the global features into the weight coefficients of the multiple groups of feature maps respectively in the following way: where e includes the weight coefficients of the respective groups of feature maps, represents the normalization result of V, and V represents the global feature, represents passing through a fully connected layer to convert to a c×1×1 vector, c corresponds to the number of the groups of feature maps, and σ represents the sigmoid function; Fusing the multiple groups of feature maps with the reference feature map of the image to be recognized into an enhanced feature map according to the weights of the multiple groups of feature maps respectively; where the reference feature map carries the spatial structure information of the image to be recognized; Decoding the enhanced feature map to obtain the affordance detection result of the image to be recognized.
2. The availability detection method according to claim 1, characterized in that The fusing the multiple groups of feature maps with the reference feature map of the image to be recognized into an enhanced feature map according to the weights of the multiple groups of feature maps respectively includes: Fusing the multiple groups of feature maps into a fused feature map in the following way according to the weights of the multiple groups of feature maps respectively: In the formula, represents the fused feature map, and F includes the multiple groups of feature maps. represents the weighted summation of the multiple groups of feature maps and their respective weight coefficients. Fusing the fused feature with the reference feature map to obtain the enhanced feature map.
3. The affordance detection method according to claim 1, wherein The affordance detection model further includes a residual network layer, the residual network layer including multiple residual units corresponding one by one to the multiple feature extraction branches, and the obtaining multiple groups of feature maps of the image to be recognized through the multiple feature extraction branches includes: Obtaining the image to be recognized; Inputting the image to be recognized into the residual network layer to obtain multiple groups of initial feature maps output by the multiple residual units; Inputting the multiple groups of initial feature maps into the multiple feature extraction branches according to the corresponding relationship between the multiple residual units and the multiple feature extraction branches to obtain multiple groups of feature maps of the image to be recognized.
4. The affordance detection method according to claim 3, characterized in that The affordance detection model includes 4 branches, respectively denoted as b1, b2, b3, b4, where branch b1 represents a direct connection branch for transmitting the reference feature map, and b2, b3, b4 represent the multiple feature extraction branches; The inputting the multiple groups of initial feature maps into the multiple feature extraction branches according to the corresponding relationship to obtain multiple groups of feature maps of the image to be recognized includes: Inputting 3 groups of initial feature maps into feature extraction branches b2, b3, b4 respectively; For each feature extraction branch, performing feature fusion with the remaining feature extraction branches in the following way to obtain the corresponding feature map: Where y i represents a set of feature maps output by the feature extraction branch b i f ij (x j ) represents the sampling process that needs to be performed before fusing a set of feature maps x j to be fused output by the feature extraction branch b j with a set of feature maps to be fused output by the feature extraction branch b i 2×(i - j)times↓ indicates that x j needs to be downsampled; 2×(j - i)times↑ indicates that x j needs to be upsampled.
5. An affordance detection device, characterized in that, Applied to an affordance detection device, the affordance detection device being configured with a pre-trained affordance detection model, the affordance detection model including multiple feature extraction branches and a semantic encoding layer, the affordance detection device includes: An image encoding module, configured to obtain multiple groups of feature maps of an image to be recognized through the multiple feature extraction branches, wherein the multiple groups of feature maps carry the feature information of the image to be recognized from shallow to deep; The image encoding module is further configured to input the deep feature maps in the multiple groups of feature maps into the semantic encoding layer to obtain the weights of the multiple groups of feature maps respectively. Specifically, it is configured to: Encode the deep feature maps by the semantic encoding layer in the NetVLAD manner to obtain global features; Convert the global features into the weight coefficients of the multiple groups of feature maps respectively in the following manner: where e includes the weight coefficients of the respective multiple groups of feature maps, represents the normalized result of V, and V represents the global feature, represents passing through the fully connected layer to convert to a c×1×1 vector, c corresponds to the number of the multiple groups of feature maps, and σ represents the sigmoid function; An image decoding module, configured to fuse the multiple groups of feature maps with a reference feature map of the image to be recognized into an enhanced feature map according to the weights of the multiple groups of feature maps respectively; wherein the reference feature map carries the spatial structure information of the image to be recognized; The image decoding module is further configured to decode the enhanced feature map to obtain the affordance detection result of the image to be recognized.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which when executed by a processor, implements the affordance detection method according to any one of claims 1-4.
7. An affordance detection device, characterized in that, The affordance detection device includes a processor and a memory. The memory stores a computer program, which when executed by the processor, implements the affordance detection method according to any one of claims 1-4.
Citation Information
Patent Citations
Article recognition model training method, article recognition method, device and electronic equipment
CN110781973A
Unmanned aerial vehicle image semantic segmentation identification method based on hierarchical processing
CN111209808A