Method and device for detecting object affordance

The method and apparatus enhance object affordance detection by transferring action intentions from reference images to target images, addressing the challenge of accurately identifying objects capable of specific actions in unknown environments, thereby improving scene understanding and human-computer interaction.

CN115082750BActive Publication Date: 2025-07-15BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110269399.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-12
Publication Date
2025-07-15
Estimated Expiration
2041-03-12

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect the availability of unseen objects, and is not robust in scenario understanding and human-computer interaction.

Method used

By acquiring the reference image and the image to be detected, the features are extracted and the person's action intention information on the object is captured, and the action intention information is used to migrate it into the image to be detected to segment the object that can complete the intention, and the common features between different objects with the same affordance are captured through a collaborative method to improve detection robustness.

Benefits of technology

It improves the availability detection capability and robustness of objects that have not been seen, and can more accurately identify objects that can complete specific action intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082750B_ABST
    Figure CN115082750B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for detecting object affordances, relating to the field of computers. The method includes: obtaining a reference image and a to-be-detected image, where the reference image includes a first object of a person and the person's action; extracting the features of the reference image and the features of the to-be-detected image; according to the features of the reference image, extracting the action intention information of the person in the reference image with respect to the first object; according to the action intention information of the person in the reference image with respect to the first object and the features of the to-be-detected image, migrating the action intention information to the to-be-detected image, and segmenting out a second object capable of completing the action intention information from the to-be-detected image. The detection ability of the affordances of unseen objects is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and particularly to a method and device for detecting object affordances. Background Art

[0002] Affordance was proposed by psychologist Gibson in 1966. It describes how to directly perceive the intrinsic value and meaning of objects in the environment and explains how this information is related to the action possibilities provided by the environment for the organism.

[0003] In practical applications, it is very important to perceive the affordances of various objects in an unknown environment, which has important application values in aspects such as scene understanding, action recognition, and human-computer interaction. Summary of the Invention

[0004] Embodiments of the present disclosure capture a person's action intention towards an object based on a reference image including a person and the object, and migrate it to all images to be detected, and segment out all objects that can complete the action intention from them, improving the detection ability of the affordances of unseen objects. In addition, a collaborative method is used to capture the common features (i.e., internal connections) between different objects with the same affordance, and based on the common features, multiple objects with this affordance are detected, improving the robustness of object affordance detection.

[0005] Some embodiments of the present disclosure propose a method for detecting object affordances, including:

[0006] Obtain a reference image and an image to be detected, where the reference image includes a person and a first object of the person's action;

[0007] Extract the features of the reference image and the features of the image to be detected;

[0008] According to the features of the reference image, extract the action intention information of the person in the reference image towards the first object;

[0009] According to the action intention information of the person in the reference image towards the first object and the features of the image to be detected, migrate the action intention information to the image to be detected, and segment out a second object from the image to be detected that can complete the action intention information.

[0010] In some embodiments, extracting the action intention information of the person in the reference image towards the first object includes:

[0011] According to the feature representation of the person in the reference image, perform a weighted operation on the features of the reference image to obtain a first output;

[0012] According to the feature representation of the first object in the reference image, perform a weighted operation on the features of the reference image to obtain a second output;

[0013] Obtain a third output that describes the relevant positions of the person's action on the first object according to the feature representation of the person and the feature representation of the first object in the reference image;

[0014] Determine the action intention encoding of the person on the first object in the reference image according to the first output, the second output, and the third output.

[0015] In some embodiments, performing a weighted operation on the features of the reference image according to the feature representation of the person in the reference image to obtain a first output, including: performing a correlation operation on the pooled feature representation of the person in the reference image and each position of the features of the reference image, normalizing the correlation operation result to obtain the weight of each position, and multiplying the weight of each position by the features of the reference image to obtain the first output.

[0016] In some embodiments, performing a weighted operation on the features of the reference image according to the feature representation of the first object in the reference image to obtain a second output, including: performing a correlation operation on the pooled feature representation of the first object in the reference image and each position of the features of the reference image, normalizing the correlation operation result to obtain the weight of each position, and multiplying the weight of each position by the features of the reference image to obtain the second output.

[0017] In some embodiments, obtaining a third output that describes the relevant positions of the person's action on the first object according to the feature representation of the person and the feature representation of the first object in the reference image, including: performing a correlation operation on the pooled feature representation of the first object and the feature representation of the person, and performing convolution processing on the correlation operation result to obtain a third output that describes the relevant positions of the person's action on the first object.

[0018] In some embodiments, determining the action intention encoding of the person on the first object in the reference image according to the first output, the second output, and the third output, including:

[0019] Multiplying the third output by the first output and pooling to obtain the first action intention sub-information;

[0020] Multiplying the third output by the second output and pooling to obtain the second action intention sub-information;

[0021] Adding the first action intention sub-information and the second action intention sub-information to obtain the action intention encoding of the person on the first object in the reference image.

[0022] In some embodiments, the bounding box of the person in the reference image is multiplied by the features of the reference image to obtain the feature representation of the person in the reference image; the bounding box of the first object in the reference image is multiplied by the features of the reference image to obtain the feature representation of the first object in the reference image.

[0023] In some embodiments, according to the action intention information of the person in the reference image towards the first object and the features of the image to be detected, migrating the action intention information to the image to be detected, and segmenting a second object capable of completing the action intention information from the image to be detected includes:

[0024] Using the action intention information of the person in the reference image towards the first object to perform a correlation operation with each position of the features of the image to be detected, and obtaining the weight of each position after normalization;

[0025] The weight of each position is multiplied by the features of the image to be detected, and the multiplication result is added to the features of the image to be detected to obtain the second object that can complete the action intention information segmented from the image to be detected.

[0026] In some embodiments, the method further includes:

[0027] Reconstructing the first features of the second object using a set of bases that can capture the common features between different objects with the same affordance;

[0028] Determining the second features of the second object according to the first features of the second object and the reconstructed first features of the second object;

[0029] Outputting an image of the second object according to the second features of the second object.

[0030] In some embodiments, the method for obtaining the set of bases includes: randomly initializing a set of bases, using a preset optimization algorithm, iteratively updating the set of bases by reducing the gap information between the training image and the training image obtained by performing a correlation operation using the set of bases, and using the updated set of bases as a set of bases learned to capture the common features between different objects with the same affordance, where the optimization algorithm includes the expectation maximization algorithm or the gradient descent algorithm.

[0031] Some embodiments of the present disclosure propose a detection device for object affordance, characterized by including: a memory; and a processor coupled to the memory, the processor being configured to execute a detection method for object affordance based on instructions stored in the memory.

[0032] Some embodiments of the present disclosure propose a detection device for object affordance, characterized by including:

[0033] A feature extraction module, configured to obtain a reference image and an image to be detected, where the reference image includes a first object of a person and a person's action; extract features of the reference image and features of the image to be detected;

[0034] An intention learning module, configured to extract action intention information of the person in the reference image with respect to the first object according to the features of the reference image;

[0035] An intention transfer module, configured to transfer the action intention information to the image to be detected according to the action intention information of the person in the reference image with respect to the first object and the features of the image to be detected, and segment a second object capable of completing the action intention information from the image to be detected.

[0036] In some embodiments, the apparatus further includes:

[0037] A collaborative enhancement module, configured to reconstruct the first features of the second object by using a set of bases, where the set of bases can capture common features between different objects with the same affordance; determine second features of the second object according to the first features of the second object and the reconstructed first features of the second object;

[0038] A decoding module, configured to output an image of the second object according to the second features of the second object.

[0039] Some embodiments of the present disclosure propose a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for detecting object affordance are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The following will briefly introduce the drawings required for use in the embodiments or related art descriptions. According to the following detailed description with reference to the drawings, the present disclosure can be more clearly understood.

[0041] Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.

[0042] Figure 1 A flowchart showing a method for detecting object affordance according to some embodiments of the present disclosure.

[0043] Figure 2 A schematic diagram showing the extraction of action intention information of a person in a reference image with respect to a first object according to some embodiments of the present disclosure.

[0044] Figure 3 A schematic diagram showing the Element-wise Multiplication process according to some embodiments of the present disclosure.

[0045] Figure 4 Schematic diagram showing the Position-wise Dot Product process of some embodiments of the present disclosure.

[0046] Figure 5 Schematic diagram showing migrating action intention information to an image to be detected and segmenting a second object capable of completing the action intention information from the image to be detected in some embodiments of the present disclosure.

[0047] Figure 6 Schematic diagram showing the collaborative enhancement step of some embodiments of the present disclosure.

[0048] Figure 7 Schematic diagram showing a detection device for object affordance of some embodiments of the present disclosure.

[0049] Figure 8 Schematic diagram showing a detection device for object affordance of some other embodiments of the present disclosure. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure.

[0051] Unless otherwise specified, the descriptions such as "first" and "second" in the present disclosure are used to distinguish different objects and do not indicate meanings such as size or time sequence.

[0052] Figure 1 Schematic flowchart showing a detection method for object affordance of some embodiments of the present disclosure.

[0053] As Figure 1 shown, the detection method for object affordance in this embodiment includes: steps 110-160, where step 150 can be selectively executed according to needs.

[0054] In step 110, an image acquisition step: acquiring a reference image (Support image) and an image to be detected (Queryimage).

[0055] The reference image includes a first object of a person and the person's action, and the bounding box of the person and the bounding box of the first object can be marked. For example, the reference image of "a person kicking a ball" includes the "person" kicking the ball and the "ball" kicked by the person, and a rectangular box for the "person" and a rectangular box for the "ball" are marked.

[0056] The image to be detected can be one or more. If there are multiple images to be detected, each image to be detected performs the same affordance detection operation as a single image to be detected.

[0057] In step 120, the feature extraction step: Extract the features of the reference image and the image to be detected.

[0058] Use an image feature extraction network, such as Resnet (Residual Network), VGGnet, etc., to extract the features of the reference image / the image to be detected.

[0059] In step 130, the intention learning step: According to the features of the reference image, extract the action intention information of the person in the reference image towards the first object.

[0060] In some embodiments, extracting the action intention information of the person in the reference image towards the first object includes steps 130.1 - 130.4, as Figure 2 shown.

[0061] In step 130.1, according to the feature representation of the person in the reference image, perform a weighted operation on the features of the reference image to obtain a first output.

[0062] Among them, the bounding box of the person in the reference image is multiplied by the features of the reference image to obtain the feature representation of the person in the reference image.

[0063] In some embodiments, performing a weighted operation on the features of the reference image according to the feature representation of the person in the reference image to obtain a first output includes: pooling the feature representation of the person in the reference image (such as global average pooling) and performing a correlation operation (such as element - wise multiplication) with each position of the features of the reference image. After normalizing the correlation operation results (such as the Softmax method), obtain the weights for each position, and multiply the weights for each position by the features of the reference image (such as element - wise multiplication) to obtain the first output.

[0064] In step 130.2, according to the feature representation of the first object in the reference image, perform a weighted operation on the features of the reference image to obtain a second output.

[0065] Among them, the bounding box of the first object in the reference image is multiplied by the features of the reference image to obtain the feature representation of the first object in the reference image.

[0066] In some embodiments, performing a weighted operation on the features of the reference image according to the feature representation of the first object in the reference image to obtain a second output, including: pooling the feature representation of the first object in the reference image (such as Global Average Pooling) and performing a correlation operation (such as Element-wise Multiplication) with each position of the features of the reference image. After the result of the correlation operation is normalized (such as by the Softmax method), the weight of each position is obtained, and the weight of each position is multiplied by the features of the reference image (such as Element-wise Multiplication) to obtain the second output.

[0067] In step 130.3, according to the feature representation of the person and the feature representation of the first object in the reference image, obtain a third output that describes the relevant position of the person's action on the first object.

[0068] In some embodiments, according to the feature representation of the person and the feature representation of the first object in the reference image, obtaining a third output that describes the relevant position of the person's action on the first object includes: pooling the feature representation of the first object (such as Global Average Pooling) and performing a correlation operation (such as Element-wise Multiplication) with the feature representation of the person, so that the relevant position of the person's action on the first object in the reference image is focused. After the result of the correlation operation is subjected to convolutional processing (conv, such as 1*1 or 3*3 convolutional processing), a third output that describes the relevant position of the person's action on the first object is obtained. The third output is a feature map of 1*H*W, where H is the height of the feature map of the reference image and W is the width of the feature map of the reference image.

[0069] In step 130.4, according to the first output, the second output, and the third output, determine the action intention encoding of the person on the first object in the reference image.

[0070] In some embodiments, according to the first output, the second output, and the third output, determining the action intention encoding of the person on the first object in the reference image includes: multiplying the third output by the first output (such as Position-wise Dot Product) and pooling to obtain the first action intention sub-information; multiplying the third output by the second output (such as Position-wise Dot Product) and pooling to obtain the second action intention sub-information; adding the first action intention sub-information and the second action intention sub-information (sum) to obtain the action intention encoding of the person on the first object in the reference image.

[0071] Such as Figure 3As shown in the figure, the Element-wise Multiplication process: Two matrices A([1, 1, C]) and B([W, H, C]) are input. The third dimension of A and B (i.e., the channel (Channel, C)) is multiplied, that is, the corresponding channels of A and B are multiplied, which is also called "channel-related operation".

[0072] As Figure 4 shown in the figure, the Position-wise Dot Product process: Two matrices A([W, H, 1]) and B([W, H, C]) are input. The first two dimensions of A and B are multiplied. The first two dimensions respectively represent the height (H) and width (W) of the feature map matrix, that is, the corresponding positions of A and B are multiplied, which is also called "position-related operation".

[0073] In step 140, the intention migration step: According to the action intention information of the person on the first object in the reference image and the features of the image to be detected, the action intention information is migrated to the image to be detected, and the second object (the first feature of the second object) that can complete the action intention information is segmented from the image to be detected.

[0074] In some embodiments, as Figure 5 shown in the figure, migrating the action intention information to the image to be detected and segmenting the second object that can complete the action intention information from the image to be detected includes: Using the action intention information of the person on the first object in the reference image (i.e., the action intention encoding), performing a correlation operation (such as Element-wise Multiplication) with each position of the features of the image to be detected (step 140.1), and obtaining the weight of each position after normalization (such as the Softmax method) (step 140.2); multiplying the weight of each position by the features of the image to be detected (such as Element-wise Multiplication) (step 140.3), adding the multiplication result to the features of the image to be detected (sum) (step 140.4), to obtain the second object (the first feature of the second object) that can complete the action intention information segmented from the image to be detected, that is, obtaining the features of the relevant region activated by the action intention information.

[0075] That different objects can complete the same action intention means that these different objects have the same affordance corresponding to this action intention.

[0076] In step 150, the collaborative enhancement step: Using a set of bases to reconstruct the first feature of the second object. This set of bases can capture the common features between different objects with the same affordance. According to the first feature of the second object and the reconstructed first feature of the second object, the second feature of the second object is determined.

[0077] In some embodiments, as Figure 6As shown, the collaborative enhancement step specifically includes: performing a correlation operation (such as Position-wise Dot Product) on the first feature of the second object with this set of bases for reconstruction (step 150.1). After the first feature of the reconstructed second object undergoes convolution processing (conv) (step 150.2), it is added (sum) to the first feature of the second object (step 150.3) to obtain the second feature of the second object.

[0078] This set of bases can capture the internal connection between different objects with the same affordance (reflected by common features), suppress irrelevant background regions during the detection process based on this set of bases, and obtain better detection effects.

[0079] The method for obtaining this set of bases includes: randomly initializing a set of bases, using a preset optimization algorithm, and iteratively updating this set of bases by continuously reducing the gap information between the training image and the training image after performing a correlation operation (such as Position-wise Dot Product) with this set of bases. The updated set of bases is used as the set of bases that can capture the common features between different objects with the same affordance. Among them, the optimization algorithm includes the Expectation-Maximum (EM) algorithm or the gradient descent algorithm. The number of this set of bases is usually much smaller than the product of the height (H) and width (W) of the image. The number of this set of bases is, for example, several or more than a dozen.

[0080] For example, randomly initialize a set of bases. The form of each base is [1, 1, C], where C represents the channel. Each base performs a correlation operation (such as Position-wise Dot Product) with the features of a training image with a size of [W, H, C]. W and H respectively represent the width and height of the feature map of the training image. Combine the corresponding correlation operation results of multiple bases to obtain the features of the training image after performing a correlation operation with this set of bases. The features of the training image after the correlation operation are transformed into features with a size of [W, H, C] through convolution processing. Use the expectation maximization optimization method to iteratively update this set of bases so that the gap between the features of the training image and the features of the training image after the correlation operation and convolution processing continuously shrinks until the preset number of iterations is reached or the gap is less than the preset value, and then stop the iteration. The updated set of bases is the set of bases that can capture the common features between different objects with the same affordance.

[0081] In step 160, the decoding output step: Output the image of the second object through decoding.

[0082] If following step 140, output the image of the second object through decoding according to the first feature of the second object. If following step 150, output the image of the second object through decoding according to the second feature of the second object.

[0083] Decoding means restoring the image features to the corresponding images. Decoding can be achieved, for example, through deconvolution processing, or through processing of first upsampling and then convolution.

[0084] In the above embodiments, based on a reference image including a person and an object, the action intention of the person towards the object is captured and migrated to all images to be detected, and all objects that can complete the action intention are segmented therefrom, improving the detection ability of the affordance of unseen objects. In addition, in the above embodiments, the common features (i.e., internal relationships) between different objects with the same affordance are captured through a collaborative method, and multiple objects with this affordance are detected based on the common features, improving the robustness of object affordance detection.

[0085] In some application examples, for example, given a reference image of "a person kicking a ball", the action intention of "a person kicking a ball" is captured from the reference image, and based on the action intention of "a person kicking a ball" captured from the reference image, it is migrated to all images to be detected, and all spherical objects that meet the action intention are segmented therefrom, improving the detection ability of the affordance of unseen objects; the common features between different objects that meet the action intention, such as an arc-shaped appearance, can also be captured through a collaborative method, and multiple objects that meet the action intention are detected based on the common features, improving the robustness of object affordance detection.

[0086] Figure 7 A schematic diagram of a detection device for object affordance showing some embodiments of the present disclosure. The detection device for object affordance is also referred to as a detection network for object affordance.

[0087] As Figure 7 shown, the detection device 700 for object affordance of this embodiment includes: modules 710-750, wherein module 740 is selectively configured or executed.

[0088] The feature extraction module 710 is configured to obtain a reference image and images to be detected, the reference image including a person and a first object of the person's action; and extract the features of the reference image and the images to be detected.

[0089] The intention learning module 720 is configured to extract the action intention information of the person towards the first object in the reference image according to the features of the reference image.

[0090] The intention migration module 730 is configured to migrate the action intention information to the images to be detected according to the action intention information of the person towards the first object in the reference image and the features of the images to be detected, and segment a second object (the first feature of the second object) that can complete the action intention information from the images to be detected.

[0091] The decoding module 750 is configured to output an image of the second object according to the first feature of the second object.

[0092] In some embodiments, the object availability detection device 700 further includes: a collaborative enhancement module 740 configured to reconstruct the first feature of the second object by using a set of bases that can capture the common features among different objects having the same availability; determine the second feature of the second object according to the first feature of the second object and the reconstructed first feature of the second object. At this time, the decoding module 750 is configured to output an image of the second object according to the second feature of the second object.

[0093] For the specific processing of the operations performed by the above modules, reference can be made to the foregoing embodiments, which will not be elaborated here.

[0094] The object availability detection device 700 needs to be trained before use. However, the object availability detection device 700 can be pre-trained and ready for direct use.

[0095] The training process of the object affordance detection device 700 includes: obtaining a data set; dividing the data set into a training set and a test set. Both the training set and the test set include reference images and images to be detected. One reference image can correspond to one or more images to be detected. The reference images in the training set are labeled with the bounding boxes of people and the first object, and the images to be detected in the training set are pre-labeled with the second object that can satisfy the action intention of the person in the reference image for the first object. Input the reference images and images to be detected in the training set into the object affordance detection device 700 for detection. The detected object is set as the third object. Determine the loss according to the gap information between the detected third object and the pre-labeled second object and a loss function (such as cross entropy). Use an optimization function (such as adam, sgd (Stochastic Gradient Descent)) to optimize the network parameters (such as various parameters in the convolution process of each module) in the detection device 700 to reduce the loss to a certain extent, and the training is completed. Then, use the test set to test the trained detection device 700. The reference images in the test set are labeled with the bounding boxes of people and the first object, and the images to be detected in the test set are pre-labeled with the second object that can satisfy the action intention of the person in the reference image for the first object. Input the reference images and images to be detected in the test set into the object affordance detection device 700 for detection. The detected object is set as the third object. Determine the accuracy of the detection according to whether the detected third object belongs to the pre-labeled second object. If the accuracy of the detection is higher than a certain degree, it is considered that the detection device 700 passes the test and is qualified. If the detection device 700 fails the test, the detection device 700 can be continuously trained by increasing the training samples or increasing the number of training iterations. During training, for example, select the data of 1 / 3 of the affordance categories as the test set, and the data of the remaining affordance categories as the training set for training to improve the training effect.

[0096] Figure 8 Schematic diagram of an object affordance detection device showing other embodiments of the present disclosure.

[0097] As Figure 8 shown, the object affordance detection device 800 of this embodiment includes: a memory 810 and a processor 820 coupled to the memory 810. The processor 820 is configured to execute the object affordance detection method in any of the foregoing embodiments based on instructions stored in the memory 810.

[0098] Among them, the memory 810 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.

[0099] The detection device 800 may further include an input / output interface 830, a network interface 840, a storage interface 850, etc. These interfaces 830, 840, 850, the memory 810, and the processor 820 may be connected, for example, through a bus 860. Among them, the input / output interface 830 provides connection interfaces for input / output devices such as a display, a mouse, a keyboard, a touch screen, etc. The network interface 840 provides connection interfaces for various networking devices. The storage interface 850 provides connection interfaces for external storage devices such as an SD card and a USB flash drive.

[0100] Some embodiments of the present disclosure propose a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method for detecting object availability in any of the foregoing embodiments.

[0101] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more non-transitory computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer program code.

[0102] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0103] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or boxes Figure 1 in one or more processes and / or boxes Figure 1 steps for the functions specified in one box or a plurality of boxes.

[0105] The foregoing are only preferred embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for detecting object affordances, characterized in that, Including: Obtaining a reference image and an image to be detected, where the reference image includes a first object of a person and the person's action; Extracting the features of the reference image and the features of the image to be detected; According to the features of the reference image, extracting the action intention information of the person in the reference image for the first object, where the features of the reference image include the feature representation of the person in the reference image and the feature representation of the first object; According to the action intention information of the person in the reference image for the first object and the features of the image to be detected, migrating the action intention information to the image to be detected, and segmenting out a second object capable of completing the action intention information from the image to be detected, including: using the action intention information of the person in the reference image for the first object to perform a correlation operation with each position of the features of the image to be detected to determine the weight of each position, and segmenting out a second object capable of completing the action intention information from the image to be detected according to the weight of each position.

2. The method according to claim 1, wherein Extracting the action intention information of the person in the reference image for the first object includes: Performing a weighted operation on the features of the reference image according to the feature representation of the person in the reference image to obtain a first output; Performing a weighted operation on the features of the reference image according to the feature representation of the first object in the reference image to obtain a second output; Obtaining a third output describing the relevant positions of the person's action on the first object according to the feature representation of the person and the feature representation of the first object in the reference image; Determining the action intention encoding of the person in the reference image for the first object according to the first output, the second output, and the third output, including: determining the action intention encoding of the person in the reference image for the first object according to the position correlation operation result between the third output and the first output and the position correlation operation result between the third output and the second output.

3. The method according to claim 2, wherein Performing a weighted operation on the features of the reference image according to the feature representation of the person in the reference image to obtain a first output, including: Pooling the feature representation of the person in the reference image and performing a correlation operation with each position of the features of the reference image. After normalizing the correlation operation result, the weight of each position is obtained, and the weight of each position is multiplied by the features of the reference image to obtain a first output.

4. The method according to claim 2, characterized in that, Performing a weighted operation on the features of the reference image according to the feature representation of the first object in the reference image to obtain a second output, including: Pooling the feature representation of the first object in the reference image and performing a correlation operation with each position of the features of the reference image. After normalizing the correlation operation result, the weight of each position is obtained, and the weight of each position is multiplied by the features of the reference image to obtain a second output.

5. The method according to claim 2, characterized in that Obtaining a third output describing the relevant positions of the person's action on the first object according to the feature representation of the person and the feature representation of the first object in the reference image, including: Pooling the feature representation of the first object and performing a correlation operation with the feature representation of the person. After convolution processing of the correlation operation result, a third output describing the relevant positions of the person's action on the first object is obtained.

6. The method according to claim 2, characterized in that, Determine the action intention encoding of the person in the reference image for the first object according to the first output, the second output, and the third output, including: Multiply the third output by the first output and perform pooling to obtain first action intention sub-information; Multiply the third output by the second output and perform pooling to obtain second action intention sub-information; Add the first action intention sub-information and the second action intention sub-information to obtain the action intention encoding of the person in the reference image for the first object.

7. The method according to claim 2, wherein Multiply the bounding box of the person in the reference image by the features of the reference image to obtain the feature representation of the person in the reference image; Multiply the bounding box of the first object in the reference image by the features of the reference image to obtain the feature representation of the first object in the reference image.

8. The method according to claim 1, wherein According to the action intention information of the person in the reference image for the first object and the features of the image to be detected, transfer the action intention information to the image to be detected, and segment a second object that can complete the action intention information from the image to be detected, including: Use the action intention information of the person in the reference image for the first object to perform a correlation operation with each position of the features of the image to be detected, and obtain the weight of each position after normalization; Multiply the weight of each position by the features of the image to be detected, and add the multiplication result to the features of the image to be detected to obtain the second object that can complete the action intention information segmented from the image to be detected.

9. The method according to claim 1, characterized in that Further include: Reconstruct the first feature of the second object using a set of bases that can capture the common features between different objects with the same affordance. The method for obtaining this set of bases includes: randomly initialize a set of bases, use a preset optimization algorithm, and iteratively update the set of bases by reducing the gap information between the training image and the training image after performing a correlation operation using this set of bases, and use the updated set of bases as the learned set of bases that can capture the common features between different objects with the same affordance; Determine the second feature of the second object according to the first feature of the second object and the reconstructed first feature of the second object; Output the image of the second object according to the second feature of the second object.

10. The method according to claim 9, characterized in that, Wherein, The optimization algorithm includes the expectation maximization algorithm or the gradient descent algorithm.

11. A detection device for object affordance, characterized in that, Include: A memory; And A processor coupled to the memory, the processor being configured to execute the method for detecting object affordance according to any one of claims 1-10 based on instructions stored in the memory.

12. A detection device for object affordance, characterized in that, Include: A feature extraction module configured to obtain a reference image and an image to be detected, the reference image including a person and a first object of the person's action; Extract the features of the reference image and the features of the image to be detected; An intention learning module configured to extract the action intention information of the person in the reference image for the first object according to the features of the reference image, the features of the reference image including the feature representation of the person in the reference image and the feature representation of the first object. An intention transfer module, configured to transfer the action intention information to the image to be detected according to the action intention information of a person on a first object in the reference image and the features of the image to be detected, and segment a second object capable of completing the action intention information from the image to be detected, including: performing a correlation operation on the action intention information of a person on the first object in the reference image and each position of the features of the image to be detected to determine the weight of each position, and segmenting a second object capable of completing the action intention information from the image to be detected according to the weight of each position.

13. The device according to claim 12, characterized in that Further comprising: A collaborative enhancement module, configured to reconstruct the first features of the second object by using a set of bases, where the set of bases can capture the common features between different objects with the same affordance. The method for obtaining the set of bases includes: randomly initializing a set of bases, using a preset optimization algorithm, iteratively updating the set of bases by reducing the gap information between the training image and the training image after performing a correlation operation using the set of bases, and using the updated set of bases as a set of bases learned to capture the common features between different objects with the same affordance; determining the second features of the second object according to the first features of the second object and the first features of the reconstructed second object; A decoding module, configured to output an image of the second object according to the second features of the second object.

14. A non-transitory computer-readable storage medium, having stored thereon a computer program, which when executed by a processor, implements the steps of the method for detecting object affordance according to any one of claims 1-10.

Citation Information

Patent Citations

  • Complex product availability characteristic identification method designed for maintainability

    CN105868917A

  • System and Apparatus for Information Retrieval

    US20140280179A1