Host-guest interaction recognition model

By generating environment-independent subject, object, and environment weights using deep convolutional neural networks (DCN) and three-stream CNN, the problem of low accuracy in recognizing subject-object interactions when separated from the environment is solved, and robust subject-object interaction recognition is achieved.

CN113574533BActive Publication Date: 2026-02-17QUALCOMM TECHNOLOGIES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080020352.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-22
Filing Date
2020-03-23
Publication Date
2026-02-17
Estimated Expiration
2040-03-23

AI Technical Summary

Technical Problem

Existing convolutional neural networks are not very accurate in recognizing subject-object interactions out of context, and cannot effectively distinguish between subject, object and environment, resulting in decreased recognition accuracy in non-training environments.

Method used

A deep convolutional neural network (DCN) is used for subject-object interaction recognition. The model is trained to generate relative weights of the subject, object, and environment that are independent of the environment. A three-stream convolutional neural network (CNN) is used to generate D-dimensional image features of the subject-object-environment triples for classification.

Benefits of technology

It improves the accuracy of subject-object interaction recognition under environmental conditions, and can robustly identify subjects and objects in different environments, reducing the impact of the environment on recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113574533B_ABST
    Figure CN113574533B_ABST
Patent Text Reader

Abstract

A method for processing an image is presented. The method locates a subject and an object of a human-object interaction in the image. The method determines relative weights of the subject, the object, and an environmental region for classification. The method further classifies the human-object interaction based on a classification of a weighted representation of the subject, a weighted representation of the object, and a weighted representation of the environmental region.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Patent Application No. 16 / 827,592, filed March 23, 2020, entitled “ITERATIVE REFINEMENT OF PHYSICSSIMULATIONS,” and the benefit of Greek Patent Application No. 20190100141, filed March 22, 2019, entitled “SUBJECT-OBJECT INTERACTION RECOGNITION,” the disclosures of which are expressly incorporated herein by reference in their entirety.

[0003] background

[0004] field

[0005] The aspects disclosed herein generally involve the identification of host-guest interactions. Background Technology

[0006] Artificial neural networks can include a group of interconnected artificial neurons (e.g., neuron models). Artificial neural networks can be computing devices or represented as methods performed by computing devices. Convolutional neural networks (such as deep convolutional neural networks) are a type of feedforward artificial neural network. Convolutional neural networks can include individual neuron layers that can be configured in tiled receptive fields.

[0007] Deep convolutional neural networks (DCNs) are used in various technologies, such as vision systems, speech recognition, autonomous driving, and Internet of Things (IoT) devices. Vision systems can identify interactions between subjects (e.g., actors) and objects. During an interaction, the subject acts on the object within a scene (e.g., an environment). Subject-object interactions can be represented as noun-verb pairs, such as horse-ride or milk-milk. Improving the accuracy of interaction recognition systems is desirable.

[0008] Overview

[0009] In one aspect of this disclosure, a method for processing an image locates the subject and object of a subject-object interaction within the image. The method also determines the relative weights of the subject, the object, and an ambient region for classification. The method further classifies the subject-object interaction based on a classification of weighted representations of the subject, the object, and the ambient region.

[0010] In another aspect of the disclosure, an apparatus for processing an image includes at least one processor coupled to a memory and configured to locate a subject and an object of a human-object interaction in the image. The processor(s) are also configured to determine relative weights of the subject, the object, and an environmental region for classification. The processor(s) are further configured to classify the human-object interaction based on a classification of a weighted representation of the subject, a weighted representation of the object, and a weighted representation of the environmental region.

[0011] In yet another aspect of the disclosure, an apparatus for processing an image includes means for locating a subject and an object of a human-object interaction in the image. The apparatus also includes means for determining relative weights of the subject, the object, and an environmental region for classification. The apparatus further includes means for classifying the human-object interaction based on a classification of a weighted representation of the subject, a weighted representation of the object, and a weighted representation of the environmental region.

[0012] In yet another aspect of the disclosure, a computer-readable medium storing program code for processing an image. The program code is executed by at least one processor and includes program code to locate a subject and an object of a human-object interaction in the image. The computer-readable medium also stores program code to determine relative weights of the subject, the object, and an environmental region for classification. The computer-readable medium further stores program code to classify the human-object interaction based on a classification of a weighted representation of the subject, a weighted representation of the object, and a weighted representation of the environmental region.

[0013] These and other features and advantages of the present disclosure can be better understood with respect to the following detailed description and drawings, of which: BRIEF DESCRIPTION OF DRAWINGS

[0015] The features, nature, and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings in which like reference characters identify correspondingly throughout and wherein:

[0016] Figure 1An example implementation of designing a neural network using a system on a chip (SOC), including a general purpose processor, is illustrated in accordance with certain aspects of the present disclosure.

[0017] Figure 2A 、 2B and 2C are diagrams illustrating a neural network in accordance with aspects of the present disclosure.

[0018] Figure 2D is a diagram illustrating an example deep convolutional network (DCN) in accordance with aspects of the present disclosure.

[0019] Figure 3 is a block diagram illustrating an example deep convolutional network (DCN) in accordance with aspects of the present disclosure.

[0020] Figure 4 Examples of in-context images and out-of-context images are illustrated in accordance with aspects of the present disclosure.

[0021] Figure 5 An example of a model for classifying human-object interactions from out-of-context images is illustrated in accordance with aspects of the present disclosure.

[0022] Figure 6 An example network that produces activations for each subject, object, and context region is illustrated in accordance with aspects of the present disclosure.

[0023] Figure 7 An example of context-agnostic image feature learning is illustrated in accordance with aspects of the present disclosure.

[0024] Figure 8 An example framework for identifying subject-object-context image features for classification is illustrated in accordance with aspects of the present disclosure.

[0025] Figure 9 A flowchart of a method for classifying human-object interactions from images is illustrated in accordance with aspects of the present disclosure.

[0026] DETAILED DESCRIPTION

[0027] The detailed description set forth below, in connection with the appended drawings and specifications, is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein can be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without

[0028] Based on this teaching, it is to be understood that the scope of the disclosure is intended to cover any aspect of the disclosure and any combinations of them, whether implemented independently of, or in combination with, any other aspect of the disclosure. For example, an apparatus can be implemented or a method can be practiced using any number of the aspects set forth. In addition, the scope of the disclosure is intended to cover such an apparatus or method which is practiced using, as substitute for, or in addition to, some of the features of the aspects of the disclosure described. It is understood that any aspect of the disclosure disclosed can be implemented by one or more elements of a claim.

[0029] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects.

[0030] While certain aspects are described herein, changes and modifications can be made to those aspects without departing from the scope of the present disclosure. Although benefits and advantages of preferred aspects are described herein, the scope of the disclosure should not be limited to particular benefits or advantages. It should be understood that aspects of the disclosure can be used in other without departing from the scope of the present disclosure. Conversely, various changes and modifications can be understood within the scope of the aspects. The detailed description and drawings are merely illustrative of the disclosure rather than limiting, the scope of the disclosure being defined by the appended claims and equivalents thereof.

[0031] An interaction recognition system identifies interactions between a subject (e.g., an actor) and an object. The subject can be referred to as an interacting actor and the object can be referred to as an interactee. For example, the subject can be a human and the object can be a horse. During an interaction, the subject acts on the object in a particular context (e.g., environment) to achieve a purpose, such as a transportation purpose (e.g., horse riding). An interaction can be represented as a noun-verb pair, such as horse-ride or milk-pour. Accurate noun-verb pair identification can improve various applications, such as visual search, indexing, and image tagging.

[0032] Convolutional neural networks (CNNs) have improved the accuracy of interaction recognition systems. In some instances, when the target interaction is documented in the training set, the recognition performance can be on par with that of a human performing the interaction recognition. However, in real-world contexts, human-object interactions are not limited to the interactions documented in the training set.

[0033] For example, horse riding often occurs in rural settings or equestrian centers. Thus, a training set for horse riding can be limited to examples from rural settings or equestrian centers. Nonetheless, horse riding can also occur in cities. For example, a police officer can ride a horse in a city. Horse riding in a city can be considered an out-of-context interaction. While a human can identify the horse riding interaction in the city, an image recognition system can not be able to recognize this out-of-context interaction.

[0034] Off-environment interactions refer to principal-agent interactions that have limited or zero examples in a training set based on environment-based training samples. Environment-based training samples can include, for example, a skier in the snow, a swimmer in a pool, a basketball player on a basketball court, etc. Aspects of the disclosure are not limited to human-agent interaction identification. Other types of interactions (e.g., principal-agent interactions) can be identified.

[0035] Conventional CNNs are unable to accurately identify representations off-environment. Moreover, conventional CNNs rely on an environment to some extent to correctly identify principals, agents, and / or interactions. Recognition of off-environment interactions can improve many tasks. For example, conflict avoidance for an autonomous vehicle can be improved by accurately identifying a horseback rider within a city. It is desirable for an image interaction recognition system to identify principal-agent interactions without having an environment.

[0036] Aspects of the disclosure relate to environment-agnostic principal-agent interaction classification. In one configuration, a model (e.g., an interaction recognition model) classifies principal-agent interactions off-environment. The identification of principal-agent interaction regions by the model can be robust to environmental changes. Representations from the principal-agent interaction regions can be used to classify images having principal-agent interactions.

[0037] Figure 1 An example implementation of a system on a chip (SOC) 100 is illustrated that can include a central processing unit (CPU) 102 or multi-core CPU configured to classify principal-agent interactions from images in accordance with certain aspects of the disclosure. Variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., neural networks with weights), delays, frequency bin information, and task information can be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or can be distributed across multiple blocks. Instructions executed at the CPU 102 can be loaded from a program memory associated with the CPU 102 or can be loaded from the memory block 118.

[0038] The SOC 100 can also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which can include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that, for example, can detect and recognize gestures. In one implementation, the NPU is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SOC 100 can also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120 (which can include a global positioning system).

[0039] The SOC 100 can be based on an ARM instruction set. In an aspect of the disclosure, instructions loaded into the general purpose processor 102 can include code to locate a subject and an object in an image while ignoring at least one of a bystander or a background object. The general purpose processor 102 can further include code to identify relative weights of the subject, the object, and an environmental region for classification. The general purpose processor 102 can still further include code to classify the subject-object interaction based on a classification of a weighted representation of the subject, a weighted representation of the object, and a weighted representation of the environmental region.

[0040] A deep learning architecture can perform an object recognition task by learning to represent inputs in each layer with successively higher levels of abstraction, thereby building a useful feature representation of the input data. In this way, deep learning addresses a major bottleneck of traditional machine learning. Prior to the advent of deep learning, a machine learning approach to an object recognition problem can have relied heavily on human-engineered features, perhaps in combination with a shallow classifier. A shallow classifier can be a two-class linear classifier, for example, in which a weighted sum of the feature vector components can be compared to a threshold to predict which class the input belongs to. Human-engineered features can be templates or kernels tailored to a specific problem domain by engineers with domain expertise. In contrast, a deep learning architecture can learn to represent features similar to what a human engineer might design, but it learns through training. Moreover, a deep network can learn to represent and recognize new types of features that a human might not have considered.

[0041] Deep learning architectures can learn hierarchies of features. For example, if visual data is presented to a first layer, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if auditory data is presented to a first layer, the first layer can learn to recognize spectral power in certain frequencies. A second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as recognizing simple shapes for visual data or sound combinations for auditory data. Higher layers can learn to represent complex shapes in visual data or words in auditory data, for example. Still higher layers can learn to recognize common visual objects or spoken phrases.

[0042] Deep learning architectures can perform particularly well when applied to problems that have a natural hierarchical structure. For example, classification of motor vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher layers to recognize cars, trucks, and airplanes.

[0043] Neural networks can be designed with various connectivity patterns. In feed-forward networks, information is passed from lower to higher layers, with each neuron in a given layer communicating to neurons in higher layers. As described above, hierarchical representations can be built in successive layers of a feed-forward network. Neural networks can also have recurrent or feedback (also called top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can be helpful in recognizing patterns that span more than one chunk of input data delivered to the neural network in sequence. Connections from a neuron in a given layer to a neuron in a lower layer are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when recognition of high-level concepts can aid in discriminating particular low-level features of the input.

[0044] Connections between layers of a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In a fully connected neural network 202, a neuron in a first layer can communicate its output to every neuron in a second layer, so that every neuron in the second layer will receive input from every neuron in the first layer. Figure 2BAn example of a locally connected neural network 204 is illustrated. In a locally connected neural network 204, neurons in a first layer can be connected to a limited number of neurons in a second layer. More generally, a locally connected layer of a locally connected neural network 204 can be configured so that each neuron in a layer will have the same or a similar connectivity pattern, but its connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connectivity pattern can result in spatially distinct receptive fields in higher layers, since higher layer neurons in a given region can receive inputs that are tuned through training to properties of a restricted portion of the total input to the network.

[0045] One example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. A convolutional neural network 206 can be configured so that the connection strengths associated with inputs for each neuron in a second layer are shared (e.g., 208). Convolutional neural networks can be well suited for problems in which the spatial location of inputs is meaningful.

[0046] One type of convolutional neural network is a deep convolutional network (DCN). Figure 2D A detailed example of a DCN 200 designed to recognize visual features from image 226 inputs from an image capture device 230, such as a vehicle-mounted camera, is illustrated. The DCN 200 of the current example can be trained to identify traffic signs and the numbers provided on traffic signs. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or identifying traffic lights.

[0047] The DCN 200 can be trained with supervised learning. During training, an image, such as image 226 of a speed limit sign, can be presented to the DCN 200, and then a "forward pass" can be computed to produce an output 222. The DCN 200 can include a feature extraction section and a classification section. Upon receiving the image 226, a convolutional layer 232 can apply convolutional kernels (not shown) to the image 226 to generate a first set of feature maps 218. As an example, the convolutional kernels of the convolutional layer 232 can be 5x5 kernels that generate 28x28 feature maps. In the present example, since four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to the image 226 at the convolutional layer 232. The convolutional kernels can also be referred to as filters or convolutional filters.

[0048] The first set of feature maps 218 can be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 220. The max pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (such as 14x14) is smaller than the size of the first set of feature maps 218 (such as 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0049] In Figure 2D In the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Further, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 can include a number corresponding to a possible feature of the image 226, such as "sign," "60," and "100." A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. As such, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.

[0050] In the present example, the probabilities for "sign" and "60" in the output 222 are higher than the probabilities of other features (such as "30," "40," "50," "70," "80," "90," and "100") of the output 222. Prior to training, the output 222 produced by the DCN 200 is likely to be incorrect. As such, an error between the output 222 and a target output can be computed. The target output is the ground truth of the image 226 (e.g., "sign" and "60"). The weights of the DCN 200 can then be adjusted so that the output 222 of the DCN 200 more closely aligns with the target output.

[0051] To adjust the weights, a learning algorithm can compute a gradient vector for the weights. The gradient can indicate the amount by which the error will increase or decrease if the weights are adjusted. At the top layer, the gradient can directly correspond to the value of the weight connecting the activation neuron in the penultimate layer to the neuron in the output layer. In lower layers, the gradient can depend on the value of the weight and the computed error gradient of the higher layer. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be referred to as "backpropagation" because it involves a "backward pass" through the neural network.

[0052] In practice, the error gradient of the weights can be computed on a small number of examples, so that the computed gradient approximates the true error gradient. This approximation method can be referred to as stochastic gradient descent. Stochastic gradient descent can be repeated until the error rate that the entire system can achieve has stopped decreasing or until the error rate has reached a target level. After learning, the DCN can be presented with new images and the forward pass through the network can produce an output 222, which can be considered an inference or prediction of the DCN.

[0053] A deep belief network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. A DBN can be used to extract a hierarchical representation of a training data set. A DBN can be obtained by stacking multiple layers of restricted Boltzmann machines (RBMs). An RBM is a class of artificial neural network that can learn a probability distribution over a set of inputs. Because an RBM can learn a probability distribution without information about which class each input should be classified into, RBMs are often used for unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBMs of a DBN can be trained in an unsupervised manner and can be used as feature extractors, while the top RBMs can be trained in a supervised manner (on the joint distribution of inputs and target classes from the previous layer) and can be used as classifiers.

[0054] A deep convolutional network (DCN) is a network of convolutional networks configured with additional pooling and normalization layers. DCNs have achieved state-of-the-art performance on many tasks. DCNs can be trained using supervised learning, where both the input and the output target are known for many exemplars and are used to modify the weights of the network by using a gradient descent method.

[0055] A DCN can be a feedforward network. In addition, as described above, the connections from neurons in a first layer of the DCN to a group of neurons in the next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of a DCN can be exploited for fast processing. The computational burden of a DCN can be much less than, for example, a similarly sized neural network that includes recurrent or feedback connections.

[0056] The processing of each layer of a convolutional network can be thought of as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be thought of as three-dimensional, with two spatial dimensions along the axes of the image and a third dimension capturing color information. The output of a convolutional connection can be thought of as forming a feature map in a subsequent layer, with each element in the feature map (e.g., 220) receiving input from a range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values in the feature map can be further processed with a nonlinearity, such as the rectified max(0, x). Values from neighboring neurons can be further pooled, which corresponds to down-sampling, and can provide additional local invariance and dimensionality reduction. Normalization can also be applied through lateral inhibition between neurons in a feature map, which corresponds to whitening.

[0057] The performance of deep learning architectures can improve as more labeled data points become available or as computing power improves. Modern deep neural networks are routinely trained with many orders of magnitude more computing resources than were available to a typical researcher just fifteen years ago. New architectures and training paradigms can further push the performance of deep learning. Rectified linear units can reduce a training problem known as vanishing gradients. New training techniques can reduce over-fitting and thus enable larger models to achieve better generalization. Encapsulation techniques can abstract data in a given receptive field and further boost overall performance.

[0058] Figure 3 is a block diagram illustrating a deep convolutional network 350. The deep convolutional network 350 can include multiple different types of layers based on connectivity and weight sharing. As shown in Figure 3 The deep convolutional network 350 includes convolutional blocks 354A, 354B. Each of the convolutional blocks 354A, 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max-pooling layer (MAX POOL) 360.

[0059] The convolutional layer 356 can include one or more convolutional filters that can be applied to input data to generate a feature map. Although only two convolutional blocks 354A, 354B are shown, the present disclosure is not so limited, but rather any number of convolutional blocks 354A, 354B can be included in the deep convolutional network 350 according to design preference. The normalization layer 358 can normalize the output of the convolutional filters. For example, the normalization layer 358 can provide whitening or lateral inhibition. The max-pooling layer 360 can provide a down-sampling aggregation in space to achieve local invariance and dimensionality reduction.

[0060] For example, a parallel filter bank of a deep convolutional network can be loaded onto CPU 102 or GPU 104 of SOC 100 to achieve high performance and low power consumption. In alternative embodiments, the parallel filter bank can be loaded onto DSP 106 or ISP 116 of SOC 100. Additionally, deep convolutional network 350 can access other processing blocks that can be present on SOC 100, such as sensor processor 114 and navigation module 120 that are specialized for sensors and navigation, respectively.

[0061] Deep convolutional network 350 can also include one or more fully connected layers 362 (FC1 and FC2). Deep convolutional network 350 can further include a logistic regression (LR) layer 364. Between each layer 356, 358, 360, 362, 364 of deep convolutional network 350 are weights (not shown) to be updated. The output of each layer (e.g., 356, 358, 360, 362, 364) can be used as input to a subsequent layer (e.g., 356, 358, 360, 362, 364) in deep convolutional network 350 to learn hierarchical feature representations from input data 352 (e.g., images, audio, video, sensor data, and / or other input data) supplied at first convolutional block 354A. The output of deep convolutional network 350 is a classification score 366 for input data 352. Classification score 366 can be a set of probabilities, where each probability is a probability that input data includes a feature from a set of features.

[0062] If the interaction recognition model discussed, the interaction between the subject and the object identified in the image can be classified. The interaction can be a decontextualized interaction. The interaction recognition model can locate the interaction between the subject and the object in the image while ignoring bystanders. In one configuration, the representations of the subject and the object can be context-free.

[0063] That is, the interaction between the subject and the object can be identified without considering the context. The model can be trained to represent the subject and the object that are context-free. The localization of the subject and the object can be based on prior training. The term context-free can also be referred to as context-neutral or context-agnostic.

[0064] The relative weights of the subject, the object, and the context can be identified for classification. The interaction between the subject and the object can be classified based on the weights. For example, the interaction is sampled to create a subject-object-context triple regarding the entity weights. The entity weights provide the relative importance of the subject, the object, and the context.

[0065] The model can generate subject-object-environment triplets via a CNN that receives only subject image regions, only object image regions, and only environment image regions. The CNN can be a three-stream CNN. The CNN outputs D-dimensional image features or embeddings for each region (e.g., subject, object, and environment regions). The D-dimensional image features include a D-dimensional activation vector for each subject, object, and environment image region.

[0066] For example, only subject image regions and only object image regions are masked to generate only environment image regions. As another example, only subject image regions and only environment image regions are masked to generate only object image regions. In yet another example, only environment image regions and only object image regions are masked to generate only subject image regions.

[0067] In one aspect of the disclosure, the model receives D-dimensional image features for all subjects and / or objects within an image and produces an Nx1 dimensional vector representing the contribution of each region to the final response. N represents the number of detected objects. Detected objects refer to subjects and objects detected in the image.

[0068] The model computes a weighted aggregation of all subject and / or object activations within the input. For example, object activations are representations of objects detected in the image. An image feature function can produce a D-dimensional activation vector for each object or subject detected in the image. The environment is sampled to create subject-object-environment triplets for entity weighting to determine the relative importance of subjects, objects, and environments for the final classification.

[0069] In one example, the model receives 3xD-dimensional image features for subject-object-environment triplets. The model can generate a 3x1 dimensional vector representing the relative importance of subjects, objects, and environments for the final classification.

[0070] Figure 4 Examples in environment images 402a-402d and out-of-context images 404a-404b are illustrated in accordance with aspects of the disclosure. As discussed, an image interaction recognition model can learn interactions from in-context images 402a-402d. For example, a first in-context image 402a depicts snowboarding on a snow mountain, a second in-context image 402b depicts snowboarding on a snow mountain, a third in-context image 402c depicts horseback riding in an equestrian center, and a fourth in-context image 402d depicts sailing on water. The in-context images 402a-402d represent subject-object interactions in regular contexts.

[0071] At the time of testing, the image interaction recognition model can observe out-of- environment interactions. For example, a first out-of-environment image 404a depicts a snowboarder carving on a sand dune. The regular environment for a snowboarder is a snow mountain. As such, a snowboarder on a sand dune is not in the environment (e.g., out-of-environment). In another example, a second out-of-environment image 404b depicts a skier on water. The regular environment for a skier is on a snow mountain. As such, a skier on water is not in the environment.

[0072] In yet another example, a third out-of-environment image 404c depicts a horseback rider in a city. The regular environment for a horseback rider is in an equestrian center. As such, a horseback rider in a city is not in the environment. In another example, a fourth out-of-environment image 404d depicts a sailboat on ice. The regular environment for a sailboat is water. As such, a sailboat on ice is not in the environment.

[0073] According to aspects of the disclosure, the image interaction recognition model identifies interactions from images regardless of the environment of the images. The model can be trained with in-environment images 402a-402d and the trained model can identify principal-agent interactions in out-of-environment images 404a-404d.

[0074] Figure 5 An image interaction recognition model 500 for classifying principal-agent interactions independent of an environment is illustrated in accordance with aspects of the disclosure. The image interaction recognition model 500 can include a first framework 502, a second framework 504, and a third framework 506. Each framework 502, 504, 506 can be a different subnetwork in a principal-agent interaction classification model

[0075] In one configuration, the first framework 502 identifies a principal and an agent of an interaction while ignoring bystanders. In some cases, a scene includes bystanders that are not observed with the principal-agent pair during training. The bystanders can include humans and objects in various locations of the image (such as the background). The bystanders do not contribute to classifying the principal-agent interaction. That is, the bystanders add unnecessary noise to the classification. Accordingly, the first framework 502 distinguishes the principal and the agent from the bystanders.

[0076] In Figure 5 In the example, the first framework 502 identifies principal-agent interaction regions 502a, 502e and bystander regions 502b, 502c, 502d of an input image. As shown in Figure 5 In the example, the first framework 502 identifies principal-agent interaction regions 502a, 502e and bystander regions 502b, 502c, 502d of an input image. As shown in

[0077] The second framework 504 obtains an environment-agnostic representation of the subject 504b and the object 504c identified in the first framework 502. The subject 504b and the object 504c can be referred to as a subject-object pair (504b, 504c). It is desirable to obtain a representation of the subject-object pair (504b, 504c) that is robust to the disengaged environment scene. In some cases, the environment can modify the subject-object pair. For example, the environment can modify the camera angle, lighting, scene pixels within the subject-object bounding box, and / or visible portions of the subject and / or object. As an example, the environment can occlude (e.g., hide) portions of the subject and / or object.

[0078] To improve robustness, the image interaction recognition model 500 identifies image features that are invariant across different environments for the same interaction. In Figure 5 In the example, the interaction in the subject-object pair (504b, 504c) is horseback riding. In one configuration, the image interaction recognition model 500 is trained on subject-object pair training data (504a, 504d) to identify invariant image features for the horseback riding interaction. For example, the interaction recognition model 500 can be trained to identify one or more features of the horse or the rider that are invariant to the environment. As an example, the feet, the horse tail, or the saddle can be invariant to the environment. The invariant features learned from the subject-object pair training data (504a, 504d) can identify a subject-object representation that is agnostic to the environment in the subject-object pair (504b, 504c).

[0079] The third framework 506 identifies relative weights of the entities (e.g., the subject, the object, and the environment) of the subject-object-environment triple. In one configuration, the image interaction recognition model 500 dynamically adjusts the weights of the subject, the object, and / or the environment. The adjusted weights influence the final classification of the interaction.

[0080] As discussed, portions of the subject or the object can be occluded. For example, the horse in the river can be partially hidden. The third framework 506 identifies whether the environment contributes to the classification of a given subject-object pair and / or whether only the subject and / or only the object contributes to the classification. If the environment corresponds to being in the environment scene, then the environment can contribute to the classification.

[0081] For example, for the first subject-object pair (506a, 506b), an additional weight can be assigned to the first environment 506c (e.g., a horseback riding center) because the first environment 506c contributes to the classification. In contrast, for the second subject-object pair (508a, 508b), the weight assigned to the second environment 508d can be reduced because the second environment 508d (e.g., a snowing environment) does not contribute to the classification of the first subject-object pair (506a, 506b) (e.g., horseback riding). To focus on the first and second environments 506c, 508d, regions 506d, 508c corresponding to the subject and object can be masked from the first and second environments 506c, 508d.

[0082] The relative importance of each entity can be modeled with a weakly supervised submodel that learns to weight subject-object-environment triples given an interaction classification. In one configuration, in- environment interactions are augmented with out-of- environment images manually to improve the process of weighting entities.

[0083] As discussed, input images can be classified according to subject-object pairs (e.g., human-object interaction classes). The classification can be based on a combination of subject, object, and environment features. Figure 6 A model 600 for generating features for subject regions, object regions, and environment regions is illustrated in accordance with aspects of the present disclosure. As shown in Figure 6 The model 600 receives an image 602 depicting a human-object interaction. The input image 602 can be subdivided into subject regions 606a, 606b, object regions 608a, 608b, 608c, and an environment region 610b.

[0084] The image feature function f(.) can generate a D-dimensional feature vector for each subject region 606a, 606b, object region 608a, 608b, 608c, and environment region 610b. In Figure 6 In the example, the image feature function f(.) generates features hi and h2 for the subject regions 606a, 606b, features o1-o3 for the object regions 608a, 608b, 608c, and a feature c for the environment region 610b. Based on image feature learning independent of the environment (see Figure 7 The highest weighted features can be selected for the features of the subject and object.

[0085] In the current example, hi represents the subject and feature o1 represents the object. For clarity, the subject feature hi and the object feature o1 are bolded in Figure 6 The features h2 and o2-o3 represent bystanders. These features can also be referred to as activations of the feature function f(.).

[0086] In one configuration, the image feature function f(.) generates a D-dimensional feature vector via pooling of the region of interest. The pooling can be at the last layer before application to the classifier to obtain region-specific image features.

[0087] Each set of regions (e.g., subject regions 606a, 606b, object regions 608a, 608b, 608c, and context region 610b) can be obtained by masking other regions from image 602. For example, subject regions 606a, 606b can be obtained by masking object regions 608a, 608b, 608c and context region 610b. A region can be masked by setting its value to zero.

[0088] Masking decouples the regions. As such, the model 600 can be prevented from exploiting the co-occurrence of entity regions. As discussed in Figure 6 As shown in FIG. 6B, region 610a corresponding to subject regions 606a, 606b and object regions 608a, 608b, 608c is masked from context region 610b. When determining features for context region 610b, the masked region 610a in context region 610b can prevent model 600 from exploiting subject regions 606a, 606b and / or object regions 608a, 608b, 608c. The improved decoupling provides subject regions 606a, 606b and object regions 608a, 608b, 608c with environment-agnostic representations because the model does not observe context region 610b.

[0089] Figure 7 An example 700 of environment-agnostic image feature learning is illustrated in accordance with aspects of the present disclosure. In one configuration, for each image, features from the model are aggregated to obtain similar representations across different images of the same interaction. Aspects of the present disclosure discuss human-object interactions. Aspects of the present disclosure can also apply to other interacting parties (e.g., other objects or living beings).

[0090] As discussed, the model produces a D-dimensional feature vector for each subject region (h i ), object region (o i ), and context region (c) per image. That is, D is based on the number of subject regions, object regions, and context regions. In an image, it can be unclear which region(s) correspond to the interaction as opposed to bystanders and / or background objects. The obtained representations can be sensitive to environment-specific transformations of the subject and object regions. Environment-specific transformations can include, for example, changes in viewpoint, pose, and / or lighting. As such, it is desirable to weight the subject or object features to improve subject and / or object identification.

[0091] In one configuration, subnetwork g(.) receives subject region (h i) and object region (o i The subnetwork g(.) generates an Nx1-dimensional vector (where N is the number of detections) representing the contribution of each region to the final response. The weighted aggregation of all subject or object features within the input can be determined based on this Nx1-dimensional vector.

[0092] like Figure 7 As shown, from Figure 6 The subject regions 606a and 606b of the input image 602, depicting humans, can be fed into the subnetwork g(.) 706. Subnetwork g(.) 706 generates a weighted sum of subject regions 606a and 606b. The training image 702 (e.g., the target image) is sampled from the training set. The training image 702 depicts the same interaction as subject regions 606a and 606b and also depicts the same subject or object. For example, both the training image 702 and the subject regions 606a and 606b depict humans, and the interaction in one of the subject regions 606a is riding a horse.

[0093] Image feature function f(.)( Figure 7 (not shown in the image) and subnetwork g(.) 706 can be applied to training image 702. Training image 702 may have an environment different from that of subject regions 606a, 606b, resulting in different subject and object appearances. Still, training image 702 and subject regions 606a, 606b share the same subjects and objects.

[0094] During training, the Euclidean loss is determined across the outputs 704a and 704b of each subnetwork g(.)706. Output 704 can be a regularized representation. This loss achieves a weighted aggregation of subject and object activations across different environments with the same interaction. That is, subnetwork g(.)706 is trained to assign higher weights (e.g., probabilities) to the first region 606a based on its similarity to the training image 702. This similarity is determined from the similarity between the first region 606a and the subject and interaction of the training image 702. By adjusting the weights of each region, subject-object similarity representations can be identified across different environments. This provides environment-independent representations.

[0095] based on Figure 7 In weakly supervised training, the model learns to appropriately reduce the weight of bystanders. For example, ... Figure 7 As shown, bystanders in the second subject region 606b are assigned a weight of 0.2. Conversely, interacting subjects in the first subject region 606a are assigned a weight of 0.8. The training image 702 is a ground reality image and has a weight of 1. The weights are adjusted so that the representation of the first subject region 606a is similar to the representation of the training image 702.

[0096] That is, while the environment of the first subject region 606a is different from the environment of the training image 702, the first subject region 606a and the subject of the training image 702 are similar (e.g., a person on a horse). The bystander of the second subject region 606b is not similar to the subject of the training image 702. Thus, through the comparison loss, the network learns to assign a lower weight to the bystander of the second subject region 606b. Figure 7 Weakly supervised training can also be applied to objects.

[0097] In one aspect, the subnetwork g(.) 706 is implemented as a one layer neural network of size Dx 1. The output of the subnetwork g(.) 706 modulates the subject or object representation of size (N x D) with an inner product ((N x D) x (D x 1)). The inner product can be a pooling mean across the first dimension to obtain a D x 1 dimensional feature representation of the input image for the subject object pair.

[0098] A softmax function can be applied to the weight activations of the subject or object representation. Through the application of the function activations, the weight activations of the subject or object can be set to 1. For example, the weighted activations can assign a higher value to the true subject-object pair of the interaction compared to the values associated with other subjects (e.g., bystanders) of the object.

[0099] Entity reweighting can improve the environment-agnostic classification. For the environment-agnostic classification, the subject, object, and environment are jointly detected. The model obtains a representation of the human, object, and environment regions, and aggregates the subject or object representation to make them similar across different scenes. This results in a three-dimensional (3D) activation vector for each image, summarizing the subject, object, and environment appearance in the input.

[0100] In one aspect of the disclosure, assuming the observations can be subject-object-environment triplets, the informative input is dynamically determined for the final classification. That is, the model dynamically reweights the subject, object, and environment to obtain the final classification decision for the input image.

[0101] Figure 8 A framework 800 for identifying subject-object-environment image features for classification is illustrated in accordance with aspects of the present disclosure. The framework 800 identifies relative weights of subject-object-environment triplets 802, 804 for classification. A (3, 1) dimensional vector can represent the relative importance of the subject-object-environment triplets 802, 804 for classification. The framework 800 can be implemented prior to the classification procedure.

[0102] As Figure 8The subject-object-environment triplets 802, 804 are generated by an entity weighting module h(.) 810, shown in FIG. 8. In one configuration, the entity weighting module h(.) 810 receives the three 3xD dimensional image features of the subject-object-environment triplet 802, 804 and produces a 3xl dimensional vector representing the relative importance of the subject, object, and environment for classification. Since the model observes subject-objects coupled with a regular environment, the model can fail to learn to suppress the contribution of the environment when the environment is not informative for the observed interaction (e.g., a snow mountain for a horse riding interaction). To suppress the contribution of non-informative environments, the model relies on the mining of unexpected environments.

[0103] In the mining of unexpected environments, given the image features of a subject-object (e.g., a horse rider and a horse), the original environment is replaced by sampling environment features from a different interaction image (e.g., an image of a snowboarder). The entity weighting module h(.) 810 produces an input vector for the classification module k(.) 812, such that when the subject, object, and environment are combined, the correct classification is obtained via the classification module k(.) 812.

[0104] The mining of unexpected environments is performed during training to simulate unexpected subject-object-environment triplets. As a result of the training, the model learns the improved concept of dynamic weighting. To this end, the entity weighting module h(.) 810 learns to identify subject-object pairs that are detached from the environment, and assigns a lower weight to the environment compared to a regular environment (e.g., a horse racing club). Figure 8 Example values 806, 808 for the importance of subject-object-environment for classification accuracy are shown in FIG. 8. This final re-weighted representation is input to a classifier to label the input image.

[0105] For example, a first subject-object-environment triad 802 of riding a horse at an equestrian center can have relative weights 806 corresponding to subject (0.3), object (0.4), and environment (0.3). A second subject-object-environment triad 804 of riding a horse on a snow-covered mountain can have relative weights 808 corresponding to subject (0.5), object (0.4), and environment (0.1). In this example, the horse on the snow-covered mountain can be considered to be out of context. Thus, the environment for the snow-covered mountain ride (relative weight 0.1) is deemphasized relative to the equestrian center ride, where the relative weight of the environment is 0.3. The emphasis for the snow-covered mountain ride is placed on the subject of the second subject-object-environment triad 804. For example, the relative weight of the subject of the second subject-object-environment triad 804 (0.5) is increased relative to the relative weight of the subject of the first subject-object-environment triad 802 (0.3). A Softmax function is applied so that the sum of all importance values is 1 (e.g., 0.5 + 0.4 + 0.1 = 1).

[0106] Figure 9 A method 900 according to an aspect of the disclosure is illustrated. As shown in Figure 9 The neural network locates the subject and object of the human-object interaction in the image (block 902). In one configuration, the neural network ignores bystanders and / or background objects. The image can be an out-of-context image. Additionally, the subject can be identified as a human.

[0107] In one configuration, the neural network receives only the subject image region, only the object image region, and only the environment image region. The neural network can generate image features corresponding to each of the subject, object, and environment regions. The neural network can be trained to represent the subject and object in a manner that is independent of the environment, and locate the subject and object based on this learning. That is, the neural network can obtain a subject representation and an object representation of the image that are independent of the environment.

[0108] As shown in Figure 9 The neural network determines relative weights of the subject, object, and environment regions for classification (block 904). In an optional configuration, the neural network masks only the subject image region and only the object image region to obtain only the environment image region; masks only the subject image region and only the environment image region to obtain only the object image region; and masks only the environment image region and only the object image region to obtain only the subject image region. The relative weights of the subject, object, and environment regions can be determined based on relative importance to the human-object interaction classification, which is determined based on the image features.

[0109] As shown in Figure 9As shown in the middle, the neural network classifies the agent-object interaction based on the classification of the weighted representation of the agent, the weighted representation of the object, and the weighted representation of the environment region (block 906). That is, the entity weighting module produces an input weight vector for the agent, the object, and the environment region. The classification module classifies the weighted environment- independent object, the weighted environment-independent agent, and the weighted environment region, where the weights are obtained from the entity weighting module. When the weighted environment-independent object, the weighted environment-independent agent, and the weighted environment region are combined, the correct classification is obtained via the classification module.

[0110] The various operations of methods described above can be performed by any suitable means depending on the functionality of the means. These means can include various hardware and / or software components such as a circuit, an application specific integrated circuit (ASIC), or processor. Generally, where there are operations illustrated in figures, those operations can have corresponding counterpart means-plus-function components with similar numbering.

[0111] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure), ascertaining and the like. Additionally, “determining” can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory) and the like. Furthermore, “determining” can include resolving, selecting, choosing, establishing and the like.

[0112] As used herein, the term “at least one of’ a list of items refers to any combination of one or more of the items in the list. As an example, “at least one of a, b, or c” is intended to mean: a; b; c; a-b; a-c; b-c; and a-b-c.

[0113] The various illustrative logical blocks, modules, and circuits described in connection with the disclosure can be implemented or performed with a general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the processor can be any commercially available processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0114] The steps of a method or algorithm described in connection with the present disclosure can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in any form of storage medium that is known in the art. Some examples of storage media that can be used include random access memory (RAM), read only memory (ROM), flash memory, erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, a hard disk, a removable disk, a CD-ROM, and so forth. A software module can comprise a single instruction, or many instructions, and can be distributed over several different code segments, 0 different programs, and across multiple storage media. A storage medium can be coupled to a processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor.

[0115] The methods disclosed herein comprise one or more steps or actions for achieving the described method. The method steps and / or actions can be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims.

[0116] The functions described can be implemented in hardware, software, firmware or any combination thereof. If implemented in hardware, an example hardware configuration can include a processing system in a device. The processing system can be implemented with a bus architecture. The bus can include any number of interconnecting buses and bridges depending on the specific application of the processing system and the overall design constraints. The bus can link together various circuits including processing circuitry, machine-readable medium, and bus interface. The bus interface can be used to connect a network adapter to the processing system via the bus. The network adapter can be used to implement signal processing functionality. For certain aspects, a user interface (e.g., keypad, display, mouse input, joystick, etc.) can also be connected to the bus. The bus can also link various other circuits such as timing sources, peripherals, voltage regulators, power management circuits, and the like, which are well known in the art, and therefore, will not be described any further.

[0117] The processor can be responsible for managing a bus and general processing, including the execution of software stored on the machine-readable media. The processor can be implemented with one or more general-purpose and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuitry that can execute software. Software shall be construed broadly to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Machine-readable media can include, by way of example, random access memory (RAM), flash memory, read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard drives, or any other suitable storage medium, or any combination thereof. The machine-readable media can be embodied in a computer- program product. The computer-program product can comprise packaging materials.

[0118] In a hardware implementation, the machine-readable media can be part of the processing system separate from the processor. However, as would be appreciated by those skilled in the art, the machine-readable media, or any portion of it, can be external to the processing system. As an example, the machine-readable media can include a transmission line, a carrier wave modulated by the data, and / or a computer product separate from the device, all

[0119] The processing system can be configured as a general- purpose processing system with one or more microprocessors providing processor functionality and external memory providing at least a portion of the machine-readable media all linked together with other supporting circuitry through an external bus architecture. Alternatively, the processing system can include one or more neuromorphic processors for implementing the neuron models and nervous system models described herein. As another alternative, the processing system can be implemented with an application specific integrated circuit (ASIC) with the processor, bus interface, user interface, supporting circuitry, and at least a portion of machine-readable media integrated into a single chip, or with one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuitry that can perform the various functionality described throughout this disclosure. Those skilled in the art will recognize how best to implement the described functionality for the processing system, depending on the particular application and the overall design constraints imposed on the overall system.

[0120] The machine-readable media can comprise a number of software modules. The software modules include instructions that, when executed by the processor, cause the processing system to perform various functions. The software modules can include a transmission module and a receiving module. Each software module can reside in a single storage device or be distributed across multiple storage devices. By way of example, a software module can be loaded into RAM from a hard drive when a triggering event occurs. During execution of the software module, the processor can load some of the instructions into cache to increase access speed. One or more cache lines can then be loaded into a general register file for execution by the processor. When referring to the functionality of a software module below, it will be understood that such functionality is implemented by the processor when executing instructions from that software module. Furthermore, it should be appreciated that aspects of the present disclosure result in improvements to the functioning of the processor, computer, machine, or other system implementing such aspects.

[0121] If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage media can be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Thus, in some aspects computer readable medium can comprise non-transitory computer readable medium (e.g., tangible media). In addition, for other aspects computer readable medium can comprise transitory computer readable medium (e.g., signals). Combinations of the above should also be included within the scope of computer readable media.

[0122] Thus, certain aspects can comprise a computer program product for performing the operations presented herein. For example, such a computer program product can comprise a computer readable medium having instructions stored thereon (and / or encoded therein), the instructions being executable by one or more processors to

[0123] For example, such a device can be coupled to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, various methods described herein can be provided via a storage means (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or floppy disk, etc.), such that a user terminal and / or base station can obtain the various methods upon coupling or providing the storage means to the device. Moreover, any appropriate programming language can be used to implement the routines of the above described aspects.​

[0124] It will be understood that the claims are not limited to the precise configuration and components illustrated above. Various modifications, changes, and adaptations will be apparent to others skilled in the art with the benefit of this disclosure.

Claims

1. A method for processing an image, comprising: locating a subject and an object of a human-object interaction in the image, wherein the image is decontextualized, wherein the decontextualized image represents the human-object interaction in an atypical context; determining a scene context, the scene context indicating a semantic category of a scene in which the human-object interaction occurs; determining a relative weight of the subject, a relative weight of the object, and a relative weight of an environmental region corresponding to the scene context to classify the human-object interaction based on whether a human-object context associated with the human-object interaction corresponds to the scene context, wherein the environmental region is a region of the scene in the image, wherein the image is subdivided into at least one subject region comprising the subject, at least one object region comprising the object, and the environmental region, wherein determining the relative weight of the environmental region comprises assigning a lower weight to the environmental region compared to a typical context; and classifying the human-object interaction based on a classification of a weighted representation of the subject, a weighted representation of the object, and a weighted representation of the environmental region.

2. The method of claim 1, wherein the image represents a human-object interaction having a number of examples in a training set that is less than a threshold.

3. The method of claim 1, wherein the subject is identified as a human.

4. The method of claim 1, further comprising: learning to represent the subject and the object in a context-agnostic manner; and locating the subject and the object based on the learning.

5. The method of claim 1, further comprising: receiving, at a convolutional neural network, a subject-only image region, an object-only image region, and an environment-only image region; and generating, by the convolutional neural network, image features corresponding to each subject, object, and environmental region.

6. The method of claim 5, further comprising: masking the subject-only image region and the object-only image region to obtain the environment-only image region; masking the subject-only image region and the environment-only image region to obtain the object-only image region; and masking the environment-only image region and the object-only image region to obtain the subject-only image region.

7. The method of claim 5, further comprising: determining the relative weight of the subject, the relative weight of the object, and the relative weight of the environmental region based on a relative importance of classifying human-object interactions, the relative importance being determined based on the image features.

8. An apparatus for processing an image, comprising: a memory; and at least one processor coupled to the memory, the at least one processor configured to: locate a subject and an object of a human-object interaction in the image, wherein the image is decontextualized, wherein the decontextualized image represents the human-object interaction in an atypical context; determine a scene context, the scene context indicating a semantic category of a scene in which the human-object interaction occurs; determining a relative weight of the subject, a relative weight of the object, and a relative weight of an environmental region corresponding to the scene environment to classify the subject-object interaction based on whether a subject-object environment associated with the subject-object interaction corresponds to the scene environment, wherein the environmental region is a region of a scene in the image, wherein the image is subdivided into at least one subject region including the subject, at least one object region including the object, and the environmental region, wherein the at least one processor configured to determine the relative weight of the environmental region is further configured to assign a lower weight to the environmental region compared to a regular environment; and classifying the subject-object interaction based on the classification of the weighted representation of the subject, the weighted representation of the object, and the weighted representation of the environmental region.

9. The apparatus of claim 8, wherein the image represents a subject-object interaction having a number of examples in a training set that is less than a threshold.

10. The apparatus of claim 8, wherein the subject is identified as a human.

11. The apparatus of claim 8, wherein the at least one processor is further configured to: learn to represent the subject and the object in an environment-agnostic manner; and locate the subject and the object based on the learning.

12. The apparatus of claim 8, wherein the at least one processor is further configured to: receive, at a convolutional neural network, a subject-only image region, an object-only image region, and an environment-only image region; and generate, by the convolutional neural network, image features corresponding to each subject, object, and environmental region.

13. The apparatus of claim 12, wherein the at least one processor is further configured to: mask the subject-only image region and the object-only image region to obtain the environment-only image region; mask the subject-only image region and the environment-only image region to obtain the object-only image region; and mask the environment-only image region and the object-only image region to obtain the subject-only image region.

14. The apparatus of claim 12, wherein the at least one processor is further configured to determine the relative weight of the subject, the relative weight of the object, and the relative weight of the environmental region based on a relative importance of classifying subject-object interactions, the relative importance being determined based on the image features.

15. A device for processing an image, comprising: means for locating a subject and an object of a subject-object interaction in the image, wherein the image is decontextualized, wherein the decontextualized image represents the subject-object interaction in a non-regular environment; means for determining a scene environment, the scene environment indicating a semantic class of a scene in which the subject-object interaction occurs; apparatus for determining a relative weight of the subject, a relative weight of the object, and a relative weight of an environmental region corresponding to the scene environment to classify the subject-object interaction based on whether a subject-object environment associated with the subject-object interaction corresponds to the scene environment, wherein the environmental region is a region of a scene in the image, wherein the image is subdivided into at least one subject region including the subject, at least one object region including the object, and the environmental region, wherein the apparatus for determining the relative weight of the environmental region comprises an apparatus for assigning a lower weight to the environmental region compared to a regular environment; and an apparatus for classifying the subject-object interaction based on a classification of a weighted representation of the subject, a weighted representation of the object, and a weighted representation of the environmental region.

16. The apparatus of claim 15, wherein the image represents a subject-object interaction having a number of examples in a training set that is less than a threshold.

17. The apparatus of claim 15, wherein the subject is identified as a human.

18. The apparatus of claim 15, further comprising: an apparatus for learning to represent the subject and the object in an environment-agnostic manner; and an apparatus for localizing the subject and the object based on the learning.

19. The apparatus of claim 15, further comprising: an apparatus for receiving, at a convolutional neural network, a subject-only image region, an object-only image region, and an environment-only image region; and an apparatus for generating, by the convolutional neural network, image features corresponding to each subject, object, and environment region.

20. The apparatus of claim 19, further comprising: an apparatus for masking the subject-only image region and the object-only image region to obtain the environment-only image region; an apparatus for masking the subject-only image region and the environment-only image region to obtain the object-only image region; and an apparatus for masking the object-only image region and the environment-only image region to obtain the subject-only image region.

21. The apparatus of claim 19, further comprising: an apparatus for determining a relative weight of the subject, a relative weight of the object, and a relative weight of the environmental region based on a relative importance of a classification of a subject-object interaction, the relative importance being determined based on the image features.

22. A non-transitory computer-readable medium having recorded thereon program code for processing an image, the program code being executed by a processor and comprising: program code for localizing a subject and an object of a subject-object interaction in the image, wherein the image is decontextualized, wherein the decontextualized image represents the subject-object interaction in a non-regular environment; program code for determining a scene environment, the scene environment indicating a semantic class of a scene in which the subject-object interaction occurs; program code for determining a relative weight of the subject, a relative weight of the object, and a relative weight of an environmental region corresponding to the scene environment to classify the subject-object interaction based on whether a subject-object environment associated with the subject-object interaction corresponds to the scene environment, wherein the environmental region is a region of a scene in the image, wherein the image is subdivided into at least one subject region including the subject, at least one object region including the object, and the environmental region, wherein the program code for determining the relative weight of the environmental region includes program code for assigning a lower weight to the environmental region compared to a regular environment; and program code for classifying the subject-object interaction based on a classification of a weighted representation of the subject, a weighted representation of the object, and a weighted representation of the environmental region.

23. The non-transitory computer-readable medium of claim 22, wherein the image represents a subject-object interaction having a number of examples in a training set that is less than a threshold.

24. The non-transitory computer-readable medium of claim 22, wherein the subject is identified as a human.

25. The non-transitory computer-readable medium of claim 22, wherein the at least one processor is further configured to: learn to represent the subject and the object in an environment-agnostic manner; and locate the subject and the object based on the learning.

26. The non-transitory computer-readable medium of claim 22, wherein the at least one processor is further configured to: receive, at a convolutional neural network, a subject-only image region, an object-only image region, and an environment-only image region; and generate, by the convolutional neural network, image features corresponding to each subject, object, and environment region.

27. The non-transitory computer-readable medium of claim 26, wherein the at least one processor is further configured to: mask the subject-only image region and the object-only image region to obtain the environment-only image region; mask the subject-only image region and the environment-only image region to obtain the object-only image region; and mask the environment-only image region and the object-only image region to obtain the subject-only image region.

28. The non-transitory computer-readable medium of claim 26, wherein the at least one processor is further configured to determine the relative weight of the subject, the relative weight of the object, and the relative weight of the environmental region based on a relative importance of a subject-object interaction classification, the relative importance being determined based on the image features.

Citation Information

Patent Citations

  • Subject-object interaction recognition model

    US11481576B2