Interaction Relationship Recognition Method and Apparatus

By using an interaction relationship recognition model, the problem of accuracy in recognizing the interaction relationships between people and between people and things in sales scenarios is solved, enabling efficient transaction behavior analysis and abnormal transaction early warning in complex environments.

CN121545106BActive Publication Date: 2026-04-21BEIJING WENAN INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING WENAN INTELLIGENT TECH CO LTD
Filing Date
2026-01-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify interactions between people and between people and objects in sales scenarios, resulting in low accuracy in transaction behavior recognition. In particular, under the influence of factors such as changes in lighting and human obstruction in complex environments, existing methods fail to delve into the interactions between people and objects in transactions.

Method used

An interactive relationship recognition model is adopted, including a feature extraction module, a subject-object pair matching module, and a relationship recognition module. Feature information is extracted from video frames to generate subject-object pair information and action relationship categories, and to identify the interactive relationships between people and between people and objects.

Benefits of technology

It improves the accuracy of transaction behavior identification, can accurately identify the interaction between people and between people and things in complex environments, supports real-time analysis and provides early warning of abnormal transactions, and reduces losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545106B_ABST
    Figure CN121545106B_ABST
Patent Text Reader

Abstract

This invention provides an interaction relationship recognition method and apparatus. A specific embodiment of the method includes: acquiring a video of a target area within a sales venue; inputting video frames from the video into an interaction relationship recognition model to obtain the interaction relationship between a subject and an object in the video frames, where the subject includes a human body and the object includes both human bodies and objects; the interaction relationship recognition model includes a feature extraction module, a subject-object pair matching module, and a relationship recognition module; the interaction relationship recognition model performs the following recognition operations on the video frames: inputting the video frames into the feature extraction module, which extracts feature information from the video frames; inputting learnable query vectors and feature information into the subject-object pair matching module, which outputs multiple subject-object pair information; inputting each subject-object pair information and feature information into the relationship recognition module, which generates the action relationship category between the subject and object in each subject-object pair information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to an interaction relationship recognition method and apparatus. Background Technology

[0002] In the field of transaction behavior recognition in sales scenarios (e.g., retail scenarios), past research and applications have mainly focused on statistical analysis of transaction data and behavior judgment based on fixed rules. With the rise of artificial intelligence technology, some studies have introduced machine learning algorithms, such as Support Vector Machines (SVM) and decision trees, to classify and identify transaction behaviors, achieving classification of transaction behaviors through model training. However, these traditional methods heavily rely on manual feature extraction, which has significant limitations in complex and ever-changing sales scenarios. It is difficult for humans to comprehensively and accurately extract features that reflect the essence of various transaction behaviors, resulting in a significant decrease in the accuracy of model recognition.

[0003] In recent years, some studies have utilized computer vision technology to capture customer behavior in stores using cameras, and to assist in identifying transaction behavior through movement trajectories and areas of stay. However, in practical applications, complex environments, lighting changes, and human occlusion significantly impact the accuracy of visual data collection and analysis. Furthermore, these methods typically focus only on the overall external manifestation of customer behavior, failing to delve into the interaction between people and objects in actual transactions, such as the interaction between people and transaction tools. The interactions between people and between people and objects in actual transactions are crucial for transaction behavior analysis. Therefore, how to identify the interactions between people and between people and objects in sales scenarios is a pressing issue that needs to be addressed. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide an interaction relationship identification method and apparatus to eliminate or improve one or more defects existing in the prior art.

[0005] According to a first aspect, an interaction relationship recognition method is provided, the method comprising the following steps: acquiring a video of a target area within a sales venue, wherein the target area includes the area where the cashier is located; inputting video frames from the video into a pre-trained interaction relationship recognition model to obtain the interaction relationship between a subject and an object in the video frames, wherein the subject includes a human body, and the object includes a human body and an object; the interaction relationship recognition model includes a feature extraction module, a subject-object pair matching module, and a relationship recognition module; the interaction relationship recognition model performs the following recognition operations on the video frames: inputting the video frames into the feature extraction module, which extracts feature information from the video frames; inputting a learnable query vector and the feature information into the subject-object pair matching module, which outputs multiple subject-object pair information, wherein each subject-object pair information includes a subject detection box, an object detection box, the category corresponding to the object detection box, and a relationship confidence score; inputting each subject-object pair information and the feature information into the relationship recognition module, which generates the action relationship category of the subject and object in each subject-object pair information.

[0006] According to a second aspect, an interaction relationship recognition device is provided, comprising: an acquisition unit for acquiring video of a target area within a sales venue, wherein the target area includes the area where the cashier is located; and a recognition unit for inputting video frames from the video into a pre-trained interaction relationship recognition model to obtain the interaction relationship between a subject and an object in the video frames, wherein the subject includes a human body, and the object includes a human body and objects; the interaction relationship recognition model includes a feature extraction module, a subject-object pair matching module, and a relationship recognition module; and the interaction relationship recognition model performs the following recognition operations on the video frames. The video frame is input into the feature extraction module, which extracts feature information from the video frame. The learnable query vector and the feature information are input into the subject-object pair matching module, which outputs multiple subject-object pairs, each of which includes a subject detection box, an object detection box, the corresponding category of the object detection box, and a relationship confidence score. The subject-object pairs and the feature information are input into the relationship recognition module, which generates the action relationship category between the subject and the object in each subject-object pair.

[0007] According to a third aspect, a computing device is provided, including a processor, a memory, and a computer program / instructions stored in the memory, the processor being configured to execute the computer program / instructions, and when the computer program / instructions are executed, the computing device performing the steps of the method as described in any of the first aspects.

[0008] According to a fourth aspect, a computer-readable storage medium is provided that stores a computer program / instructions thereon, which, when executed by a processor, implement the steps of the method as described in any of the first aspects.

[0009] The interaction relationship recognition method and apparatus provided in the embodiments of this specification can acquire video of a target area within a sales venue and input video frames from the video into a pre-trained interaction relationship recognition model to obtain the interaction relationship between the subject and object in the video frames. Here, the subject can include a human body, and the object can include both a human body and an object. The interaction relationship recognition model can include a feature extraction module, a subject-object pair matching module, and a relationship recognition module. Based on this, the interaction relationship recognition model can perform the following recognition operations on the video frames: First, the video frames are input into the feature extraction module, which extracts feature information from the video frames. Then, the learnable query vector and feature information are input into the subject-object pair matching module, which outputs multiple subject-object pair information, where each subject-object pair information includes a subject detection box, an object detection box, the category corresponding to the object detection box, and the relationship confidence. Then, each subject-object pair information and feature information are input into the relationship recognition module, which generates the action relationship category between the subject and object in each subject-object pair information. Thus, the interaction relationship recognition between people and between people and objects in a sales scenario is realized.

[0010] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.

[0011] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description

[0012] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to limit the scope of the invention.

[0013] Figure 1 A flowchart of an interaction relationship identification method according to one embodiment is shown;

[0014] Figure 2 This diagram illustrates an example of data annotation for an image.

[0015] Figure 3A schematic diagram is shown illustrating one application scenario in which the embodiments of this specification can be applied;

[0016] Figure 4 A schematic block diagram of an interaction relationship recognition device according to one embodiment is shown. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.

[0018] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.

[0019] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0020] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0021] It is understood that the ordinal numbers such as "first" and "second" mentioned in this specification are only used to distinguish multiple objects of the same or different categories (such as components, steps, parameters, etc.), and do not indicate the priority, importance or order relationship between objects, nor do they constitute a limitation on the technical features.

[0022] As mentioned earlier, the interactions between people and between people and things in actual transactions are extremely important for the identification and analysis of transaction behavior. However, existing methods of transaction behavior identification often neglect the acquisition of person-to-person and person-to-thing relationships. Although computer vision technology can capture the position and actions of people, apart from the most basic and standard transaction actions, it is difficult to deeply analyze the substantive relationships between people for most actions, making it impossible to accurately determine whether a transaction has occurred. Abnormal transactions also become more difficult to identify accurately. The main reason for this situation is the lack of identification of person-to-person and person-to-thing relationships, which makes transaction behavior identification lack crucial contextual information.

[0023] Therefore, embodiments of this specification provide an interaction relationship recognition method that can identify the interaction relationships between people and between people and objects in a sales scenario.

[0024] Please see Figure 1 , Figure 1A flowchart of an interaction relationship identification method according to one embodiment is shown. It will be understood that this method can be executed by any device, apparatus, platform, or cluster of devices with computing and processing capabilities. Figure 1 As shown, the interaction relationship identification method may include the following steps 100 and 200, specifically:

[0025] Step 100: Obtain video of the target area within the sales venue.

[0026] In this embodiment, video of a target area within a sales location (e.g., a shopping mall, supermarket, etc.) can be acquired. Here, the target area may include the area where the checkout counter is located. For example, video of the checkout counter area can be captured by a camera installed diagonally above it, to capture video of staff and customers, as well as handheld mobile phones, POS (Point of Sale) machines, cash, etc., with as few obstructions as possible.

[0027] Step 200: Input the video frames from the video into the pre-trained interaction relationship recognition model to obtain the interaction relationship between the subject and the object in the video frames.

[0028] In this embodiment, video frames, such as all video frames or frames extracted (e.g., one frame extracted every preset second), can be input into a pre-trained interaction relationship recognition model. This model can output the interaction relationship between the subject and object in the video frame. The subject can include the human body, and the object can include both the human body and objects. The interaction relationship recognition model can include a feature extraction module, a subject-object pair matching module, and a relationship recognition module. The feature extraction module can be used to extract feature information from the video frames. For example, it can include various image feature extraction networks, such as convolutional neural networks, MobilNetV4 (MobileNet Version 4), etc. The subject-object pair matching module can be used to generate subject-object pair information. This module can be various machine learning networks, such as convolutional neural networks, Transformer-based networks, etc. The relationship recognition module can be used to generate the action relationship categories between the subject and object in the subject-object pair information. This module can also be various machine learning networks, such as convolutional neural networks, Transformer-based networks, etc.

[0029] In some examples, the feature extraction module may include a feature extraction network and an encoder, the subject-object pair matching module may include a subject-object pair matching decoder, and the relationship recognition module may include an action-relationship decoder. For instance, the encoder may include a Transformer encoder, and the subject-object pair matching decoder and relationship recognition module may include a Transformer decoder. The Transformer encoder and Transformer decoder consist of multiple stacked layers.

[0030] In some examples, the interaction relationship recognition model can be trained through steps one through three, specifically:

[0031] Step 1), obtain the training sample set.

[0032] In this example, each training sample in the training sample set can include sample images and corresponding annotation data. The annotation data can include annotation data corresponding to each subject-object pair, and can include subject detection boxes, object detection boxes, the corresponding sample category, sample relationship confidence, and sample action relationship category. In this example, the subject can include a human body, and the object can include both human bodies and objects. As an example, the subject and object detection boxes can include center point coordinates (x, y), width w, and height h. The sample category corresponding to the object detection box can refer to the category of the object, which can include human body, mobile phone, POS machine, cash, shopping bag, etc. The sample relationship confidence indicates the probability that there is a relationship between the subject corresponding to the subject detection box and the object corresponding to the object detection box. For example, the sample relationship confidence can be a value between 0 and 1, with a higher value indicating a higher probability of a relationship between the subject and object. The sample action relationship category refers to the action relationship category between the subject and object.

[0033] In this example, a large amount of transaction video data from retail scenarios can be collected. After frame extraction, the images are labeled with bounding boxes for people and objects, as well as for human-object and human-human relationships that exhibit interaction. By using labeled data from a large amount of real-world retail transaction behavior, the model's ability to detect human-object interactions in retail scenarios can be significantly improved.

[0034] Please continue reading Figure 2 , Figure 2 This diagram illustrates an example of data annotation for an image. Figure 2In the example shown, detection boxes can be labeled for each human body, as well as for objects such as mobile phones, POS machines, and shopping bags. The labeled action relationships can be "take" for human-phone pairs, "scan" for human-POS machine pairs, and "transaction" for human-human pairs. This is understandable. Figure 2 The data labeled in the image is merely illustrative and not a limitation on the data to be labeled. In practice, various data can be labeled in the image according to actual needs.

[0035] Step 2) Input the sample image into the interaction relationship recognition model, which will output prediction information.

[0036] In this example, the prediction information output by the interaction relationship recognition model can include the prediction information corresponding to each prediction subject-predicted object pair. The prediction information can include the prediction subject detection box, the prediction object detection box, and the prediction action relationship category between the prediction subject and the prediction object.

[0037] Step 3) Based on the difference between the predicted information and the labeled information, adjust the parameters of the interaction relationship recognition model.

[0038] In this example, a preset loss function can be used to calculate the difference between the predicted information output by the interaction relationship recognition model and the labeled data. Then, based on the calculated difference, the parameters of the interaction relationship recognition model can be adjusted. For example, the BP (Back Propagation) algorithm or the SGD (Stochastic Gradient Descent) algorithm can be used to adjust the model parameters.

[0039] In some examples, a phased training approach can be used to train the interaction relationship recognition model. The first phase adjusts the parameters of the feature extraction module and the subject-object pair matching module, while the second phase adjusts the parameters of the feature extraction module, the subject-object pair matching module, and the relationship recognition module. Different learning rates can be set for different parameters in the second phase. Specifically, the parameters of the interaction relationship recognition model can be divided into two groups: Group 1 includes the parameters of the feature extraction module and the subject-object pair matching module (the parameters of Group 1 have already been trained in the first phase and only need fine-tuning in the second phase); Group 2 includes the parameters of the relationship recognition module. In the second phase, a smaller learning rate can be set for the parameters of Group 1 (e.g., 0.1 times the initial learning rate), and a larger learning rate can be set for the parameters of Group 2 (e.g., the initial learning rate). In this example, step three above can specifically include steps 1) and 2), where step 1) corresponds to the first phase of the phased training, and step 2) corresponds to the second phase of the phased training. Specifically:

[0040] Step 1): Based on the difference between the first predicted information and the labeled information output by the subject-object pair matching module, adjust the parameters of the feature extraction module and the subject-object pair matching module.

[0041] In this example, the first prediction information may include the predicted subject detection box, the predicted object detection box, the prediction category corresponding to the predicted object detection box, and the prediction relationship confidence.

[0042] Step 2), based on the difference between the predicted action relationship category output by the relationship recognition module and the sample action relationship category, adjust the parameters of the feature extraction module, the subject-object pair matching module, and the relationship recognition module.

[0043] In this example, in the second stage, the learning rates of the parameters for the feature extraction module and the subject-object pair matching module can differ from those of the parameters for the relationship recognition module. The learning rates of the parameters for the feature extraction module and the subject-object pair matching module can be smaller than those of the parameters for the relationship recognition module. In some approaches, a dynamic weighting mechanism can be used during training to address the uneven distribution of categories in the data. The model can be trained first using a pre-defined loss function, then the parameters of the feature extraction module can be fixed, and the subject-object pair matching module and the relationship recognition module can be trained using a loss function with a small learning rate and a dynamic weighting strategy. The dynamic weighting is applied to the object category and the action relationship category, respectively.

[0044] The above describes the training process of the interaction relationship recognition model. The resulting interaction relationship recognition model can process the input video frames to obtain the interaction relationship between the subject and the object in the video frame.

[0045] Next, in this embodiment, the interaction relationship recognition model can perform the following recognition operation steps 201 to 203 on the video frame, specifically:

[0046] Step 201: Input the video frame into the feature extraction module, which then extracts feature information from the video frame.

[0047] In this embodiment, the interaction relationship recognition model can input video frames into the feature extraction module, which will then extract features from the video frames to obtain their feature information.

[0048] In some examples, the feature extraction module described above may include an encoder, for example, a Transformer encoder. Based on this, step 201, in which the feature extraction module extracts feature information from the video frame, may include steps 1) to 3), specifically:

[0049] Step 1) Extract features from the video frames to obtain feature sequences, and add position encoding information to the feature sequences to obtain feature sequences with position information.

[0050] In this example, convolutional neural networks (CNNs), such as VGG (Visual Geometry Group) and ResNet (Residual Network), can be used to extract features from video frames and obtain feature maps. Here, the feature maps are three-dimensional; for example, the dimension of the feature map could be... ,in, It can represent the number of channels. and These can represent height and width, respectively. Since the Transformer decoder expects a sequence as input, i.e., a two-dimensional matrix, for example, its shape could be... ,in, It can represent the sequence length (i.e., the number of feature vectors). The dimension of each feature vector can be represented; therefore, the feature map output by the CNN needs to be... Transform into In other words, each spatial location (total) The feature vectors are used as elements of the sequence to obtain the feature sequence. Since the Transformer encoder itself does not have positional information, positional encoding information needs to be added to the feature sequence. The positional encoding information can be a matrix with the same shape as the feature sequence. By adding it to the feature sequence, a feature sequence with positional information can be obtained. For example, the positional encoding information can be a fixed code generated by sine and cosine functions, or it can be learnable parameters.

[0051] In some examples, step 1) above, which involves extracting features from video frames to obtain a feature sequence, can specifically include the following: First, a deep residual network is used to extract features from the video frames to obtain feature maps. Then, a feature sequence is generated based on the feature maps. In this example, the deep residual network can include ResNet34, ResNet50, etc.

[0052] Step 2) Detect human body key points in the video frame to obtain human body key point information.

[0053] In this example, human keypoint detection can be performed on human bodies detected in video frames. For example, various human keypoint detection networks, such as CNN and HRnet (High-Resolution Network), can be used to perform human keypoint detection on human bodies detected in video frames and obtain human keypoint information.

[0054] Step 3) Input the feature sequence with location information and human key point information into the encoder, and the encoder outputs feature information.

[0055] In this example, feature sequences with location information and human keypoint information can be fused (e.g., concatenated) and input into the encoder, which then outputs feature information. This yields feature information containing human keypoint information. By adding human keypoint information, the model's ability to detect people is improved, and its ability to understand human behavior in occluded, blurred, and easily confused retail scenarios is significantly enhanced. This is beneficial for classifying relationships between people and objects, and between people themselves.

[0056] Step 202: Input the learnable query vector and feature information into the subject-object pair matching module, and the subject-object pair matching module outputs multiple subject-object pairs.

[0057] In this embodiment, the interaction relationship recognition model may include a set of learnable query vectors, which can be learned during the training process of the interaction relationship recognition model. That is, the learnable query vectors are model parameters, learned through training, and used in the subject-object pair matching module to query subject-object pair information. For example, the subject-object pair matching module may include a Transformer decoder. The learnable query vectors and feature information can interact in the Transformer decoder through a cross-attention mechanism, ultimately predicting multiple subject-object pairs. Here, each subject-object pair may include a subject detection box, an object detection box, the corresponding category of the object detection box, and the relationship confidence. In this example, the subject may include a human body, and the object may include both human bodies and objects. As an example, the subject detection box and the object detection box may include center point coordinates (x, y), width w, and height h, and the category corresponding to the object detection box may include human body, mobile phone, POS machine, cash, shopping bag, etc. Relationship confidence can be used to indicate the probability that there is a relationship between the subject corresponding to the subject detection box and the object corresponding to the object detection box. For example, the relationship confidence can be a value between 0 and 1, and the larger the value, the higher the probability that there is a relationship between the subject and the object.

[0058] Step 203: Input the subject-object pair information and feature information into the relationship recognition module, and the relationship recognition module generates the action relationship category of the subject and object in each subject-object pair information.

[0059] In this embodiment, the subject-object pair information output by the subject-object pair matching module can be used as input to the relationship recognition module. By utilizing the prior knowledge learned by the subject-object pair matching module, the relationship recognition module can obtain the action relationship category. For example, the subject-object pair matching module may include a Transformer decoder. The subject-object pair information output by the subject-object pair matching module can be used as an initialized action query vector. The action query vector and feature information can interact in the Transformer decoder through a cross-attention mechanism, ultimately outputting the action relationship category between the subject and object in the subject-object pair information. For example, the action relationship category between human bodies may include transactions, conversations, hugs, etc. The action relationship category between a human body and an object may include taking, operating, scanning, displaying, passing, receiving, etc. As an example, the relationship recognition module can output a data tuple containing <subject detection box, object detection box, object category, action relationship>.

[0060] In some examples, the relationship recognition module may include an action relationship decoder, which may include a self-attention mechanism and a cross-attention mechanism. For example, the action relationship decoder may be a Transformer decoder. Based on this, step 203 above may specifically include the following: inputting the subject-object pair information and feature information into the action relationship decoder, and having the action relationship decoder generate the action relationship categories of the subject and object in each subject-object pair information. The self-attention mechanism of the action relationship decoder may apply a preset attention mask to the attention score to suppress unreasonable action relationship categories between the subject and object.

[0061] In this example, an attention mask, such as a mask matrix, can be pre-set (e.g., manually set). This attention mask can suppress unreasonable action relationships between the subject and object. Here, unreasonable action relationships can refer to actions that would not occur in the real world; for example, actions like "take" or "lift" would not occur between people. For instance, in a self-attention mechanism, the mask matrix can be used to combine the mask matrix with the original attention score matrix through mathematical operations before calculating the attention weights, thereby controlling the distribution of attention weights and masking certain positions, causing the weights of these positions to change after softmax. In this way, the action relationship decoder will not focus on those masked positions. Thus, by applying an attention mask, unreasonable action relationships between the subject and object can be prevented from being learned, and reasonable action relationships between the subject and object can be focused on.

[0062] For example, the formula for calculating the attention outcome in the self-attention mechanism can be as follows:

[0063] ,

[0064] in, It can represent subject-object pair information of the action relationship category to be matched; It can carry the semantic features of subject-object interactions (such as the semantics of "holding" in the context of a human body and a mobile phone), and be used for computation. The degree of matching with the interaction semantics; V can represent the core feature carrier of the subject-object interaction semantics, which is finally weighted by attention weights to output the refined interaction features for action relationship classification. M can represent a mask matrix, which can be used to constrain... Search At that time, we only focus on the reasonable subject-object interaction semantics. It can represent element-wise multiplication, that is, multiplying corresponding elements of two matrices. This can represent the dimension of the feature vector, used to normalize the attention score and prevent numerical instability caused by excessively large or small values. In this example, the mask matrix M can be a validity mask for the semantic dimension of subject-object interaction. Elements in the mask matrix can include 0 and 1; an element of 1 indicates a reasonable action relationship, and an element of 0 indicates an unreasonable action relationship. In this example, the mask matrix can be manually set. By adding a mask matrix, incorrect action relationship categories between the subject and object can be effectively avoided, improving the accuracy of interaction relationship recognition.

[0065] In practice, identifying the action relationship between the subject and object can provide a basis for transaction behavior. Complex and diverse transaction behaviors are accompanied by a variety of changing action relationships between people and between people and objects. For example, there are action relationships between customers and items such as mobile phones, cash, and shopping bags; and action relationships between cashiers and items such as POS machines, mobile phones, cash, and shopping bags. Action relationships between people and objects can include taking, handing over, operating, scanning, etc. Furthermore, action relationships between people and objects can also be mapped to action relationships between people, where action relationships between people can include transactions.

[0066] In some examples, the above-described interaction relationship identification method may further include the following steps S1 to S3, specifically:

[0067] Step S1: Analyze the interaction relationship between the subject and object identified from multiple video frames.

[0068] In this example, various analyses can be performed on the interaction relationships between the subject and object identified from multiple video frames. Here, the multiple video frames can be video frames captured within a target time period, which can be any time period. For example, suppose we want to identify whether a transaction occurred within a certain time period, then that time period can be used as the target time period.

[0069] Step S2: Based on the analysis, determine whether any transaction occurred within the target time period.

[0070] In this example, the analysis of the action relationships between subjects and objects across multiple video frames can determine whether a transaction occurred within a target time period. For instance, firstly, multiple object tracking (MOT) technology can be used to correlate detected human bodies and objects across each video frame, forming trajectories and analyzing their temporal interactions. Then, the tracking results and pre-defined logical rules can be used to determine if a transaction occurred. For example, logical rules can be set such that if a series of actions is detected, such as human A transferring cash, human B receiving cash, human B transferring a shopping bag, and human A receiving a shopping bag, a transaction can be determined to have occurred. Alternatively, logical rules can be set such that if a series of actions is detected, such as human A displaying a mobile phone, human B operating a POS machine, human B transferring a shopping bag, and human A receiving a shopping bag, a transaction can be determined to have occurred.

[0071] Step S3: In response to determining that a transaction has occurred, generate transaction-related information based on the analysis.

[0072] In this example, if a transaction is determined to have occurred, transaction-related information can be generated based on the analysis of the action relationships between the subject and object identified in multiple video frames. This transaction-related information can include at least one of the following: transaction method, transaction tool, and whether goods have been picked up. The transaction method can include cash payment, POS machine payment, mobile payment, etc. The transaction tool can include cash, a POS machine, a mobile phone, etc. In this example, information such as the transaction method, transaction tool, and whether goods have been picked up can be generated based on the analysis of the action relationships between the subject and object identified in multiple video frames. For example, when determining the transaction method and transaction tool, if multiple video frames identify a series of actions such as a person passing cash and a person receiving cash, the transaction method can be determined to be cash payment and the transaction tool to be cash; if multiple video frames identify a series of actions such as a person displaying a mobile phone and a person operating a mobile phone, the transaction method can be determined to be mobile payment and the transaction tool to be a mobile phone; if multiple video frames identify a person operating a POS machine, the transaction method can be determined to be POS machine payment and the transaction tool to be a POS machine. When determining whether goods have been picked up, if multiple video frames identify the action of a customer picking up a shopping bag, then it is determined that goods have been picked up. This implementation method can determine whether a transaction has occurred based on the action relationships between the subject and object identified in multiple video frames, and generate transaction-related information when a transaction occurs. Based on the interaction relationship recognition model, it identifies transaction behavior in sales scenarios, going beyond the surface-level action states between customers and cashiers to delve into the detailed relationships between people, cash registers, and goods within the transaction scene. This makes the identification of various behaviors, such as normal transactions, abnormal transactions, and potential transactions, more accurate. Therefore, it can handle complex and diverse transaction behavior patterns in real life. By acquiring and analyzing video frames in real time through the model, it can provide timely technical support to merchants. When abnormal transactions occur (e.g., picking up goods without payment), it can provide timely warnings and allow merchants to take measures to reduce losses, effectively ensuring their safety. Furthermore, it detects multi-frame human-object relationships and human-human relationships and progressively identifies transaction behavior based on the results of multiple frames. This method can obtain more accurate information about the transaction process at a deeper level, including whether a transaction occurred between people, the transaction tools used, and whether goods were picked up. It avoids the misjudgment of behavior caused by traditional behavior recognition that only focuses on the recognition of actions, and provides a lot of information about the transaction process.

[0073] Please continue reading Figure 3 , Figure 3 This diagram illustrates one application scenario to which the embodiments of this specification can be applied. Figure 3In the application scenario shown, video of the location of the cashier in a store can be acquired. The video frames can be input into an interaction relationship recognition model, which may include a feature extraction module, a subject-object pair matching module, and a relationship recognition module. The feature extraction module may include a feature extraction network (CNN) and an encoder; the subject-object pair matching module may include a subject-object pair matching decoder; and the relationship recognition module may include an action relationship decoder. The video frames are processed by the CNN to extract features, resulting in feature maps. Then, a feature sequence is generated based on the feature maps. Location encoding information is added to the feature sequence. The process begins by obtaining a feature sequence with location information. This feature sequence is then fused with human keypoint information and input into an encoder to obtain feature information. Next, the learnable query vector and feature information are input into a subject-object pair matching module, which outputs multiple subject-object pairs. Each subject-object pair includes a subject detection box, an object detection box, the corresponding category of the object detection box, and a relationship confidence score. Finally, the subject-object pair information and feature information output by the subject-object pair matching module are input into a relationship recognition module, which generates the action relationship category between the subject and object in each subject-object pair.

[0074] According to another embodiment, an interaction relationship identification device is provided. This interaction relationship identification device can be deployed in any device, platform, or device cluster with computing and processing capabilities.

[0075] Figure 4 A schematic block diagram of an interaction relationship recognition device according to one embodiment is shown. Figure 4 As shown, the interaction relationship recognition device 400 includes:

[0076] Acquisition unit 401 is used to acquire video of a target area within the sales venue, wherein the target area includes the area where the cashier is located;

[0077] The recognition unit 402 is used to input video frames from the aforementioned video into a pre-trained interaction relationship recognition model to obtain the interaction relationship between the subject and object in the video frame, wherein the subject includes the human body, and the object includes the human body and objects; the aforementioned interaction relationship recognition model includes a feature extraction module, a subject-object pair matching module, and a relationship recognition module; the aforementioned interaction relationship recognition model performs the following recognition operations on the video frame:

[0078] The video frame is input into the feature extraction module, which then extracts feature information from the video frame.

[0079] The learnable query vector and the aforementioned feature information are input into the subject-object pair matching module. The subject-object pair matching module outputs multiple subject-object pair information, wherein each subject-object pair information includes a subject detection box, an object detection box, the category corresponding to the object detection box, and the relationship confidence.

[0080] The subject-object pair information and the aforementioned feature information are input into the relationship recognition module, which then generates the action relationship category between the subject and object in each subject-object pair information.

[0081] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program / instructions are stored, which, when executed by a processor, implement... Figure 1 The method described.

[0082] According to another embodiment, a computing device is also provided, including a processor, a memory, and a computer program / instructions stored in the memory, wherein the processor is configured to execute the computer program / instructions, and when the computer program / instructions are executed, implement... Figure 1 The method described.

[0083] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.

[0084] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.

[0085] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0086] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying interaction relationships, characterized in that, The method includes the following steps: Acquire video of a target area within the sales venue, wherein the target area includes the area where the cashier is located; The video frames from the video are input into a pre-trained interaction relationship recognition model to obtain the interaction relationships between the subject and object in the video frames. The subject includes the human body, and the object includes both the human body and objects. The interaction relationship recognition model includes a feature extraction module, a subject-object pair matching module, and a relationship recognition module. The interaction relationship recognition model performs the following recognition operations on the video frames: The video frame is input into the feature extraction module, which extracts feature information from the video frame. The learnable query vector and the feature information are input into the subject-object pair matching module, and the subject-object pair matching module outputs multiple subject-object pair information, wherein each subject-object pair information includes a subject detection box, an object detection box, the category corresponding to the object detection box, and the relationship confidence. The subject-object pair information and the feature information are input into the relationship recognition module, and the relationship recognition module generates the action relationship category of the subject and object in each subject-object pair information. The action relationship category of the subject and object includes the action relationship category between human bodies and the action relationship category between human and object. The action relationship between the subject and the object identified from multiple video frames is analyzed, wherein the multiple video frames are video frames collected during the target time period; Based on the analysis, it is determined whether any transactions occurred within the target time period; In response to determining that a transaction has occurred, transaction-related information is generated based on analysis, wherein the transaction-related information includes at least one of the following: transaction method, transaction instrument, and whether goods have been picked up.

2. The method according to claim 1, characterized in that, The relationship recognition module includes an action relationship decoder, which includes a self-attention mechanism and a cross-attention mechanism; and the step of inputting the subject-object pair information and the feature information into the relationship recognition module, and having the relationship recognition module generate the action relationship categories of the subject and object in each subject-object pair information, includes: The subject-object pair information and the feature information are input into the action relationship decoder, which generates the action relationship category of the subject and object in each subject-object pair information. The action relationship decoder applies a preset attention mask to the attention score in its self-attention mechanism.

3. The method according to claim 1, characterized in that, The feature extraction module includes an encoder, and the extraction of feature information from video frames by the feature extraction module includes: Feature extraction is performed on video frames to obtain feature sequences, and position encoding information is added to the feature sequences to obtain feature sequences with position information; Human key point detection is performed on the human body in the video frame to obtain human key point information; The feature sequence with location information and the human body key point information are input into the encoder, and the encoder outputs feature information.

4. The method according to claim 3, characterized in that, The step of extracting features from video frames to obtain a feature sequence includes: A deep residual network is used to extract features from video frames to obtain feature maps. A feature sequence is generated based on the feature map.

5. The method according to claim 1, characterized in that, The interaction relationship recognition model was trained in the following way: Obtain a training sample set, wherein the training samples include sample images and their corresponding annotation data. The annotation data includes sample subject detection boxes, sample object detection boxes, sample category corresponding to the sample object detection boxes, sample relationship confidence, and sample action relationship category. Input sample images into the interaction relationship recognition model, which then outputs prediction information. Based on the difference between predicted and labeled information, the parameters of the interaction relationship recognition model are adjusted.

6. The method according to claim 5, characterized in that, Based on the difference between predicted and labeled information, the parameters of the interaction relationship recognition model are adjusted, including: Based on the difference between the first prediction information and the annotation information output by the subject-object pair matching module, the parameters of the feature extraction module and the subject-object pair matching module are adjusted. The first prediction data includes the predicted subject detection box, the predicted object detection box, the prediction category corresponding to the predicted object detection box, and the prediction relationship confidence. Based on the difference between the predicted action relationship category output by the relationship recognition module and the sample action relationship category, the parameters of the feature extraction module, the subject-object pair matching module, and the relationship recognition module are adjusted. The learning rate of the corresponding parameters of the feature extraction module and the subject-object pair matching module is different from the learning rate of the corresponding parameters of the relationship recognition module.

7. The method according to claim 1, characterized in that, The feature extraction module includes a feature extraction network and an encoder, the subject-object pair matching module includes a subject-object pair matching decoder, and the relationship recognition module includes an action relationship decoder.

8. A computing device, comprising a processor, a memory, and computer programs / instructions stored in the memory, characterized in that, The processor is configured to execute the computer program / instructions, and when the computer program / instructions are executed, the computing device implements the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Human-object interaction relationship identification method, model training method and corresponding device

    CN112633159A