Memory network enhancement method for experimental interaction behavior recognition and application

By fusing the global and local features of experimental videos in memory networks and introducing memory embedding losses, the problem of unstable interactive behavior recognition performance in chemical experiments is solved, especially in the case of uneven data distribution, and higher recognition accuracy and robustness are achieved.

CN120220033AActive Publication Date: 2025-06-27HUNAN NORMAL UNIVERSITY

Patent Information

Application Number
CN202510521621.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Existing video-based interactive behavior recognition methods are not stable enough during chemical experiments, especially when data distribution is uneven, it is difficult to accurately identify interactive actions with scarce data.

Method used

The memory network enhancement method is adopted to extract the global features of experimental videos and the regional graph network through the video-level graph network, and the memory terms in the memory network are integrated and enhanced, and the inter-class distance between similar samples is expanded by using memory embedding loss to improve the ability to identify scarce data.

Benefits of technology

In the case of uneven data distribution, it is realized that the ability to identify scarce data is improved, the intra-class distance is narrowed and the inter-class distance is expanded, and the robustness and accuracy of the experimental interactive behavior recognition model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220033A_ABST
    Figure CN120220033A_ABST
Patent Text Reader

Abstract

The invention discloses a memory network enhancement method for experimental interaction behavior recognition and application, and belongs to the character interaction behavior recognition technology in video image recognition, a video level graph network is used for extracting global features of each category in a data set, and memory items of the memory network are initialized to serve as supplementary information to reduce the intra-class distance; extracting local features of a single video by using a region-level graph network, capturing detail information of an interaction behavior, and fusing the detail information with global features in a memory network to enrich sample features; and designing memory embedding loss to enlarge the inter-class distance between similar samples, thereby realizing accurate identification of the interaction behavior, and further improving the model robustness. According to the invention, accurate identification of experimental interaction behavior categories in the experimental video is realized, and the overall performance of the model is excellent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a memory network enhancement method and application for experimental interaction behavior recognition, belonging to the application of human interaction behavior recognition technology in teaching practice. Background Art

[0002] Existing human interaction behavior recognition methods can be classified into image-based human interaction behavior recognition methods and video-based human interaction behavior recognition methods according to data processing types. Among them, the image-based human interaction behavior recognition method refers to recognizing the interaction relationship between each pair of people and objects in an image. Currently, this type of method mainly uses multi-stream networks or graph networks, etc., to understand the context information in the scene. However, this method cannot capture the dynamic changes and time series information of actions, and it is difficult to be applied to real scenarios that require detecting continuous actions. The video-based human interaction behavior recognition method can identify and locate the time periods when all human interaction behaviors occur in a video and their specific interaction behavior categories. This type of method constructs and analyzes the spatio-temporal interaction graph structure between people in the video, and uses technologies such as graph networks and recurrent neural networks to capture the spatio-temporal semantic information in video frames, so as to identify and locate the time periods and interaction behavior categories of human interaction behaviors in the video.

[0003] Currently, the human interaction behavior recognition method is mainly based on graph networks. For example, in order to simultaneously model spatio-temporal correlation and the long-term time dynamics of objects, Ning W, Guangming Z, Hongsheng L, etc. proposed the STGC network, which captures intra-frame and inter-frame dependencies by simultaneously capturing spatial and temporal correlations. In order to further improve the model's ability to capture context information and complex relationships in the scene, Hong H S, Lee J C, Kumar A, etc. proposed DT-HOI, which enhances the model's recognition ability for ambiguous actions by introducing extended text descriptions as supplementary inputs to visual features. However, existing video-based interaction behavior recognition methods mainly start from deeply mining the spatio-temporal relationships in videos, and improve the model's understanding ability of interaction videos by constructing spatio-temporal network feature maps of individual videos, without considering the impact of data distribution on the model.

[0004] In addition, current research on the problem of uneven data distribution mainly focuses on image-based interactive behavior recognition, including methods based on training strategies and combination learning. Among them, the method based on training strategies refers to improving the performance of interactive behavior recognition through strategies such as data augmentation. Fang S, Liu S, Li J et al. proposed using the label-to-image generation method to generate virtual images to solve the problem of uneven data distribution. The method based on combination learning alleviates the data distribution problem by recombining different representations of human-object pairs and interactions to form new triples. For example, Hou Z, Peng X, Qiao Y et al. decomposed the HOI representation into object and verb-specific features and composed new interactive samples in the visual feature space. Moreover, in a virtual-real fusion chemistry experiment platform for real-time feedback based on multimodal perception disclosed in a Chinese patent application with the application number CN202411636115.X, during the student experiment process, the interactive behavior between the student and the experimental equipment is recognized through HOI, and it is judged whether it aligns with the key points in the preset experimental steps. Although the above methods have achieved good results in some application scenarios, they mainly focus on solving the problem of uneven distribution of interactive behavior data in the image field, and there is little research on the recognition method for judging interactive actions during the chemistry experiment process based on video.

[0005] In the real interactive scenario of a chemistry experiment, there is often a situation of uneven distribution of category sample data, that is, the number of samples in some categories is much larger than that in other categories, resulting in the model tending to learn samples with rich data, having a high recognition accuracy for common behaviors, but having limited learning ability for samples with scarce data and insufficient recognition ability for few-sample behaviors, making the model deviate in the recognition of intra-class homogeneity and inter-class heterogeneity during the optimization process. For example, in the eight major chemistry experiment scenarios that students must do as stipulated in the junior high school compulsory education chemistry curriculum standard, some key actions (such as 'oscillation') only appear in a few experiments, while other common actions (such as 'pouring') appear frequently. And the accurate recognition of key actions such as 'oscillation' with scarce data is very important for the subsequent evaluation of chemistry experiment operation ability. Although the existing video-based interactive behavior recognition methods have achieved good results, there is little research on the problem of uneven data distribution. Therefore, how to improve the recognition ability for scarce data, narrow the intra-class distance and widen the inter-class distance under the condition of uneven data distribution is an urgent problem to be solved in the video-based interactive behavior recognition task. Summary of the Invention

[0006] The technical problem solved by the present invention is: aiming at the problem that the recognition performance of the existing human interactive behavior recognition method is not stable enough during the chemistry experiment process, a memory network enhancement method and application for experimental interactive behavior recognition are provided.

[0007] The present invention is implemented by adopting the following technical solutions:

[0008] The present invention first discloses a memory network enhancement method for experimental interaction behavior recognition, including the following steps:

[0009] S1. Obtain a video dataset containing multiple experimental videos, construct a global video graph of the experimental videos through a video-level graph network, and extract the global features of each experimental interaction behavior category in all the experimental videos of the video dataset;

[0010] S2. Construct a frame-level local interaction graph of the experimental videos through a region-level graph network to obtain the local features of a single experimental video,

[0011] S3. Initialize the memory items of the memory network with the global features, project the local features into the feature space of the memory items through an attention mechanism for fusion enhancement with the global features, and expand the inter-class distance of similar experimental interaction behaviors between the memory items through a memory embedding loss.

[0012] In the memory network enhancement method for experimental interaction behavior recognition provided by the present invention, further, in the step S1, the video-level graph network processes the time series of the frame-level features of all the experimental videos using a BiRNN to obtain the visual features of each experimental video, instantiates the video nodes of the experimental videos with the visual features, and connects every two video nodes pairwise to obtain an initialized global video graph.

[0013] In the memory network enhancement method for experimental interaction behavior recognition provided by the present invention, further, in the step S1, the extraction steps of the global features of each experimental interaction behavior category are as follows:

[0014] First, update the video nodes of the global video graph through a graph attention network, and obtain the attention coefficient a between the video nodes through an attention mechanism ij ;

[0015] Secondly, update to obtain the final feature x i ’

[0016] ,

[0017] of a single experimental video through the attention coefficient, where x i is the video node feature, W video is the attention mechanism weight matrix, and σ is the activation function;

[0018] Finally, after updating the visual features of all the experimental videos using the graph attention network, group the visual features of the experimental videos according to the label information of the experimental interaction behaviors, and calculate the global features of each experimental interaction behavior category through the following formula

[0019] ,

[0020] where c k is the global feature of the k-th experimental interaction behavior category, N video is the number of all experimental videos, y i is the visual feature of the i-th experimental video, and the global feature dataset C of all experimental interaction behavior categories is obtained, C = {c k} k=1,...,K , where K represents the number of experimental interaction behavior categories.

[0021] In the memory network enhancement method for experimental interaction behavior recognition provided by the present invention, further, in the step S2, the region-level graph network uses Faster R-CNN or YOLOv8 to locate the instances of the operator or experimental equipment on each frame picture of the experimental video, extracts the visual features, spatial features, and semantic features of all instances for splicing, constructs the frame-level local interaction graph of the experimental video, and connects the features spliced within and between frames to obtain the local features of a single experimental video.

[0022] In the memory network enhancement method for experimental interaction behavior recognition provided by the present invention, further, the instance located by Faster R-CNN has a four-dimensional coordinate box and a target category,

[0023] Each instance crops the region of interest from the original experimental video frame according to its four-dimensional coordinate box and inputs it into the pre-trained ResNet-50 model to extract the visual feature F v ,

[0024] The multi-layer perceptron MLP is used to convert the four-dimensional coordinates of each instance into the spatial feature F of the corresponding dimension s ,

[0025] The learnable word embedding model is used to generate the semantic feature F of the instance w .

[0026] In the memory network enhancement method for experimental interaction behavior recognition provided by the present invention, further, the frame-level local interaction graph includes a set of nodes of the operator or experimental equipment in the frame picture, and a set of edges of the operator or experimental equipment in the frame picture. For the edges connecting the operator and the experimental equipment, the edge weights are initialized to 1, and the rest are 0. The graph attention network is used to analyze the intra-frame edge connection graph and the inter-frame edge connection graph to obtain the intra-frame edge connection graph attention coefficient and the inter-frame connection graph attention coefficient , and the features spliced within and between frames are connected through the following formula to obtain the final local features of a single experimental video ,

[0027] ,

[0028] Among them, W inra and W inter are the intra-frame edge weight matrix and the inter-frame edge weight matrix, H represents the concatenated features of the visual, spatial, and semantic features of the instance, and [] represents the concatenation operation.

[0029] In the memory network enhancement method for experimental interaction behavior recognition provided by the present invention, further, in the step S3, the process of fusing and enhancing the local features and the global features is as follows:

[0030] S31. Construct a memory network M using the global features of all experimental videos. This memory network has K memory items, corresponding to the number of experimental interaction behavior categories, and each memory item is initialized with the global features of the corresponding experimental interaction behavior category;

[0031] S32. Input the local features of a single experimental video into the memory network, and project the local features into the feature space of the memory item through the attention mechanism method to obtain a new feature representation ;

[0032] S33. The memory item enhances the projected local features through the following formula

[0033] ,

[0034] where is the local feature projected into the feature space of the memory item, is the feature after the memory item is enhanced, c k is the global feature of the k-th experimental interaction behavior category, c j represents the global feature corresponding to the j-th memory item of the memory network;

[0035] S34. Obtain the label of the predicted experimental interaction behavior category through the softmax function.

[0036] In the memory network enhancement method for experimental interaction behavior recognition provided by the present invention, further, in the step S3, the cosine similarity loss L sim and the Euclidean distance loss L euc are jointly used as the memory embedding loss L mem of the memory network;

[0037] ,

[0038] where λ is a hyperparameter;

[0039] The cosine similarity loss L sim is obtained by minimizing the cosine similarity between the experimental interaction behavior features of different memory items through the L2 norm,

[0040] ,

[0041] Among them, s ij represents the cosine similarity between memory item c i and memory item c j . s is the similarity matrix of all cosine similarities between memory items of the memory network, and K is the number of memory items of the memory network;

[0042] Calculate the Euclidean distance between memory items to obtain the distance matrix of all Euclidean distances between memory items of the memory network. Sort the elements of each row of the distance matrix in ascending order to get D’=(d ij ’)∈R K×K , d ij ’ represents the Euclidean distance between the sorted memory item c i and memory item c j . Select the k2 minimum values in each row of the distance matrix through the following formula to increase the Euclidean distance between memory items.

[0043] ,

[0044] Among them, d — is the increased Euclidean distance, and the Euclidean distance loss L euc =max(0, -d — +γ), where γ is a hyperparameter for adjusting the distance between memory items.

[0045] The present invention also discloses an experimental interaction behavior recognition method. Input the experimental video to be recognized into the memory network enhanced by the above memory network enhancement method of the present invention, retrieve the memory items corresponding to the experimental interaction behavior from the memory network in combination with the currently input video data, and output the experimental interaction behavior category in the experimental video to be recognized.

[0046] Furthermore, the present invention also discloses an experimental interaction behavior recognition system, including:

[0047] An input module for inputting the experimental video to be recognized;

[0048] A memory module for storing the memory network enhanced by the above memory network enhancement method of the present invention;

[0049] A query module for retrieving the memory items corresponding to the experimental interaction behavior from the memory network in combination with the currently input video data;

[0050] An output module for outputting the experimental interaction behavior category in the experimental video to be recognized.

[0051] The present invention has the following beneficial effects:

[0052] (1) First, the present invention uses a video-level graph network to extract the global features of each experimental interaction category in the experimental video dataset, and initializes the memory items of the memory network as supplementary information to narrow the intra-class distance. Secondly, it uses a region-level graph network to extract the local features of a single experimental video, captures the detailed information of the experimental interaction behavior, and fuses it with the global features in the memory network to enrich the sample features. Finally, a memory embedding loss is designed to expand the inter-class distance between similar samples. By constructing a global video graph and a frame-level local interaction graph, global features and local features are obtained. A memory network is designed to fuse the global features and local features, supplement the context information of the experimental interaction behavior of scarce data to narrow the intra-class distance, and introduce a memory embedding loss to address the uneven distribution of human interaction actions in the experimental video, so as to achieve accurate recognition of experimental interaction behaviors, and further improve the robustness of the experimental interaction behavior recognition model.

[0053] (2) The present invention adopts a global feature and local feature fusion framework. In view of the fact that existing methods mostly focus on the construction of spatio-temporal relationships within a single experimental video and do not supplement the context information of scarce category interaction behaviors from a global perspective, the present invention extracts the global features and local features of human interaction behaviors in the experimental video through independent global feature branches (video-level graph networks) and local feature branches (region-level graph networks). The global feature branch uses a video-level graph network to model all experimental videos, mines the relationships between different video nodes through a graph attention network, updates the node features, and groups the visual features according to the label information of the experimental interaction behaviors, calculates the global features of each interaction behavior category, aggregates the sample features of the entire video dataset, mines the internal connections and rules of the video dataset, and enriches the experimental interaction behavior feature representation of the model. The local feature branch constructs a frame-level local interaction graph through a region-level graph network, extracts the local features of a single experimental video, including the visual features, spatial features and semantic features of all instances, and uses a graph attention network to analyze the intra-frame edge connection graph and the inter-frame edge connection graph to capture the overall semantics of the frame-level interaction behavior. The memory items in the memory network are initialized as the global features of various experimental interaction behaviors. When the local features of a single experimental video are input, the local visual features are projected into the feature space of the memory items through a multi-layer perceptron, and the features are enhanced by the memory items, so that the final visual features fuse the global features and local features similar to it in the entire dataset, especially enriching the context information of scarce sample interaction behaviors. The synergy of the two makes up for the defect that local features are difficult to handle similar categories and effectively addresses the challenge of uneven distribution of experimental interaction behavior sample data in experimental videos.

[0054] (3) To solve the problem that in the recognition experiment, some interaction behaviors have a large semantic similarity and a small feature distance between different categories, the present invention introduces a memory network to store the prototype features of various experimental interaction behaviors in the dataset, and uses the combined cosine similarity loss and Euclidean distance loss as the memory embedding loss. The cosine similarity loss is used to minimize the cosine similarity between memory items, so that each pair of memory items is different in the semantic space; the Euclidean distance loss is used to further expand the distance between memory items in the Euclidean space to ensure that each memory item maintains a sufficient distance from other memory items, thereby expanding the inter-class distance of similar experimental interaction behaviors and improving the discrimination ability of the recognition model for similar experimental interaction behaviors. By using the memory embedding loss to expand the inter-class distance between similar experimental interaction behavior samples, the model's understanding ability for data-scarce experimental interaction behaviors is enhanced, and it can reallocate resources for data-scarce experimental interaction behavior samples and data-rich experimental interaction behavior samples during training to perform better feature discrimination, resulting in a significant improvement in the discrimination ability of the recognition model for similar experimental interaction behaviors and optimizing the overall performance of the recognition model.

[0055] In summary, the memory network enhancement method and application for experimental interaction behavior recognition provided by the present invention achieve accurate recognition of the categories of experimental interaction behaviors in experimental videos, and the overall performance of the model is excellent, which can be used for automatic recognition of videos in the teaching experiment assessment of various science and engineering disciplines.

[0056] The following further describes the present invention in conjunction with the drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic flowchart of the memory network enhancement method for experimental interaction behavior recognition in the embodiment.

[0058] Figure 2 It is a sample distribution table of the CE-HIRD dataset in the embodiment

[0059] Figure 3 It is a sample distribution table of the CAD120 dataset in the embodiment.

[0060] Figure 4 It is a sample distribution table of the Something-else dataset in the embodiment.

[0061] Figure 5 It is a confusion matrix diagram of the chemical experiment interaction behavior recognition of the CE-HIRD dataset for the baseline model (left) and the present embodiment (right).

[0062] Figure 6 It is a confusion matrix diagram of the object function usability of the CE-HIRD dataset for the baseline model (left) and the present embodiment (right).

[0063] Figure 7 Feature visualization diagram of the baseline model.

[0064] Figure 8 Feature visualization diagram of this embodiment. Detailed implementation manners

[0065] Embodiment

[0066] Refer to Figure 1 , taking the chemical experiment interaction behavior in the chemical experiment video as an example, in this embodiment of the present invention, a global feature branch and a local feature branch are adopted to process the chemical experiment interaction behavior features in the chemical experiment video, and data augmentation is performed on the memory network for chemical experiment interaction behavior recognition, which specifically includes the following steps:

[0067] S1. Obtain a video data set containing multiple chemical experiment videos. The global feature branch first constructs a global video graph of the chemical experiment videos through a video-level graph network, extracts the global features of each chemical experiment interaction behavior category in all chemical experiment videos of the video data set, and initializes the memory items of the memory network with the global features as supplementary information to reduce the intra-class distance between chemical experiment interaction behaviors.

[0068] The global features can provide more comprehensive chemical video data information and context knowledge. By introducing the global features, more comprehensive data information and context knowledge can be obtained, reducing the increase in the intra-class distance between chemical experiment interaction behavior features caused by local feature differences, and improving the understanding ability of data-scarce interaction behavior categories. Therefore, this method introduces a global feature branch, constructs a global video graph of chemical experiment videos, and extracts the global features of each category in the data set. Specifically, in the global feature branch, the video-level graph network of this embodiment uses BiRNN to process the time series of all chemical experiment video frame-level features to obtain the visual features Y = {y i}i=1,...,N video , where N video is the number of all chemical experiment videos in the data set, and constructs a global video graph G video =(V video ,E video ), where V video represents the node set of the chemical experiment videos, V video ={x i |i=1,2,...,N video}, and instantiates each video node of the chemical experiment video with a single visual feature y i , E video represents the edges connecting different video nodes, and every two video nodes are connected pairwise to obtain an initialized fully connected global video graph.

[0069] The extraction steps of the global features of each chemical experiment interaction behavior category in this step are as follows:

[0070] First, the video nodes of the global video graph are updated through a graph attention network. The graph attention network (GAT) refers to using an attention mechanism to weight and sum the features of neighboring video nodes. The weights of the features of neighboring video nodes depend entirely on the node features and are independent of the graph structure. The attention coefficient a between video nodes is obtained through the attention mechanism. ij The specific calculation formula is as follows:

[0071] .

[0072] Where W video is the attention mechanism weight matrix, [x i , x j means concatenating the feature vectors of video nodes x i and x j , then multiplying by the attention mechanism weight matrix W. LeakyReLU is a non-linear activation function, and softmax is a normalized activation function used to map the attention coefficient to the range of [0, 1].

[0073] Secondly, the graph attention network updates to obtain the final feature x i ' of a single chemical experiment video.

[0074] .

[0075] Where x i is the video node feature, W video is the attention mechanism weight matrix, and σ is the activation function.

[0076] Finally, after using the graph attention network to update the visual features of all chemical experiment videos, the visual features of chemical experiment videos are grouped according to the label information of chemical experiment interaction behaviors. The global features of each chemical experiment interaction behavior category are calculated through the following formula.

[0077] .

[0078] Where c k is the global feature of the k-th chemical experiment interaction behavior category, N video is the number of all chemical experiment videos, y i is the visual feature of the i-th chemical experiment video, and the global feature dataset C of all chemical experiment interaction behavior categories is obtained, C = {c k} k=1,...,K, where \(K\) represents the number of categories of chemical experiment interaction behaviors. At this time, the global feature dataset \(C\) aggregates the sample features of the entire chemical experiment dataset to explore the internal connections and laws of the chemical experiment dataset and enrich the feature representation of the model.

[0079] Global feature extraction uses a video-level graph network to model all chemical experiment videos, explores the relationships between different video nodes through a graph attention network, updates the node features, groups the visual features according to the label information, calculates the global features of each interaction behavior category, aggregates the sample features of the entire dataset, explores the internal connections and laws of the dataset, and enriches the feature representation of the model.

[0080] S2. The local feature branch constructs a frame-level local interaction graph of the chemical experiment video through a region-level graph network, obtains the local features of a single chemical experiment video, captures the detailed information of the interaction behavior, and inputs the local features into a memory network to fuse and enhance with the global features to enrich the sample features.

[0081] The global feature branch mainly focuses on the features of the entire dataset, while the local feature branch captures the detailed information of the interaction behavior by extracting the frame-level features of a single chemical experiment video, making up for the deficiency of the global features in describing the details of specific interaction behaviors in a single chemical experiment video. In this step, the region-level graph network uses Faster R-CNN or YOLOv8 to locate the instances of operators or chemical experiment equipment in each frame picture of the chemical experiment video, extracts the visual features, spatial features, and semantic features of all instances for splicing, constructs a frame-level local interaction graph of the chemical experiment video, and connects the features spliced within and between frames to obtain the local features of a single chemical experiment video.

[0082] In this embodiment, taking Faster R-CNN as an example, the instances of operators or chemical experiment equipment located by Faster R-CNN carry a four-dimensional coordinate box and a chemical target category. Each instance crops the region of interest from the original chemical experiment video frame according to its four-dimensional coordinate box and inputs it into a pre-trained ResNet-50 model to extract the visual feature \(F\) v , uses a multi-layer perceptron MLP to convert the four-dimensional coordinates of each instance into spatial features \(F\) of the corresponding dimension s , and uses a learnable word embedding model to generate the semantic feature \(F\) of the instance w .

[0083] Specifically, given the chemical experiment video \(\{X\) m , c_m\}\) m=1,...,M , where \(X\) m = \(\{\) }\) t=1,...,TAll frames of the video are represented, T represents the total number of frames of the video, M represents the number of video nodes of the operator or chemical experimental equipment in the experimental video, and c m represents the interaction behavior or object function participation label corresponding to the video node, that is, the label information of the chemical experiment interaction behavior within the corresponding video node. Through the Faster R-CNN object detection network in each video frame image locate the instances of the operator or chemical experimental equipment, as follows:

[0084] .

[0085] where each instance has a four-dimensional coordinate box and a chemical target category. In the formula, T is the chemical target category of the instance, B is the four-dimensional coordinate box (x, y, W, H) of the instance, x and y are the abscissa and ordinate of the center point of the four-dimensional coordinate box, W and H are the width and height of the four-dimensional coordinate box, RPN refers to the Region Proposal Network, which aims to generate a set of candidate regions, ROIPooling refers to Region of Interest Pooling to map each candidate region to a fixed-size feature vector, FC refers to the fully connected layer, and CNN( ) refers to inputting the t-th frame image of the m-th node into the convolutional neural network CNN.

[0086] For extracting the chemical experiment visual features in the local features, in this embodiment, for each video frame image, the region of interest (ROI) of the operator or chemical experimental equipment is cropped from the original video frame according to its four-dimensional coordinate box, and is input into the pre-trained ResNet-50 model to extract the visual feature F v , and the calculation formula is as follows:

[0087] .

[0088] where, represents the t-th video frame image of the m-th chemical experiment video, and Crop refers to cropping the target region corresponding to the four-dimensional coordinates (x, y, W, H) of the frame image .

[0089] For extracting the spatial features in the local features, in this embodiment, a multi-layer perceptron (MLP) is used to convert the four-dimensional coordinates of the instance into spatial features F s to represent the position information between the operator and the chemical experimental equipment.

[0090] For extracting the semantic features in the local features, in this embodiment, a learnable word embedding model is used to generate the semantic features F of the operator and the chemical experimental equipment w, to represent the specific category information of students or chemical experiment equipment. The word embedding model can adopt the Linear128 - BatchNorm - ReLU - Linear256 - BatchNorm - ReLU model.

[0091] Next, the obtained visual feature F v , spatial feature F s and semantic feature F w are concatenated (Concat) by the following formula:

[0092] .

[0093] H is the concatenated feature of the visual, spatial, and semantic features of the instance.

[0094] Next, construct the frame - level local interaction graph G local of the chemical experiment video to explicitly model the human - object interaction relationship. The frame - level local interaction graph is defined as G local =(V ho , E ho ), where V ho represents the set of nodes of operators or chemical experiment equipment in the frame picture, and E ho represents the set of edges of operators or chemical experiment equipment in the frame picture. E ho is composed of intra - frame edges and inter - frame edges . The intra - frame edges are the edges composed of different operators or chemical experiment equipment within the same frame t, and the inter - frame edges represent the edges composed of the same operator or chemical experiment equipment in two adjacent frames t and t + 1. For the edges connecting operators and chemical experiment equipment, their edge weights are initialized to 1, and the rest are 0. Use the graph attention network to analyze the intra - frame edge connection graph and the inter - frame edge connection graph to obtain the intra - frame edge connection graph attention coefficient and the inter - frame connection graph attention coefficient . The final local feature of a single chemical experiment video is obtained by connecting the intra - frame and inter - frame concatenated features through the following formula , capturing the overall semantics of the frame - level interaction behavior.

[0095] .

[0096] Among them, W inra and W inter are the intra - frame edge weight matrix and the inter - frame edge weight matrix. H represents the concatenated feature of the visual, spatial, and semantic features of the instance, and [] represents the connection operation.

[0097] Local feature extraction constructs a frame-level local interaction graph of a chemical experiment video through a region-level graph network, extracts local features of a single video, including visual features, spatial features, and semantic features of all instances, and uses a graph attention network to analyze the intra-frame edge connection graph and the inter-frame edge connection graph to capture the overall semantics of frame-level interaction behaviors.

[0098] S3. Initialize the memory items of the memory network with global features, input the local features into the memory network, project the local features into the feature space of the memory items through the attention mechanism for fusion and enhancement with the global features, and expand the inter-class distance of similar chemical experiment interaction behaviors between the memory items through the memory embedding loss to achieve accurate recognition of interaction behaviors in chemical experiment videos, thereby improving the robustness of the model.

[0099] In the step S3, the process of fusing and enhancing the local features and the global features is as follows:

[0100] S31. To make up for the context information of data-scarce interaction behaviors and fully fuse the global features and local features of chemical experiment videos, use the global features of all chemical experiment videos to construct a memory network to enrich the features of data-scarce interaction behaviors. The memory network M = {c k} k=1,...,K constructed by the global features has K memory items, corresponding to the number of K chemical experiment interaction behavior categories. Each memory item is initialized with the global feature c k corresponding to the chemical experiment interaction behavior category.

[0101] S32. Input the local features of a single chemical experiment video into the memory network. When inputting the local features of a single chemical experiment video, the memory network projects the local features into the feature space of the memory items through the attention mechanism method to obtain a new feature representation . In this embodiment, the formula for projecting the local features by a multi-layer perceptron is as follows:

[0102] .

[0103] where MLP refers to a multi-layer perceptron.

[0104] S33. The memory item enhances the projected local features through the following formula:

[0105] .

[0106] where is the local feature projected into the feature space of the memory item, is the feature of the memory item after enhancement, and c kis the global feature of the k-th chemical experiment interaction behavior category, c j represents the global feature corresponding to the j-th memory item of the memory network, where j is the index of the memory item, and the value range is from 1 to K. Here, K is the total number of memory items in the memory network, that is, the number of interaction behavior categories.

[0107] S34. Obtain the label of the predicted chemical experiment interaction behavior category through the softmax function, that is, the result label of the chemical interaction action prediction, corresponding to the finally predicted chemical experiment interaction action category such as "oscillation".

[0108] Through the above method, the final visual feature fuses the global features and local features similar to it in the entire chemical experiment video dataset. Especially for interaction behaviors with scarce samples, the memory network enriches its context information and improves the model's understanding ability of scarce interaction behaviors.

[0109] The memory items in the memory network are initialized as the global features of various interaction behaviors in the chemical experiment video. When the local features of a single chemical experiment video are input, the local features of the video are projected into the feature space of the memory items through a multi-layer perceptron, and the memory items are used to enhance the local features, so that the final visual feature fuses the global features and local features similar to it in the entire video dataset, especially enriching the context information of scarce experimental interaction behaviors.

[0110] In order to further solve the problem that the semantic similarity between some interaction behaviors is large and the feature distance between different categories is small, this embodiment introduces a memory embedding loss to enlarge the inter-class distance of similar interaction behaviors in the chemical experiment video and improve the discriminative ability of the model for such similar chemical experiment interaction behaviors. Specifically, this embodiment combines the cosine similarity loss L sim and the Euclidean distance loss L euc as the memory embedding loss L mem .

[0111] For the cosine similarity loss L sim , first calculate the cosine similarity between the memory items of similar chemical experiment interaction behaviors to obtain the similarity matrix s, and the calculation formula is as follows:

[0112] .

[0113] Among them, is the normalized memory item matrix, where each row represents the feature vector of a memory item, represents transpose, s ij represents the memory item c i and the memory item c jThe cosine similarity between them, where K is the number of memory items in the memory network. The cosine similarity loss L is obtained by minimizing the cosine similarity between the chemical experiment interaction behavior features of different memory items through the L2 norm. sim ,

[0114] .

[0115] Among them, s is the similarity matrix of all cosine similarities between the memory items of the memory network, || ||2 is the L2 norm calculation, and s ij represents the cosine similarity between memory item c i and memory item c j , where K is the number of memory items in the memory network.

[0116] However, simply differentiating memory items by minimizing the cosine similarity is not enough. Some memory items still cannot be effectively distinguished from other features, resulting in similar chemical experiment interaction behaviors being difficult to effectively distinguish. In this embodiment, the distance between memory items in the Euclidean space is further expanded so that the model can further distinguish their features and expand the inter-class distance of chemical experiment interaction behaviors. The Euclidean distance between two memory items is calculated by the following formula:

[0117] .

[0118] Among them, d ij represents the Euclidean distance between memory item c i and memory item c j . By calculating the Euclidean distance between memory items, the distance matrix D of all Euclidean distances between the memory items of the memory network is obtained. The elements of each row of the distance matrix are sorted in ascending order to obtain D’=(d ij ’)∈R K×K , and d ij ’ represents the Euclidean distance between memory item c i and memory item c j after sorting, where K represents the total number of memory items in the memory network, that is, the number of interaction behavior categories. Each memory item corresponds to an interaction behavior category. Therefore, the rows and columns of matrix D represent the pairwise Euclidean distances between all memory items. The sorted matrix D’ is still K×K, and the elements of each row are arranged in ascending order. The k2 minimum values of each row of the distance matrix are selected by the following formula to increase the Euclidean distance between memory items:

[0119] .

[0120] Among them, d — is the increased Euclidean distance, and the Euclidean distance loss L euc =max(0, -d —+(γ), where γ is a hyperparameter for adjusting the distance between memory items, k2 is an integer less than K, and the value of k2 needs to be determined experimentally according to specific circumstances. If the semantic similarity in the dataset is relatively high, a larger value can be selected from the integers less than K. If the semantic similarity of the dataset is low, a smaller integer value is selected. In this embodiment, γ is taken as 0.8 and k2 is taken as 3.

[0121] The joint cosine similarity loss L sim and the Euclidean distance loss L euc are used to obtain the memory embedding loss L of the memory network mem,

[0122] .

[0123] Among them, λ is a hyperparameter. In this embodiment, when λ is taken as 0.6, the model has the best effect. Through the update and iteration of the memory network, the existing memory items represent the prototype features of each type of interaction behavior. By constraining the feature distance of the memory items, the distance between the features of different chemical experiment interaction behaviors can be effectively restricted.

[0124] The memory embedding loss is composed of the cosine similarity loss and the Euclidean distance loss. The cosine similarity loss is used to minimize the cosine similarity between memory items, so that each pair of memory items is different in the semantic space; the Euclidean distance loss is used to further expand the distance between memory items in the Euclidean space to ensure that each memory item maintains a sufficient distance from other memory items, thereby expanding the inter-class distance of similar experimental interaction behavior features and improving the discriminative ability of the model for similar experimental interaction behaviors.

[0125] This embodiment also discloses a method for identifying chemical experiment interaction behaviors by applying the above memory network enhancement method. The chemical experiment video to be identified is input into the memory network enhanced by the above memory network enhancement method, and the memory items corresponding to the chemical experiment interaction behaviors are retrieved from the memory network in combination with the currently input video data, and the category of the chemical experiment interaction behavior in the chemical experiment video to be identified is output.

[0126] The memory network constructed by the memory network enhancement method of this embodiment is applied to the chemical experiment interaction behavior recognition method. The global features of each category in the dataset are extracted using a video-level graph network, and the memory items of the memory network are initialized as supplementary information to reduce the intra-class distance. The local features of a single video are extracted using a region-level graph network to capture the detailed information of the interaction behavior and fuse them with the global features in the memory network to enrich the sample features. The independent video-level graph network and region-level graph network extract the global features and local features of the chemical experiment video respectively, and their synergistic effect makes up for the defect that local features are difficult to handle similar categories and effectively cope with the challenge of uneven data distribution. Finally, a memory embedding loss is designed to increase the inter-class distance between similar samples. By combining the cosine similarity loss and the Euclidean distance loss as the memory embedding loss, the feature distance of the memory items is constrained, which significantly improves the discriminative ability of the model for similar interaction behaviors, realizes the accurate recognition of interaction behaviors, and further improves the robustness of the model. It can accurately predict the interaction instances between the operator and chemical experiment equipment in the chemical experiment video, solve the problem that the model tends to learn samples with rich data due to unbalanced data distribution, enhance the learning ability for samples with scarce data, avoid the deviation in the recognition of intra-class homogeneity (large intra-class distance) and inter-class heterogeneity (small inter-class distance) during the recognition process, and improve the accuracy of chemical experiment interaction behavior recognition.

[0127] Correspondingly, this embodiment also provides a system for implementing the above chemical experiment interaction behavior recognition method, which specifically includes: an input module, a memory module, a query module, and an output module. The input module is used to input the chemical experiment video to be recognized; the memory module stores the memory network enhanced by the memory network enhancement method; the query module retrieves the memory items corresponding to the chemical experiment interaction behavior from the memory network in combination with the currently input video data; the output module outputs the chemical experiment interaction behavior category in the chemical experiment video to be recognized retrieved.

[0128] The recognition method of the chemical experiment interaction behavior in this embodiment can be used for video recognition assessment of students' chemical experiment interaction behaviors in chemical teaching experiments, and can also be used for video assessment of other science and engineering teaching experiments, which will not be elaborated here.

[0129] The performance of this embodiment is verified through specific datasets below.

[0130] In this embodiment, three datasets are used for experiments, namely the CE-HIRD dataset, the CAD120 dataset, and the Something-else dataset. Among them, the CE-HIRD dataset is a self-built dataset for recognizing chemical experiment interaction behaviors, and the CAD120 and Something-else datasets are public datasets. The datasets selected in this paper are all interaction behavior datasets in real scenarios. Through the statistical analysis of the labeled categories of the datasets, it is found that the datasets all show obvious uneven data distribution.

[0131] Specifically, as Figure 2As shown, the CE-HIRD dataset is a public dataset of human interaction behaviors in a student chemistry experiment scenario, containing 160 long videos that capture 8 experimental activities carried out by 20 human subjects in a chemistry experiment scenario. In this paper, the shooting content is based on the eight compulsory experiments in junior high school in the "Compulsory Education Chemistry Curriculum Standard", and a real chemistry experiment bench scenario is simulated and built. 20 junior high school students are recruited as experimental personnel, and a total of 682 minutes of large-scale chemistry experiment video data is shot. Due to the large amount of labeling information, a semi-automated labeling method is adopted in the labeling process to ensure the labeling quality of the dataset and significantly reduce the time cost. Specifically, in this paper, the ffmpeg tool is first used to extract frames from the chemistry experiment video at a frame rate of 25 FPS, and the resolution of each frame image is 1920×1080, with a total of 204,600 frames; secondly, the target category and position are marked for each frame. The specific target categories are beaker, narrow-mouth bottle, test tube, wide-mouth bottle, dropper, alcohol lamp, glass rod, glass slide, bottle stopper, iron stand, forceps, test tube brush, pH test paper, spatula, mortar, funnel, measuring cylinder, person, a total of 18 categories of target data. After marking 3000 frames, a chemistry experiment object detection model is trained using the yolov8 algorithm, and this model is used to identify and locate the positions of students, equipment categories and positions in subsequent video frames. Since the obtained target category indexes do not correspond within a video frame sequence, in this paper, the bytetrack object tracker is subsequently used to track and align the detected target categories within the same video frame sequence. In terms of interactive behavior annotation, the target position information obtained above is used to crop the region frame sequence of the target. If the target region category is a person, the interactive behavior category is annotated, otherwise the annotation of the object function participation target is carried out. Among them, the interactive behavior categories mainly include pouring, putting down, picking up, standing up, oscillating, stirring, covering and inverting, a total of 8 categories. The object function participation categories are 9 in total, mainly including static, pourable, put-downable, pick-upable, stand-upable, oscillatable, stirrable, coverable and invertible. The specific annotation file structure is [video_id, seg_frames, label]. Among them, video_id represents the video serial number, seg_frames represents the starting frame of the video corresponding to the behavior sequence, and label includes the interactive category information of people and the function participation category information of equipment.

[0132] As Figure 3 shown, the CAD120 dataset is a public video dataset recording human interaction behaviors, containing 120 RGB-D long videos that capture 10 activities carried out by 4 human subjects in daily indoor environments, such as making cereal, etc. Each activity consists of a sequence of interactive behavior video segments. The dataset contains 10 common interactive behaviors and 12 object function availabilities in daily life, and the sample distribution is as Figure 3 shown in.

[0133] As Figure 4 shown, the Something-else dataset is a larger public dataset containing 112,795 videos covering 174 different interaction behavior categories, and the sample distribution is as shown Figure 4 in it. This dataset was created by expanding the SomethingV2 dataset with additional hand and object bounding boxes, aiming to perform combined action recognition to identify both activities and objects simultaneously.

[0134] For the CE-HIRD chemical dataset and the CAD120 public dataset, the Sub-activity F1 value and the Affordence F1 value, i.e., the interaction behavior F1 value and the object functionality availability F1 value, are used to evaluate the model performance. The F1 value is the harmonic mean of the precision P and the recall R, which can comprehensively evaluate the accuracy and recall ability of the model. For the Something-else public dataset, the Top-1 and Top-5 accuracies are used in this paper's experiments to evaluate the method performance, where the Top-1 accuracy indicates that the predicted interaction behavior with the highest probability is the true interaction behavior category, representing a correct prediction; the Top-5 accuracy indicates that the true interaction behavior category is included among the top five predicted interaction behaviors, representing a correct prediction.

[0135] Both the CE-HIRD dataset and the CAD120 dataset are video datasets of human interactions, and each frame contains labeled information on human interaction behaviors and object functionality participation. However, the Something-else dataset does not contain labeled information on object functionality participation. Therefore, different comparison algorithms are used for these two types of datasets, including GPNN, HierGAT, LIGHTEN, STIGPN, IcH-HOI, STGC, and this embodiment.

[0136] As shown in Table 1, in the experiment of CE-HIRD dataset, this embodiment enriches the contextual information of data-scarce interactive behavior by introducing global features and memory embedding loss, and achieves the highest experimental results in both interactive behavior value (Sub-activity F1) and object function availability value (Affordence F1). Traditional GPNN lacks the ability to model the temporal sequence of interactive information, and the temporal information in video interactive behavior detection has an important impact on the model effect, so it is not as good as other methods. HierGAT extracts hierarchical temporal features from video data to capture the temporal dynamics of character interactions, and LIGHTEN also considers the spatiotemporal features of the video. Compared with the GPNN method, the interactive behavior values ​​of the two methods in the CE-HIRD dataset are improved by 1.42% and 3.01%, respectively, and the object function availability values ​​are improved by 0.35% and 0.04%, respectively. STIGPN uses graph networks to mine spatial and temporal evolution to simulate the long-term dynamics of the target, which has greatly improved the effect compared with previous methods. IcH-HOI uses the context fusion and interactive state reasoning modules to enable the model to have temporal reasoning capabilities, and achieves comparable results to STIGPN. STGC is a further improvement of STIGPN. It captures spatiotemporal correlation by introducing a spatiotemporal feature enhancement module, which is improved by 1.79% and 1.07% respectively compared with STIGPN. Since the existing methods all mine video information from the perspective of local features, the problem of uneven data distribution is not considered. This embodiment introduces the global features of the data set to make up for the defect that local features are difficult to handle similar categories. Compared with the current optimal STGC method, it is improved by 2.29% and 1.03% respectively, achieving SOTA (State-of-the-Art) results.

[0137] Table 1. Experimental results on the CE-HIRD dataset.

[0138] .

[0139] Since the scenarios in the CE-HIRD dataset are chemical experiment scenarios, the similarity between actions among classes is greater compared to the CAD120 dataset of daily life scenarios. And in this embodiment, the problem of large inter-class similarity is solved through the fusion of global and local features. Therefore, in the experiments on the CAD120 dataset, as shown in Table 2, this embodiment has improved by 0.55% compared to STGC. The object function availability value is slightly lower than that of the STGC method, but in the CE-HIRD dataset, it has improved by 2.29% and 1.03% compared to it. Since the annotation and evaluation criteria of the CE-HIRD dataset and the CAD120 dataset are the same, the comparative experiments of this article on the CAD120 dataset are similar to those on the CE-HIRD dataset, and at the same time, the comparison between this embodiment and the VHOIP and DT-HOI methods is added. Among them, VHOIP introduces the CLIP prior knowledge, but due to the lack of consideration of the data distribution problem, the effect is lower than that of this embodiment method. DT-HOI expands the signed text description through a large model, and since the label information in this dataset is relatively scarce, richer text information can be obtained after expansion, improving the effect of the model, and its result is higher than the method of this embodiment proposed in this article.

[0140] Table 2. Experimental result table on the CAD120 dataset.

[0141] 。

[0142] As shown in Table 3, in the experiments on the Something-else dataset, the graph network-based method outperforms the spatio-temporal interaction network-based method. This is because the graph network can more effectively capture and model the complex spatio-temporal dependencies in video data, while the spatio-temporal interaction network-based method can capture spatio-temporal information by taking the bounding box coordinates as input. However, the spatio-temporal interaction network-based method has limitations in spatio-temporal feature extraction and modeling of complex interaction behaviors when facing long-range dependencies and subtle motion changes in videos. After introducing I3D, the spatio-temporal interaction network-based method can capture the spatial and temporal information in videos, providing a richer feature representation and significantly improving the method's performance. And I3D, STIN+OIE+NL combines the separately trained I3D model with the trained STIN+OIE+NL method to obtain the best results in the STIN series. Among the graph network-based methods, STGCN and STGC model the spatio-temporal relationships of videos through spatio-temporal graph networks, achieving better results than the methods without combining I3D. DT-HOI relies on inputting label information into a large model to expand text descriptions, thereby improving the ability to understand human interaction behaviors. However, the data labels in the Something-else dataset are relatively complex, and the improvement effect on the method is not obvious. Since the method proposed in this embodiment of the present paper considers the problem of uneven data distribution, compared with the current state-of-the-art STGC method, the accuracy in Top-1 and Top-5 has increased by 2.33% and 1.47%, achieving the best results and further proving the effectiveness and generalization of the method. In summary, after introducing global features and memory embedding loss, the model enriches the context information of data-scarce interaction behaviors and solves the problem of large similarities between interaction behavior categories to a certain extent.

[0143] Table 3. Experimental result table on the Something-else dataset.

[0144] 。

[0145] To verify the effectiveness of the Global Feature Branch (GB) and Memory Embedding Loss (MEL) proposed in this embodiment, ablation experiments were conducted on the CE-HIRD, CAD120, and Something-else datasets. The experimental results are shown in Table 4, where "GLFFN" represents the complete method proposed in this embodiment, "GLFFN w / o All" represents removing both the global feature branch and the memory embedding loss module in this embodiment, "GLFFN w / o MEL" represents only removing the memory embedding loss in this embodiment, and "GLFFN w / o GB" represents only removing the global feature branch in this embodiment. The local feature branch is the basic framework for implementing interaction behaviors, so the experiment does not include the part of removing the local feature branch.

[0146] After introducing the global feature branch and memory embedding loss in this embodiment, the network performance has been effectively improved. Specifically, compared with the method after removing the global feature branch, the recognition of interaction behaviors and the object function participation rate in the CE-HIRD dataset of this embodiment have increased by 3.39% and 0.27% respectively. The scores. It has increased by 0.6% and 0.25% on the CAD120 dataset. On the Something-else dataset, the Top-1 accuracy and Top-5 accuracy have increased by 1.33% and 1.35% respectively, proving that the global feature branch can provide more comprehensive data information and context knowledge for the model after introducing the knowledge of the entire dataset. The memory embedding loss expands the distance between similar interaction behavior categories and improves the discriminative ability of the model for such interaction behaviors. Compared with the method after removing the memory embedding loss, the recognition of interaction behaviors and the object function participation rate in the CE-HIRD dataset of this embodiment have further increased by 0.69% and 1.83%, and it has also brought improvements of 0.99% and 0.35% on the CAD120 dataset, and 0.65% and 2.06% on the Something-else dataset, proving the effectiveness of the memory embedding loss. Finally, when the global feature branch and the memory embedding loss are combined, the network performance has been significantly improved, indicating that there is a synergistic effect between the two components, effectively solving the problem of deviation in the recognition of intra-class homogeneity and inter-class heterogeneity caused by limited training data for some interaction behaviors.

[0147] Table 4. Ablation experiment results on the CE-HIRD, CAD120, and Something-else datasets.

[0148] 。

[0149] To further analyze the improvement effect of the method on data-scarce categories, the confusion matrix of the method in this embodiment is also compared with the STIGPN as the baseline model, as Figure 5 and Figure 6 shown. In the CE-HIRD dataset, the recognition effect of the method in this embodiment for various interaction behaviors is higher than 70%, and it is better than the baseline model algorithm. For the categories of "inversion", "oscillation", and "covering" with scarce interaction behavior data, the effects are improved by 18%, 20%, and 17% respectively, achieving a large improvement. This is mainly because the baseline model does not effectively focus on data-scarce samples during training, resulting in the recognition results being easily misjudged as other behaviors with abundant data, making the effects generally low. For example, in the baseline model, the "oscillation" interaction behavior is easily recognized as the "pouring" interaction behavior because the number of "oscillation" samples is small, and when the "oscillation" amplitude is large, it is easily misjudged as "pouring". The method proposed in this paper introduces a global feature branch and a memory embedding loss, which makes the recognition performance of the "pouring" action slightly decrease, but the recognition performance of the "oscillation" action is greatly improved. After analyzing the object function participation, it is found that the categories of "oscillatable" and "coverable" with scarce data have a large improvement, increasing by 11% and 9% respectively. To sum up, this shows that the method in this embodiment can reallocate resources for data-scarce samples and data-rich samples during training, and make better feature distinctions, optimizing the overall performance.

[0150] To further reveal the impact of the method in this embodiment on the intra-class and inter-class distances, the high-dimensional features of the chemical experiment interaction behavior categories in the baseline model and the memory network in this embodiment are also mapped to the 2D space through t-SNE, so as to visually show the impact of different methods on the feature space distribution, as Figure 7 and Figure 8 shown. In Figure 7 , the baseline model is difficult to distinguish the two interaction behaviors of "pouring" and "oscillation" circled by the dashed line, and their feature spaces overlap. While in Figure 8 , the inter-class distances of the "pouring" and "oscillation" actions in this embodiment increase significantly, and the boundaries of the feature space are clearer. In addition, the feature spaces of "put down" and "cover" and "pick up" and "inversion" also become clearer in the visualization diagram of the method in this embodiment. For categories such as "erect" and "stir", their intra-class distances have also been significantly improved. This proves that on the one hand, this embodiment can enrich the context information of categories and reduce the intra-class distance, and on the other hand, it can enhance the differences of similar features and expand the distance between two interaction behaviors, ultimately resulting in a significant improvement in the method performance.

[0151] The experimental results show that, compared with the baseline model, the TOP-1 accuracy of interaction behavior classification in this embodiment on the three datasets of Something-else, CAD120, and CE-HIRD has increased by 2.73%, 1.05%, and 4.08% respectively, proving the improvement effect of the method on the recognition accuracy of interaction behavior. And it has surpassed the current SOTA method in the Something-else and CE-HIRD datasets. By expanding the inter-class distance between similar samples through the memory embedding loss, the model's ability to understand data-scarce interaction behaviors has been enhanced, and it can reallocate resources between data-scarce samples and data-rich samples during training for better feature discrimination, optimizing the overall performance and improving the robustness of the model.

[0152] In this article, the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", "vertical", "horizontal", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the sake of clarity and convenience of describing the technical solution, so it cannot be understood as a limitation of the present invention.

[0153] In this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, in addition to the listed elements, and may also include other elements not specifically listed.

[0154] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A memory network enhancement method for experimental interactive behavior recognition, characterized by The steps include: S1. Obtain a video dataset containing multiple experimental videos, construct a global video graph of the experimental videos through a video-level graph network, and extract global features of each experimental interaction behavior category in all experimental videos in the video dataset; S2. Construct a frame-level local interaction graph of the experimental video through a region-level graph network to obtain the local features of a single experimental video. S3. Use global features to initialize the memory items of the memory network, project local features into the feature space of the memory items through the attention mechanism, and enhance them by fusing with global features. And use memory embedding loss to expand the inter-class distance of similar experimental interaction behaviors between memory items.

2. The memory network enhancement method according to claim 1, characterized in that: In step S1, the video-level graph network uses BiRNN to process the time series of frame-level features of all experimental videos to obtain visual features of each experimental video, and uses the visual features to instantiate video nodes of the experimental video. Every two video nodes are connected in pairs to obtain an initialized global video graph.

3. The memory network enhancement method according to claim 1, characterized in that: In step S1, the steps for extracting the global features of each experimental interaction behavior category are as follows: First, the video nodes of the global video graph are updated through the graph attention network, and the attention coefficient a between the video nodes is obtained through the attention mechanism ij ; Secondly, the final feature x of a single experimental video is obtained by updating the attention coefficient i ', , Among them, x i is the video node feature, W video is the attention mechanism weight matrix, σ is the activation function; Finally, after updating the visual features of all experimental videos using the graph attention network, the visual features of the experimental videos are grouped according to the label information of the experimental interaction behavior, and the global features of each experimental interaction behavior category are calculated using the following formula: , where c k is the global feature of the kth experimental interaction behavior category, N video is the number of all experimental videos, y i is the visual feature of the ith experimental video, and the global feature dataset C of all experimental interactive behavior categories is obtained, C={c k } k=1,...,K , where K represents the number of experimental interaction behavior categories.

4. The memory network enhancement method according to claim 1, characterized in that: In step S2, the region-level graph network uses FasterR-CNN or YOLOv8 to locate instances of operators or experimental equipment on each frame of the experimental video, extracts visual features, spatial features, and semantic features of all instances for splicing, constructs a frame-level local interaction graph of the experimental video, and connects the spliced ​​features within and between frames to obtain local features of a single experimental video.

5. The memory network enhancement method according to claim 4, characterized in that: The FasterR-CNN localization instance has a four-dimensional coordinate box and a target category. For each instance, the region of interest is cropped from the original experimental video frame according to its four-dimensional coordinate box and input into the pre-trained ResNet-50 model to extract the visual feature F v , Use a multi-layer perceptron MLP to transform the four-dimensional coordinates of each instance into the spatial features F of the corresponding dimensions s , Generate the semantic features F of the instance using a learnable word embedding model w .

6. The memory network enhancement method according to claim 4, characterized in that: The frame-level local interaction graph includes a set of operator or experimental equipment nodes in the frame image, and a set of operator or experimental equipment edges in the frame image. For the edges connecting operators and experimental equipment, their edge weights are initialized to 1, and the rest are 0. The intra-frame edge connection graph and the inter-frame edge connection graph are parsed using a graph attention network to obtain the intra-frame edge connection graph attention coefficient and the attention coefficient of the inter-frame connection graph , the local features of the final single experimental video are obtained by connecting the intra-frame and inter-frame splicing features: , , Among them, W inra With W inter are the intra-frame edge weight matrix and the inter-frame edge weight matrix, H represents the concatenated features of the instance’s visual features, spatial features, and semantic features, and [] represents a connection operation.

7. The memory network enhancement method according to claim 1, characterized in that: In step S3, the process of fusing and enhancing the local features with the global features is as follows: S31, using the global features of all experimental videos to construct a memory network M, the memory network has K memory items corresponding to the number of experimental interaction behavior categories, and each memory item is initialized with the global features of the corresponding experimental interaction behavior category; S32, input the local features of a single experimental video into the memory network, and use the attention mechanism method to integrate the local features Project to the feature space of the memory item to obtain a new feature representation ; S33, the memory item enhances the local features after projection by the following formula: , in, is the local feature projected into the memory item feature space, is the feature after memory item enhancement, c k is the global feature of the kth experimental interaction behavior category, c j Represents the global feature corresponding to the jth memory item of the memory network; S34. Obtain the label of the predicted experimental interaction behavior category through the softmax function.

8. The memory network enhancement method according to claim 1, characterized in that: In step S3, the cosine similarity loss L is calculated by the following formula: sim and the Euclidean distance loss L euc Memory embedding loss L as a memory network mem ; , Among them, λ is a hyperparameter; The cosine similarity loss L is obtained by minimizing the cosine similarity between the experimental interaction behavior features of different memory items through the L2 norm sim , , Among them, s ij Represents memory item c i and memory item c j The cosine similarity between them, s is the similarity matrix of all cosine similarities between the memory items of the memory network, and K is the number of memory items in the memory network; Calculate the Euclidean distance between memory items, obtain the distance matrix of all Euclidean distances between memory items of the memory network, and sort each row of the distance matrix in ascending order to obtain D'=(d ij ')∈R K×K , d ij 'Indicates memory item c after sorting i and memory item c j The Euclidean distance between the two items is increased by selecting the k2 minimum values ​​in each row of the distance matrix through the following formula: , Among them, d — is the enlarged Euclidean distance, and the Euclidean distance loss L euc =max(0,-d — +γ), where γ is a hyperparameter for adjusting the distance between memory items.

9. An experimental interactive behavior recognition method, characterized in that: Input the experimental video to be identified into the memory network enhanced by the memory network enhancement method according to any one of claims 1 to 8, retrieve the memory items corresponding to the experimental interactive behaviors from the memory network in combination with the currently input video data, and output the experimental interactive behavior category retrieved in the experimental video to be identified.

10. Experimental interactive behavior recognition system, characterized by include: An input module, used to input the experimental video to be identified; A memory module storing a memory network enhanced by the memory network enhancement method according to any one of claims 1 to 8; The query module retrieves the memory items corresponding to the experimental interaction behaviors from the memory network in combination with the currently input video data; The output module outputs the experimental interaction behavior category retrieved from the experimental video to be identified.

Citation Information

Patent Citations

  • Image description method

    CN110390363A

  • Unsupervised video target segmentation method based on local and global memory mechanism

    CN113269021A

  • Image description method and device based on local representation enhancement, storage medium and terminal

    CN115131802A

  • Pedestrian re-identification method and device based on Transform network

    CN115909408A

  • Memory-enhanced difficult video motion small target detection method

    CN116912290A

Cited By

  • Multi-modal fusion character interaction recognition method and system based on large model

    CN122023915A

  • A large model-based multi-modal fusion character interaction recognition method and system

    CN122023915B