Method and device for recognizing complex action based on learnable markov logic network
Through an action reasoning framework based on a learnable Markov logic network, logical rules are automatically generated and weighted, which solves the interpretability and robustness problems of deep learning models in video action recognition, and realizes efficient and explainable action recognition and positioning.
Patent Information
- Application Number
- CN202210027024.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-11
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-01-11
AI Technical Summary
Existing deep learning models lack interpretability and robustness in video action recognition, have difficulty explicitly identifying the timing, location, and cause of actions, and are vulnerable to adversarial attacks.
A method based on learnable Markov logic networks is adopted. By designing an explainable action reasoning framework, first-order logic is used to model the temporal changes of complex actions, logical rules are automatically generated and assigned weights, and probabilistic logical reasoning is performed to identify actions.
It improves the interpretability and robustness of action recognition, can locate the position of actions in videos, and efficiently recognize actions in big data scenarios. It is compatible with existing deep models to improve performance.
Smart Images

Figure CN116469155B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of computer vision, and particularly relates to a complex action recognition method and device based on a learnable Markov logic network. BACKGROUND
[0002] Action recognition is a fundamental task in the field of video understanding, which has attracted great attention from researchers in the past few years. Recently, with the rapid development of deep learning, 3D convolutional neural networks (3D CNNs) have revolutionized this research field. By relying on various carefully designed network architectures and learning algorithms, it has become the mainstream method for video action recognition tasks. Compared with early work based on low-level features (such as trajectories, key points, etc.), the powerful representation ability of 3D CNNs enables them to better capture complex semantic dependencies across video frames.
[0003] Although these deep neural networks have achieved widespread application in video action recognition tasks, they still have some inherent defects. Generally speaking, the working process of 3D CNN is as follows: input a video segment, after the calculation of multiple layers of network, output a score, which represents the confidence of each action category. As can be seen, this black-box nature of the prediction mechanism does not explicitly provide the relevant basis for identifying an action, such as the time, location, and cause of the action in the video, etc. In addition, due to the lack of explainability, these deep neural networks are also vulnerable to adversarial attacks, which greatly limits their application in real-world scenarios with strict security requirements. In recent years, more and more research work has been devoted to exploring the explainability of deep learning. Therefore, it is particularly important to develop an action reasoning framework with high explainability.
[0004] The present application is based on some research conclusions in cognitive science and neuroscience: that is, people usually represent a complex event as a combination of some atomic units. In addition, recent studies have shown that a complex action can be decomposed into a series of spatio-temporal scene graphs, which depict how a person interacts with the surrounding objects over time. For example, a person can be represented as a node in the graph, and the surrounding objects can be represented as other nodes. The edges between nodes represent the interactions between them. For example, a person can be represented as a node in the graph, and the surrounding objects can be represented as other nodes. The edges between nodes represent the interactions between them. Figure 1The "person wakes up in bed" action is shown as an example. To complete this action, a person can initially lie on the bed, then wake up and sit on the bed. This process can be represented by the change in the visual relationship between the person and the bed over time, i.e. from "person-lying on-bed" to "person-siting on-bed". Such characteristics enable the model to explicitly identify the occurrence of an action by detecting the transition pattern of the visual relationship in the video, thereby significantly improving its interpretability and robustness. To achieve this idea, the invention needs to address two key challenges: (1) how to automatically learn this visual relationship transition pattern from data, rather than manually specify these rules by spending a lot of effort. (2) The rules generated by the model often contain some noise information, how to avoid the negative impact of these noises, so as to perform efficient action reasoning. SUMMARY
[0005] To make up for the lack of interpretability of deep models and solve the two challenges mentioned above, the invention discloses a complex action recognition method and device based on a learnable Markov logic network, which identifies complex actions in a video by designing a novel interpretable action reasoning framework. To this end, the invention uses first-order logic to model the temporal changes of complex actions in semantic states. Specifically, in each logic rule, the visual relationship serves as the corresponding atomic predicate. These logic rules contain rich information, which can be automatically generated by a rule strategy network by gradually adding action-related relational predicates. Since the rules are automatically generated and have not been carefully defined by domain experts, they are prone to errors. To solve this problem, the invention uses Markov logic networks (MLN), which are a statistical relational model that combines first-order logic and probabilistic graphical models. The model associates each logic rule with a real-valued weight, which measures the uncertainty of the logic rule. If the weight is larger, the corresponding rule is more reliable. In this way, those formulas with noise information can be assigned a lower (even negative) weight, thereby reducing their negative impact. Using the generated formulas and Markov logic networks, the invention can perform probabilistic logic reasoning to ultimately determine the probability of occurrence of each action.
[0006] The technical content of the invention includes:
[0007] A complex action recognition method based on a learnable Markov logic network, the steps of which include:
[0008] A strategy network is used to automatically learn the set of logic rules corresponding to each action from the training data;
[0009] cutting the video to be detected into several video clips, and calculating a confidence score for each <action participant, visual relation, object> triple in each video clip;
[0010] inputting the set of logical rules and the confidence scores of all triples in a video clip into an improved Markov logic network to obtain a probability of occurrence of each action in the video clip, wherein operations between Boolean variables in the Markov logic network are replaced by functions defined on continuous variables to obtain the improved Markov logic network;
[0011] obtaining an action recognition result of the video to be detected according to the probability of occurrence.
[0012] Further, the set of logical rules is obtained by the following steps:
[0013] 1) at time t, calculating an embedding feature x t-1 of a relation predicate R t-1 obtained at the last time;
[0014] 2) inputting x t-1 and a hidden state h t-1 into a gated recurrent unit (GRU);
[0015] 3) calculating a generation probability of the relation predicate R t at time t according to an output of the GRU;
[0016] 4) sampling a specific relation predicate R t using the generation probability;
[0017] 5) obtaining a sampled probability of the formula f according to the generation probability of the relation predicate R t at each time;
[0018] 6) putting one or more formulas f into the formula set of the action based on the sampled probability to obtain the set of logical rules corresponding to the action.
[0019] Further, the video to be detected is cut by the following strategies:
[0020] 1) generating sliding windows with multiple different sizes;
[0021] 2) for a sliding window with a size of L, setting a sliding step of the sliding window as L / 2;
[0022] 3) cutting the video to be detected according to the sliding step to generate video clips with a length of L.
[0023] Further, the <action participant, visual relation, object> triple is obtained by the following steps:
[0024] 1) For the video clip, uniformly sample M video frames;
[0025] 2) Use the Faster-RCNN detector with ResNet-101 as the backbone network to detect the object o in the video frame i ;
[0026] 3) Detect the object o in each video frame i The jth visual relation e with all action participants p ij , get the action participant p, visual relationship e on the sampling frame ij 、Object i >Triple.
[0027] Furthermore, the confidence score is calculated by the following steps:
[0028] 1) For the generated triple <action participant p, visual relationship e ij 、Object i >, calculate the confidence score s of the action participant p p 、Object i Confidence score of and visual relationships ij Confidence score of
[0029] 2) According to the confidence score s p , confidence score and confidence scores Compute the confidence score for the entire triplet.
[0030] Furthermore, the occurrence probability of each action in the video clip is obtained by the following steps:
[0031] 1) According to the function defined on continuous variables and the transformation criteria in first-order logic, the formula f in the formula rule is transformed into a Horn clause;
[0032] 2) Based on the Horn clause and the confidence score, calculate each formula f i The value of the instance;
[0033] 3) According to formula f i The value of the instance, get the formula f i The number of true values n i ;
[0034] 4) Based on the number n i , calculate the occurrence probability of each action in the video clip.
[0035] Furthermore, by performing a maximum pooling operation on the occurrence probability of each action in each video clip, the result of action recognition on the entire video is obtained.
[0036] Furthermore, the improved Markov logic network and the policy network that generates the rules for each action are trained through the following steps:
[0037] 1) Rule-based policy network π l Generate a set of logical rules Obtaining an improved Markov logic network by maximizing the log-likelihood method The weight of
[0038] 2) Fixed improved Markov logic network The weight of the rule policy network is obtained by using the policy gradient algorithm and maximizing the reward function to update the rule policy network parameters. l+1 , where the reward function is an action recognition evaluation index;
[0039] 3) When the rule strategy network π l and improved Markov logic network When the set conditions are met, the trained rule strategy network and improved Markov logic network are obtained.
[0040] A storage medium stores a computer program, wherein the computer program is configured to execute the method described above when running.
[0041] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer to execute the method described above.
[0042] Compared with the prior art, the advantages of the present invention are:
[0043] (1) Superior interpretability. Compared to currently popular deep 3D convolutional neural networks, the proposed action inference framework has significant interpretability because weighted logical rules can serve as an important basis for identifying specific actions. In addition, thanks to the explicit modeling of the temporal evolution of actions, the framework of the present invention can not only identify the category of the action, but also locate its position in the video clip.
[0044] (2) No need to rely on the definition of domain experts. The rule policy network proposed in the present application can automatically learn the logical rules for encoding complex actions from data without the need for artificial definition, making the entire framework more robust. Some existing work that uses Markov logic networks for reasoning often relies on the careful design of domain experts to encode event rules, which greatly limits their applicability. The feature of automatically mining rules from data also makes the reasoning framework of the present application applicable to big data scenarios.
[0045] (3) Compatibility and efficiency. The present method can be well combined with the existing deep model-based method to further improve the performance of action recognition. In addition, the model of the present application can mine the relationship change patterns corresponding to the action without too much training data, and obtain good prediction results. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 An example diagram for decomposing actions into spatio-temporal scene graphs.
[0047] Figure 2 The computational flow of the entire method.
[0048] Figure 3 Visualization of the rules and corresponding weights learned by the model.
[0049] Figure 4 User survey results for the model of the present application. DETAILED DESCRIPTION
[0050] In order to more specifically illustrate the technical details and advantages of the present application, the present application will be further described in detail below through examples and drawings.
[0051] As mentioned earlier, complex actions can usually be decomposed into the interaction of people and objects over time. Inspired by this conclusion, the present application models the evolution pattern of such visual relationships and further invents an interpretable action reasoning framework. As shown in Figure 1 The method proposed in the present application mainly consists of two main parts. The first is the rule policy network, the purpose of which is to mine the optimal formula set for each action, each formula of which explicitly represents a specific relationship transformation pattern. The second is the action reasoning module, which uses Markov logic networks to perform probabilistic logical reasoning based on the formula set generated by the policy network to calculate the probability of each action occurring. Next, the present application will describe the implementation details of each module and the training algorithm of the entire framework.
[0052] 1. Markov logic network
[0053] A Markov Logic Network (MLN) is a probabilistic graphical model that combines logic, which uses first-order logic to define the potential functions in a traditional Markov Random Field. In a Markov Logic Network, each logical formula has an associated real-valued weight, indicating the importance and reliability of the formula. A formula with a higher weight is often more important and its encoded knowledge is more reliable. Essentially, Markov Logic Networks relax the hard constraints in first-order logic, allowing some less reliable or even incorrect formulas to be included: not impossible, but less likely.
[0054] Specifically, let denote a set of logical formulas, ω i denote the weight corresponding to formula be a finite set of constants. Then, a Markov Logic Network follows the following definition: f i Each possible assignment of each atomic predicate can be viewed as a binary node in , and the value of the binary node is 1 if the corresponding assignment of the logical predicate is true, and 0 otherwise. Each possible assignment of each formula f i serves as a potential function, with a value of 1 if the formula is true, and 0 otherwise. Therefore, there is an edge between two nodes in a Markov Logic Network if and only if the logical predicates corresponding to the two nodes both appear in a formula. The set of formulas can be viewed as a template for constructing a Markov Logic Network. With this definition, the probability corresponding to a state x can be represented as
[0055]
[0056] where n i (x) is the number of formulas f i that are true in assignment x. F is the size of the set of formulas , and Z is a normalization constant with a value of
[0057] 2. Logic rule generation
[0058] Unlike some methods that use human-defined logical formulas, the goal of the present invention is to automatically generate the corresponding logical formulas for each action without relying on any human effort. Specifically, the present invention uses the following form to model the human-object interaction pattern: R1 ^... ^ R t ... ^ R T where R 1:T Represents the relationship predicates on different frames, and T represents the total number of these predicates. Furthermore, the formula f for encoding a complex action a can be expressed as:
[0059]
[0060] Where A is the predicate representation of action a. Therefore, given a specific action predicate A, only the left side of f needs to be determined. It only contains conjunction operations (∧), so it can be further expressed as a linear sequence Based on the above definitions, the generation of formula f is transformed into a sequence decision process: that is, predicting the most appropriate sequence l for each action f To achieve this goal, the present invention uses a policy network π to model this process. The policy network is used to approximately estimate the probability distribution π(f|a; θ) that all possible formulas f should satisfy for action a, where θ is the parameter corresponding to the probability distribution. Once θ is determined, the present invention can accordingly extract some samples from π(f|a; θ) to form the required set of formulas. The present invention uses a gated recurrent neural network (GRU) to express this probability distribution. Specifically, the network can be expressed as:
[0061] h t =GRU(x t , h t-1 ) (3)
[0062] where x t is the relational predicate R t The embedding feature at time t-th, h t-1 represents the hidden state corresponding to the policy network π, which integrates all the past time relation predicates {R1, ..., R t-1 In the initial step, the present invention inputs the feature vector x0 of the action predicate A into π, and then each predicate R t The generation probability of is calculated by the following formula:
[0063] p(R t |R1, ..., R t-1 , A)=softmax(W p h t ) (4)
[0064] Where W p is the parameter to be learned from the data. During the training process, the present invention can sample a sequence according to the above probability To get a specific formula f. Therefore, the probability of each formula f being sampled is:
[0065]
[0066] After training the policy network π, the present application utilizes beam search strategy to sample k best formulas for each action a from the distribution π(f|a; θ) as the generated formula set
[0067] 3. Action reasoning
[0068] This section mainly introduces the detailed probabilistic reasoning process for actions. The whole reasoning module mainly contains three steps (see Figure 1 ). Next, the present application will introduce them respectively.
[0069] (Step1) Video segment generation based on sliding window. Given an untrimmed long video v, the present application first utilizes the sliding window mechanism to process v to generate multiple video segments. Given that different types of actions often present great changes in time span, the present application sets the size of the sliding window to multiple different sizes to generate video segments of different lengths. In addition, for a sliding window of size L, the present application sets its sliding step size to L / 2, so that each video segment has L / 2 frames overlapping with the adjacent segment. Denote the set of all video segments generated by the sliding window as U, as the candidate proposals of possible actions in the video v.
[0070] (Step2) Scene graph prediction. For each video segment u ∈ U generated in the last step, the present application utilizes a pre-trained scene graph predictor to extract high-level visual information on the video frames. Specifically, the predictor first utilizes a Faster-RCNN detector with ResNet-101 as the backbone network to detect all objects in each frame. Then, it predicts all possible visual relationships between these objects and people. The generated scene graph can be denoted as G = (O, E). Here O = {o1, o2,...} is the set of objects interacting with the action performer p, and E = {{e 11 , e 12 ,...}, {e 21 , e 22 ,...}} represents the visual relationships between people and objects, where e ij represents the jth visual relationship between the action performer p and the ith object o i . Here, due to the diversity of visual interactions, there may be multiple different types of visual relationships between each participant and object. It is worth noting that each triple r ij = <p, e ij , o i > can be regarded as a specific instance of its corresponding relationship predicate on the video segment. In addition, this instance rij confidence scores of the predicted actors p, objects o are given by:
[0071]
[0072] Here s p , are the confidence scores of the predicted actors p, objects o i and their relationships e ij in the scene graph. Considering that the visual relationships between people and objects rarely change in several consecutive video frames, if a scene graph is generated for each frame in a segment u, it will cause computational redundancy. Therefore, only M frames are uniformly sampled from the segment u e U to perform the above prediction.
[0073] (Step 3) Action probability inference. Given a trained Markov network the probability of each action a in the video segment can be inferred. According to equation (1), the entire probability needs to determine the value of the formula f i on the segment u is true n i (x). In the original Markov logic network, the value of the logical formula is obtained by logically operating on the binary predicate, which can only take discrete values 0 or 1. However, the relationship predicate instance of the present application adopts the real value specified in the formula which ranges from [0, 1]. This property makes it difficult for the present application to determine whether a formula instance should take the value 1 or 0. In order to ensure compatibility with the logical operation in first-order logic , the present application uses Lukasiewicz logic to relax the operation between Boolean variables into a function defined on continuous variables. The relaxed conjunction (A), disjunction (V) and negation can be defined as: X V Y = max(0, X + Y - 1), X V Y = min(1, X + Y), Using the above relaxation, n i (x) can be effectively calculated. Taking the formula on the left side of equation 2 as an example, according to the transformation criterion in first-order logic, such a formula can be first converted into a Horn clause:
[0074]
[0075] which can be regarded as the disjunction between the positive or negative predicates.
[0076] Then, based on the predicted scene graph on u, the value of each formula instance f i (x) is:
[0077]
[0078] Here is the confidence score obtained by formula (6). x a is a binary variable that takes value 0 or 1, indicating whether action a occurs or not. Thus, n i (x) can be obtained by summing up the values f i (x) of all formula instances. Thereafter, the probability of action a occurring on video segment u is given by:
[0079]
[0080] where F a is the number of formulas related to action a, MB x (a) represents the Markov blanket of a, which is the set of triples that appear with a in all formulas.
[0081] The final prediction result of the whole video v is obtained by performing a max-pooling operation on the segment set U.
[0082] 4. Joint training algorithm
[0083] The objective of the present application is to learn the most suitable Markov logic network To this end, the training scheme designed by the present application includes two main stages: rule exploration and weight learning. Due to the discreteness of the rule exploration stage, the policy network π cannot be directly optimized by backpropagation based on the final loss function. Therefore, the present application proposes a joint training strategy. In which, the rule exploration stage is optimized by the policy gradient algorithm in reinforcement learning, and the corresponding weights of the generated rules are optimized by supervised learning.
[0084] Suppose the present application samples a formula f from π(f|a; θ), then the present application can train the rule policy network by maximizing the expectation of the reward function:
[0085] J(θ) = E f~π (f|a; θ)[H(f)] (10)
[0086] Here H(f) is an indicator of evaluating the performance of action recognition such as mAP. Further, the gradient is: which can be estimated by Monte Carlo sampling:
[0087]
[0088] Here K is the number of samplings. In addition, the present application also introduces a baseline b, which is the H(f k) is replaced by the exponential moving average of the original reward function in equation (11). Thus, the original reward function in equation (11) is replaced by H(f k In addition, to encourage the diversity of rule exploration, the entropy regularization on π(f|a; θ) is also added to the final loss function.
[0089] The weight learning stage aims to learn a proper weight for the generated formula, which can be achieved by maximizing the log-likelihood:
[0090]
[0091] Here, N denotes the size of a batch of training data, x a is a binary variable, which is 1 if action a exists in the ith video v i , and 0 otherwise.
[0092] The whole training process will be performed alternately between rule exploration and weight learning. First, the invented method generates a set of formulas f using the initialized rule policy network π, performs weight learning, and then fixes the learned weights to calculate the action recognition accuracy to estimate the gradient in equation (11) to update the parameters of π. After that, the invented method generates a new set of formulas f using the updated π, and performs weight training. These two stages will be alternated for several times.
[0093] 5. Combination with deep models
[0094] An untrimmed video usually contains multiple actions, which may have some potential connections. Take a video in Charades as an example, there may be some reasonable connections between the actions “holding a broom”, “putting the broom somewhere” and “cleaning something on the floor”: when a person is cleaning something on the floor, he may hold a broom, and then put the broom back after finishing the cleaning. Therefore, the method proposed by the invented method can be combined with the output of a deep model as a reasoning layer, so as to enhance the recognition of difficult-to-detect action classes (e.g., cleaning something on the floor) based on the prediction results of easy-to-detect actions (e.g., holding a broom). Specifically, the framework of the invented method can be used to learn some logical formulas and corresponding weights to represent the connections between these actions. Given the confidence scores output by the deep model, the invented method regards the detection results with high confidence as observed evidence, and performs probabilistic reasoning on other action classes, thereby improving the detection accuracy.
[0095] 6. Experimental results
[0096] To fully demonstrate the superiority of the technical solutions of the present application over the prior art, the present application is tested on two representative experimental data sets, Charades and CAD-120. The former is a large video data set consisting of about 9800 uncut videos, of which 7,985 are used for training and 1,863 are used for testing. These videos contain 157 complex daily activities involving 15 different indoor scenes. On average, each video contains 6.8 different action categories, usually with multiple action categories in the same frame, making identification extremely challenging. The latter is an RGB-D data set focusing on human daily life activities. It consists of 551 video clips and 32,327 frames, involving 10 different high-level activities (such as eating, assembling objects). For Charades, the present application calculates mAP (Mean Average Precision) to evaluate the detection performance of all action categories. For CAD-120, the present application uses the mAR (Mean Average Recall) indicator to measure whether the model successfully identifies the performed action.
[0097] Table 1 shows the action recognition results on Charades. As can be seen from it, the model of the present application achieves 38.4% mAP and surpasses powerful 3D CNN models such as I3D, 3D R-101 and Non-Local, etc. This shows that the model of the present application can make full use of the interaction information of actions in the time dimension through the generated formula and its weight. Benefiting from the pre-training on the large video benchmark Kinetics, the most advanced 3D model (such as X3D) achieves higher performance compared with the model of the present application, but the method of the present application exceeds the deep model pre-trained only on ImageNet (38.4% vs 21.0%). In addition, due to the accuracy limitation of the scene graph predictor, the present application also designs an Oracle version. This version assumes that the visual relationship in all video frames is correctly predicted. As shown at the bottom of Table 1, the Oracle version of the present application achieves a significant improvement in mAP performance (about 24%) and significantly exceeds all deep models, which proves the strong potential of the method of the present application. The present application also evaluates the integration with the deep model SlowFast (R-50). As can be seen, by utilizing the relationship between different actions, the model of the present application can further improve the performance of the deep model.
[0098] Table 1: Comparison of action recognition performance of different methods on Charades
[0099]
[0100] For the CAD-120 dataset, the present invention divides a long video sequence into small segments so that each segment contains only one action, and evaluates the average recall of each action. As shown in Table 2, the model of the present invention achieves the best result in mAR performance. Although the method Explainable AAR-RAR also adopts an explainable recognition framework, they are based on the domain expert-defined transition patterns and perform action reasoning by observing the specific state transition between the two adjacent consecutive frames. In contrast, the model of the present invention utilizes the logical rules learned from real data, which is more robust and efficient.
[0101] Table 2: Action recognition performance comparison of different methods on CAD-120
[0102]
[0103] The model of the present invention can provide convincing evidence to explain the reason for making such a prediction by identifying complex actions using explainable logical formulas. Therefore, according to the time when these evidences appear, the present invention can also locate the time when the action appears in the video. The present invention compares the results of this method with several advanced deep models on Charades. As can be seen from Table 3, the model of the present invention achieves superior action localization results. The performance of the present invention is better than that of the model pre-trained only on ImageNet (20.9% mAP vs 14.2% mAP). In addition, the present invention still obtains similar localization results as the model pre-trained on Kinetics. Although slightly weaker than the model in mAP performance, the action localization results of the present invention are more explainable.
[0104] Table 3: Action localization performance comparison of different methods on Charades
[0105]
[0106] To illustrate the explainability and diversity of the generated logical rules, the present invention illustrates the formulas and related weights learned by the model in Figure 3 . From Figure 3It can be observed that the formulas with higher weights usually provide better reasoning evidence for the action. For example, "holding broom→standing on floor→looking at floor" provides a clear inference evidence for detecting the action "tidying something on the floor". In addition, the present application also conducts a user survey on interpretability. In the user survey of the present application, the weights of the formulas generated by the model are uniformly divided into three categories according to the size, and the rules of each category are correspondingly denoted as good, medium and bad. Then, the present application samples 20 action categories from Charades, and randomly selects 1 formula from each type of formula as a representative of the type. 21 subjects participating in the user survey reorder the formulas in random order according to the relevance to the action. The statistical results of the user survey are shown in Table 1. Figure 4 As observed, the survey results show a high consistency between the learned formula weights and human common sense (e.g., 78.75% of the good rules are still marked as good by humans).
[0107] The above embodiments are only used to illustrate the technical solutions of the present application but not to limit the present application, and the ordinary skilled in the art can modify or equivalently replace the technical solutions of the present application without departing from the principles and scope of the present application, and the protection scope of the present application should be subject to the description of the claims.
Claims
1. A complex action recognition method based on a learnable Markov logic network, comprising the following steps: Use a policy network to automatically learn the set of logical rules corresponding to each action from the training data; The video to be detected is divided into several video segments, and the confidence score is calculated for the <action participant, visual relationship, object> triple in each video segment; Inputting the logic rule set and the confidence scores of all triples in a video clip into an improved Markov logic network to obtain the occurrence probability of each action in the video clip, wherein the improved Markov logic network is obtained by replacing the operational relaxations between Boolean variables in the Markov logic network with functions defined on continuous variables; Obtaining an action recognition result of the video to be detected according to the occurrence probability; The logical rule set is obtained by following the steps below: At time t, calculate the relational predicate R obtained at the previous time t-1 Embedded features x t-1 ; x t-1 and hidden state h t-1 Input into a gated recurrent neural network GRU; According to the output of GRU, calculate the relation predicate R at time t t The probability of generation; Use the generation probability to sample a specific relational predicate R t ; According to the relation predicate R at each moment t The generation probability of , obtain the probability of being sampled by formula f; Based on the probability of being sampled, one or more formulas f are put into the formula set of the action to obtain the logical rule set corresponding to the action; The occurrence probability of each action in the video clip is obtained by the following steps: According to the functions defined on continuous variables and the transformation criteria in first-order logic, the formula f in the formula rule is transformed into a Horn clause; Based on the Horn clause and the confidence score, calculate each formula f i The value of the instance; According to the formula f i The value of the instance, get the formula f i The number of true values n i ; Based on the number n i , calculate the occurrence probability of each action in the video clip; The modified Markov logic network and the policy network that generates the rules for each action are trained by the following steps: Rule-based policy network π l Generate a set of logical rules Obtaining an improved Markov logic network by maximizing the log-likelihood method The weight of Fixed Improved Markov Logic Network The policy gradient algorithm is used to update the parameters of the rule policy network by maximizing the reward function to obtain the rule policy network π l+1 , where the reward function is an action recognition evaluation index; When the rule policy network π l and improved Markov logic network When the set conditions are met, the trained rule strategy network and improved Markov logic network are obtained.
2. The method according to claim 1, wherein The video to be detected is segmented using the following strategy: 1) Generate sliding windows of various sizes; 2) For a sliding window of size L, the sliding step size of the sliding window is set to L / 2; 3) Cut the video to be detected according to the sliding step size to generate a video segment of length L.
3. The method according to claim 2, wherein The <action participant, visual relation, object> triple is obtained by the following steps: 1) For the video clip, uniformly sample M video frames; 2) Use the Faster-RCNN detector with ResNet-101 as the backbone network to detect the object o in the video frame i ; 3) Detect the object o in each video frame i The jth visual relation e with all action participants p ij , get the action participant p, visual relationship e on the sampling frame ij 、Object i >Triple.
4. The method according to claim 3, wherein The confidence score is calculated by the following steps: 1) For the generated triple <action participant p, visual relationship e ij 、Object i >, calculate the confidence score s of the action participant p p 、Object i Confidence score of and visual relationships ij Confidence score of 2) According to the confidence score s p , confidence score and confidence scores Compute the confidence score for the entire triplet.
5. The method according to claim 1, wherein By performing the maximum pooling operation on the occurrence probability of each action in each video clip, the result of action recognition on the entire video is obtained.
6. A storage medium storing a computer program, wherein: The computer program is configured to execute the method according to any one of claims 1 to 5 when executed.
7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 5.