Human body behavior recognition method and device based on multi-modal knowledge graph reasoning enhancement

Through the multimodal knowledge graph inference method, combined with visual and text information, a multimodal knowledge graph is constructed and feature fusion is carried out, which solves the problem of insufficient semantic understanding in the existing methods and achieves a more efficient behavior recognition effect.

CN120236331APending Publication Date: 2025-07-01XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510320959.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing deep learning-based behavior recognition methods lack effective external knowledge representation when dealing with behavior categories with complex semantics or visually similar, resulting in insufficient semantic understanding of complex actions by the model.

Method used

The multimodal knowledge graph inference method is adopted to construct a multimodal knowledge graph through a trained human behavior recognition network, combining the complementarity of visual information and text information, and a graph convolutional neural network and multi-layer perceptron are used to fusion of features to realize timing modeling under knowledge guidance.

Benefits of technology

The semantic understanding and space-time modeling capabilities of the model are improved, the accuracy of recognition of complex human behaviors is enhanced, and data costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236331A_ABST
    Figure CN120236331A_ABST
Patent Text Reader

Abstract

The invention discloses a human behavior recognition method and device based on multi-modal knowledge graph reasoning enhancement, and relates to the technical field of image processing, and the method comprises the steps: obtaining to-be-recognized video data; uniformly sampling to-be-identified video data to obtain a plurality of key frames; processing the plurality of key frames by adopting a trained human body behavior recognition network, and obtaining a category result of the video data to be recognized by utilizing complementarity between the visual information and the text information; wherein the trained human body behavior recognition network is obtained by training the initial human body behavior recognition network by taking data of a preset category as a training set. According to the method, the semantic comprehension capability and the space-time modeling capability of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a human behavior recognition method and device based on enhanced multi-modal knowledge graph reasoning. Background Art

[0002] Behavior recognition is an important research direction in the field of computer vision, and its goal is to accurately classify human behaviors in videos. In recent years, with the development of deep learning technology, behavior recognition methods based on deep neural networks have made significant progress in feature expression ability. However, these methods mainly rely on directly modeling visual information and still face challenges when dealing with behavior categories with complex semantics or similar visual appearances. To improve the model's ability to understand behavior semantics, researchers have begun to explore introducing external knowledge to assist the behavior recognition task. External knowledge can provide rich semantic information and category association information for behavior recognition, which helps the model establish a more comprehensive understanding of behaviors. Common knowledge-guided methods include using semantic information such as text descriptions and attribute annotations to enhance the expression ability of visual features.

[0003] Although existing methods attempt to enhance video representation by leveraging external knowledge, they have problems such as a single knowledge modality in constructing the representation of external action knowledge, separation of multi-modal knowledge, and lack of cross-modal interaction in the representation of external knowledge, resulting in the fact that the external knowledge representation cannot reasonably reflect the connections between motion concepts in the objective real world, restricting the model's semantic understanding ability for complex actions.

[0004] Therefore, there is an urgent need to provide a human behavior recognition method to improve the above-mentioned defects existing in the prior art. Summary of the Invention

[0005] To solve the above problems existing in the prior art, the present invention provides a human behavior recognition method and device based on enhanced multi-modal knowledge graph reasoning. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] In a first aspect, the present invention provides a human behavior recognition method based on enhanced multi-modal knowledge graph reasoning, including:

[0007] Obtain video data to be recognized;

[0008] Uniformly sample the video data to be recognized to obtain a plurality of key frames;

[0009] Process the plurality of key frames by using a trained human behavior recognition network, and utilize the complementarity between visual information and text information to obtain the category result of the video data to be recognized;

[0010] Among them, the trained human behavior recognition network is obtained by training the initial human behavior recognition network with data of preset categories as the training set.

[0011] In a second aspect, the present invention also provides a human behavior recognition device enhanced by multi-modal knowledge graph reasoning, including a processor, a memory, an input device, an output device, and a system bus. The processor, the memory, the input device, and the output device complete communication with each other through the system bus;

[0012] The memory is used to store computer programs;

[0013] The processor is used to implement the method provided above when executing the program stored in the memory.

[0014] Advantages of the present invention:

[0015] A human behavior recognition method and device enhanced by multi-modal knowledge graph reasoning provided by the present invention use a trained human behavior recognition network to process multiple key frames, and utilize the complementarity between visual information and text information to obtain the category result of the video data to be recognized, which can improve the semantic understanding ability and spatio-temporal modeling ability of the model.

[0016] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the drawings

[0017] Figure 1 is a flowchart of a human behavior recognition method enhanced by multi-modal knowledge graph reasoning provided by an embodiment of the present invention;

[0018] Figure 2 is a flowchart of obtaining knowledge-enhanced features provided by an embodiment of the present invention;

[0019] Figure 3 is a flowchart of obtaining video-level representations provided by an embodiment of the present invention;

[0020] Figure 4 is a flowchart of obtaining a preset knowledge graph provided by an embodiment of the present invention;

[0021] Figure 5 is a flowchart of preprocessing data of preset categories provided by an embodiment of the present invention;

[0022] Figure 6 is a schematic diagram of a human behavior recognition device enhanced by multi-modal knowledge graph reasoning provided by an embodiment of the present invention. Detailed implementation manners

[0023] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0024] Please refer to Figure 1 , Figure 1 which is a flowchart of a human behavior recognition method based on multi-modal knowledge graph reasoning enhancement provided by an embodiment of the present invention. A human behavior recognition method based on multi-modal knowledge graph reasoning enhancement provided by the present invention includes:

[0025] S101. Obtain video data to be recognized.

[0026] Specifically, in this embodiment, the video data to be recognized can be collected by a camera, and the video data to be recognized is a human action video; or human action videos can be collected on the Internet.

[0027] S102. Uniformly sample the video data to be recognized to obtain multiple key frames.

[0028] Specifically, in this embodiment, the video data to be recognized is uniformly sampled to obtain K key frames {f1,..., f K}.

[0029] S103. Process the multiple key frames using a trained human behavior recognition network, and utilize the complementarity between visual information and text information to obtain the category result of the video data to be recognized;

[0030] Among them, the trained human behavior recognition network is obtained by training an initial human behavior recognition network with data of a preset category as a training set.

[0031] Specifically, in this embodiment, the trained human behavior recognition network includes a pre-trained image encoder, a trained graph convolutional neural network, and a trained multi-layer perceptron; processing the multiple key frames using the trained human behavior recognition network and utilizing the complementarity between visual information and text information to obtain the category result of the video data to be recognized includes:

[0032] S1031. Process the multiple key frames using the pre-trained image encoder to obtain visual semantic anchor features.

[0033] Specifically, the multiple key frames are processed using a pre-trained image encoder (CLIP) to obtain a visual semantic anchor feature sequence Drive the dynamic reasoning of a preset knowledge graph through the visual semantic anchor features.

[0034] S1032. Add the visual semantic anchor features to a preset knowledge graph. By calculating the similarity matrix between the visual semantic anchor features and the existing node features in the preset knowledge graph, construct the connection relationship between the visual semantic anchor features and the existing node features in the preset knowledge graph, and update the preset knowledge graph. This dynamic expansion mechanism allows the model to adaptively activate relevant knowledge nodes according to the content of the input video data to be recognized.

[0035] Specifically, add the visual semantic anchor features to the preset knowledge graph;

[0036] Define the updated preset knowledge graph as G'=(V', E'), where V' = V ∪ {x1,..., x K}, V = V t ∪V v ;

[0037] Calculate the similarity matrix S between the visual semantic anchor features and the text node features in the preset knowledge graph tv , and its expression is:

[0038]

[0039] Calculate the similarity matrix S between the visual semantic anchor features and the visual node features in the preset knowledge graph vv , and its expression is:

[0040]

[0041] where X represents the visual semantic anchor feature sequence, (·) T represents transpose, C represents the action categories of the dataset, that is, there are a total of C categories of actions, represents the vector space composed of real matrices of size K×C. In the visual semantic anchor generation stage, K key frames will be uniformly sampled from the input video V;

[0042] According to the similarity matrix S tv and the similarity matrix S vv , construct the connection relationship e ik between the visual semantic anchor features and the text node features and visual node features, and its expression is:

[0043]

[0044] where τ represents the temperature parameter, which is used to adjust the smoothness of the feature similarity distribution, σ represents the sigmoid activation function. When e ik >τ n , the m-th visual semantic anchor feature and the text node feature t k and the visual node feature vk establish a connection between τ n represents a threshold value.

[0045] It should be noted that existing methods take into account the importance of external knowledge guidance, so structured external knowledge such as knowledge graphs is constructed, and knowledge reasoning mechanisms are used to promote behavior recognition tasks. However, existing methods usually treat knowledge guidance and temporal feature extraction as two independent steps, lacking an effective mechanism to integrate knowledge information into the temporal modeling process. This results in the model being difficult to capture the rich semantic information contained in the behavior evolution process and affects the understanding ability of complex temporal behaviors. This embodiment proposes a knowledge-aware temporal feature enhancement module, which organically integrates knowledge guidance into the temporal modeling process to further help the model capture spatio-temporal dynamic information.

[0046] S1033. Process the updated preset knowledge graph using the trained graph convolutional neural network to obtain the knowledge-enhanced features corresponding to the visual semantic anchor features propagated by the trained graph convolutional neural network through multiple layers, as Figure 2 shown, Figure 2 is a flowchart for obtaining knowledge-enhanced features provided by an embodiment of the present invention.

[0047] Specifically, use the trained graph convolutional neural network to perform message passing on the extended graph to achieve the propagation and aggregation of knowledge, and finally obtain the knowledge-enhanced representation of the visual semantic anchor features. Specifically, perform L-layer graph convolutional neural network propagation calculation on the extended graph G':

[0048]

[0049] where is the adjacency matrix with self-connections added to the extended graph, is 's degree matrix, H (0) is the initial node feature matrix, including the features of the original nodes and the anchor nodes, W (l) is the learnable parameter matrix of the l-th layer. Thus, after L-layer propagation, the knowledge-enhanced features {z1,..., z k} of the anchor nodes are obtained.

[0050] S1034. Process the visual semantic anchor features and the knowledge-enhanced features using the trained multi-layer perceptron, calculate the importance weights for each time step, and fuse and aggregate the visual semantic anchor features and the knowledge-enhanced features according to the importance weights for each time step to obtain the video-level representation, that is, the class result of the video data to be recognized, as Figure 3 shown, Figure 3 is a flowchart for obtaining the video-level representation provided by an embodiment of the present invention.

[0051] Specifically, in order to effectively fuse visual semantic anchor features and knowledge enhancement features, this embodiment provides a knowledge-aware temporal attention calculation mechanism to analyze the aggregation weights of video frames under the guidance of knowledge graph reasoning.

[0052] Calculate the importance weights for each time step, including:

[0053] According to the visual semantic anchor feature x k and the knowledge enhancement feature z k , calculate the importance weight α k for each time step, and its expression is:

[0054] α k = softmax(W2·ReLU(W1h k + b1)+ b2);

[0055]

[0056] Among them, W1, W2, b1, and b2 represent the learnable parameters of the multi-layer perceptron.

[0057] According to the importance weights of each time step, fuse and aggregate the visual semantic anchor features and knowledge enhancement features to obtain the video-level representation, including:

[0058] According to the importance weight α k for each time step, fuse and aggregate the visual semantic anchor feature x k and the knowledge enhancement feature z k to obtain the video-level representation v, and its expression is:

[0059]

[0060] Among them, γ represents the learnable balance parameter for adjusting the contribution degrees of the two features, and K represents the total number of visual semantic anchor features or knowledge enhancement features.

[0061] In this embodiment, as Figure 4 shown, Figure 4 is a flowchart for obtaining a preset knowledge graph provided by an embodiment of the present invention. The obtaining process of the preset knowledge graph includes:

[0062] Construct a multi-modal knowledge graph G = (V, E); optionally, in this embodiment, the obtaining of the knowledge graph utilizes more external information, and the acquisition costs of these external information (visual information and text information) are very low. Compared with similar methods, this embodiment makes full use of these low-cost acquired external information and greatly improves the performance of the model;

[0063] Define the node set V = Vt ∪V v , where V t = {t1,..., t C} represents the set of text nodes, and V v = {v1,..., v C} represents the set of visual nodes, and C represents the total number of behavior categories in the multimodal knowledge graph; the text description of the i-th behavior category is processed by a pre-trained text encoder to obtain the text node feature t i of the i-th behavior category, and its expression is:

[0064]

[0065] where d i represents the text description of the i-th behavior category, and TextEncoder(·) represents the pre-trained text encoder, represents the d-dimensional real vector space, represents the set of real numbers, that is, each component in the vector is a real number, and d indicates that the vector has d components, that is, its dimension. For example, represents the set of all real vectors of length 300. The above expression actually describes the vector space to which the text node feature t i belongs, that is, each t i is a d-dimensional real vector, which is a commonly used expression in mathematics;

[0066] For the i-th behavior category, N representative video frames {frame i1 , frame i2 , …, frame iN} are randomly selected from the training set, and the j-th video frame frame ij of the i-th behavior category is processed by a pre-trained video encoder to obtain the visual feature v ij of the j-th video frame, and its expression is:

[0067]

[0068] where ImageEncoder(·) represents the pre-trained image encoder;

[0069] The visual features of all video frames of the i-th behavior category are averaged and pooled to obtain the visual node feature v i of the i-th behavior category, and its expression is:

[0070]

[0071] Among them, as shown in Table 1, Table 1 shows the accuracy rates of the pre-trained image encoder and the pre-trained text encoder.

[0072] Table 1 Accuracy Rates of the Pre-trained Image Encoder and the Pre-trained Text Encoder

[0073] Text Encoder Image Encoder Do not enable this solution Top-1 Accuracy (%) CLIP Pre-trained Text Encoder ViT-B-16 Yes 97.17 CLIP Pre-trained Text Encoder ViT-B-16 No 96.90

[0074] Define the edge set E = E tt ∪E vv ∪E tv , where E tt represents the relationship set between text nodes, E vv represents the relationship set between video nodes, and E tv represents the relationship set between text nodes and video nodes. A direct connection will be established between the text nodes and visual nodes of the corresponding categories; among them, the relationship sim t (i, j) between text nodes and the relationship sim v (i, j) between video nodes are expressed as follows:

[0075]

[0076] When sim t (p, q) > τ t , a relationship is established between the p-th text node and the q-th text node. τ t represents the preset text modality connection activation threshold. When sim v (x, y) > τ v , a relationship is established between the x-th video node and the y-th video node. τ v represents the preset visual modality connection activation threshold; sim(·, ·) represents the cosine similarity function.

[0077] It should be noted that in the knowledge graph, the relationships between nodes are represented by the edges between the nodes. The edge set of this embodiment includes: (1) the relationships between text nodes constructed based on the cosine similarity of text features; (2) the relationships between visual nodes constructed based on the cosine similarity of visual features; (3) the cross-modal correspondence relationships between text nodes and visual nodes of the corresponding categories. This structural design enables the knowledge graph to capture both the semantic associations within the modality and the correspondence relationships between modalities.

[0078] The external knowledge represented by the knowledge graph constructed in this embodiment is not limited to the text modality, but also includes the unified visual representation inductively extracted from a series of action videos of the same type. This can give full play to the complementary advantages of visual and semantic information. Especially when dealing with human behaviors with rich visual-semantic associations, the multi-modal knowledge representation can often more completely depict the semantic connotation of the behaviors and achieve a more comprehensive semantic understanding.

[0079] In this embodiment, the training process of the trained human behavior recognition network includes:

[0080] Obtain data of multiple preset categories, and label the data of the preset categories to obtain the true labels of the data of the preset categories; optionally, the data of the preset categories can be collected by a camera or collected from the Internet, and the data of the preset categories is human action videos;

[0081] Preprocess the data of the preset categories, sample the preprocessed data of the preset categories by uniform sampling to obtain video frame data; perform data augmentation on the video frame data to obtain the processed video frame data, so as to construct a training set, a validation set and a test set;

[0082] Input part of the samples in the training set into the r-th human behavior recognition network to be trained for training, and obtain the prediction results output during the r-th training process;

[0083] Calculate the loss according to the prediction results output during the r-th training process and the true labels of the samples for training the r-th human behavior recognition network to be trained, and use it as the loss of the r-th training process;

[0084] Perform backpropagation according to the loss of the r-th training process to update the network parameters of the r-th human behavior recognition network to be trained, and obtain the (r + 1)-th human behavior recognition network to be trained; at the same time, every time a preset validation interval value is reached, input the samples in the validation set into the human behavior recognition network to be trained, obtain the output prediction results, and combine the true labels of the samples in the validation set to perform model selection and hyperparameter tuning. Iterate in this way until the number of training times or the convergence degree meets the preset conditions, and obtain the trained human behavior recognition network;

[0085] Input the samples in the test set into the trained human behavior recognition network, obtain the output prediction results, and combine the true labels of the samples in the test set to calculate the recognition accuracy of the trained human behavior recognition network to judge the performance of the trained human behavior recognition network.

[0086] In this embodiment, a contrastive cross-modal learning framework is used to train the entire model. Specifically, in video behavior recognition, in order to better utilize the semantic information contained in the text description to guide the learning of video features, a contrastive learning framework is used to construct a joint representation space for video-text. Specifically, by designing a bidirectional contrastive loss function, the model can simultaneously learn the mapping from video features to text features and the mapping from text features to video features, so as to achieve effective alignment of the two modal features.

[0087] The loss of the r-th training process including the loss from video node features to text node features and the loss from text node features to video node features whose expression is:

[0088]

[0089] where M represents the training batch size.

[0090] In this embodiment, as Figure 5 shown, Figure 5 is a flowchart of preprocessing data of a preset category provided by an embodiment of the present invention. For preprocessing data of a preset category, uniformly sample the preprocessed data of the preset category to obtain video frame data, including:

[0091] Considering the storage space limitation and processing efficiency, downsample the data of the preset category with a frame rate greater than or equal to 30fps to 30fps; keep the original frames for the data of the preset category with a frame rate less than 30fps;

[0092] For the data of the preset category with a duration of T seconds, uniformly extract 4 to 16 frames to balance the calculation efficiency and information integrity.

[0093] Optionally, a multimedia video processing tool such as FFmpeg can be used to extract frames from the data of the preset category, decode it into a continuous sequence of RGB images, store it according to the frame rate of the original data of the preset category, and save the extracted frame rate in a lossless compression format to ensure the image quality.

[0094] In this embodiment, data augmentation is performed on the video frame data to obtain the processed video frame data, including:

[0095] Perform normalization processing on the sampled video frames. First, scale the short side to 256 pixels and maintain the original aspect ratio, and then randomly crop a 224×224 area during the training phase and use center cropping during the testing phase. At the same time, introduce random horizontal flipping (probability set to 0.5) as a data augmentation means.

[0096] In summary, a human action recognition method based on multi-modal knowledge graph reasoning enhancement provided by the present invention has the following beneficial effects:

[0097] First, better semantic understanding ability; by constructing a multi-modal action semantic knowledge graph and designing and implementing knowledge reasoning driven by visual semantic anchors, the complementarity between visual and text information can be fully utilized, significantly enhancing the cross-modal semantic alignment effect and achieving more comprehensive action semantic understanding.

[0098] Second, better ability to represent action timing; by designing a knowledge-aware timing feature enhancement module, calculating the attention scores of video frames based on the original visual features and knowledge-enhanced features, and realizing the calculation of frame importance weights guided by knowledge, so as to organically integrate knowledge guidance into the timing modeling process, further helping the model capture spatio-temporal dynamic information, enhancing the model's representation ability in the spatio-temporal dimension, enabling it to effectively understand the timing relationship of complex human actions, and improving the action recognition accuracy.

[0099] Third, lower data cost and better application prospects; pre-trained vision-text models have learned rich vision-language knowledge on a large amount of data, which provides a strong prior knowledge basis for the action recognition task. The cross-modal alignment ability of the pre-trained model can help the model better understand the semantic connotation of actions, especially performing well in dealing with complex and abstract human actions; secondly, through the transfer of the pre-trained model, the demand for labeled data can be significantly reduced, and the cost of model training can be lowered.

[0100] Based on the same inventive concept, as Figure 6 shown, Figure 6 FIG. 10 is a schematic diagram of a human action recognition device enhanced by multi-modal knowledge graph reasoning provided by an embodiment of the present invention. The present invention also provides a human action recognition device enhanced by multi-modal knowledge graph reasoning, including a processor, a memory, an input device, an output device, and a system bus. The processor, the memory, the input device, and the output device complete communication with each other through the system bus;

[0101] The memory is used for storing computer programs;

[0102] The processor is used to implement the method provided by the above embodiment of the present invention when executing the program stored in the memory. For the embodiments of the method, please refer to the above, and details are not described herein again.

[0103] In this embodiment, the input device includes a keyboard for text input, input devices such as a mouse for user operation input, and also includes image input devices such as a camera.

[0104] The output device includes display devices such as a monitor and a projector for image display output, and other common output devices in the art such as a printer.

[0105] In some embodiments, the processor may be a Central Processing Unit (CPU), or it may be other processors, Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate, transistor logic devices, or discrete hardware components, etc. The processor generally serves as the control center of the electronic device and is connected to different parts of the electronic device through a data interface such as a bus.

[0106] The memory mainly includes a program storage area and a data storage area. In some embodiments, the memory may be Flash Memory, a hard disk, a multimedia card, a card-shaped memory, Random Access Memory (RAM), Static Random-Access Memory (SRAM), Read-Only Memory (ROM), electrically erasable programmable read-only memory, programmable read-only memory, magnetic memory, a disk, an optical disc, etc. The memory may also be an external storage device, such as a plug-in hard disk, a secure data card, a flash card, etc. The memory is used to store computer programs and is executed by the processor to complete the functions of the present invention.

[0107] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element. "Connection" or "connected" and other similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "up", "down", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation on the present invention.

[0108] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0109] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A human behavior recognition method based on multimodal knowledge graph reasoning enhancement, characterized in that: include: Obtaining video data to be identified; Uniformly sampling the video data to be identified to obtain a plurality of key frames; Using a trained human behavior recognition network to process the multiple key frames, and using the complementarity between visual information and text information to obtain a category result of the video data to be recognized; The trained human behavior recognition network is obtained by training an initial human behavior recognition network using data of a preset category as a training set.

2. According to claim 1, the method for human behavior recognition based on multimodal knowledge graph reasoning enhancement is characterized in that: The trained human behavior recognition network includes a pre-trained image encoder, a trained graph convolutional neural network and a trained multi-layer perceptron; the trained human behavior recognition network is used to process the multiple key frames, and the complementarity between visual information and text information is used to obtain the category result of the video data to be recognized, including: Processing the plurality of key frames using the pre-trained image encoder to obtain visual semantic anchor features; Add the visual semantic anchor feature to a preset knowledge graph, construct a connection relationship between the visual semantic anchor feature and the existing node features in the preset knowledge graph by calculating a similarity matrix between the visual semantic anchor feature and the existing node features in the preset knowledge graph, and update the preset knowledge graph; The updated preset knowledge graph is processed using the trained graph convolutional neural network to obtain knowledge enhancement features corresponding to the visual semantic anchor point features after propagation through multiple layers of the trained graph convolutional neural network; The trained multi-layer perceptron is used to process the visual semantic anchor features and the knowledge enhancement features, and the importance weight of each time step is calculated. According to the importance weight of each time step, the visual semantic anchor features and the knowledge enhancement features are fused and aggregated to obtain a video-level representation, that is, the category result of the video data to be identified.

3. The method for human behavior recognition based on multimodal knowledge graph reasoning enhancement according to claim 2 is characterized in that: The process of obtaining the preset knowledge graph includes: Construct a multimodal knowledge graph G = (V, E); Define the node set V = V t ∪V v , where V t ={t1,...,t C } represents a text node set, V v ={v1,...,v C } represents a visual node set; the pre-trained text encoder is used to process the text description of the i-th behavior category to obtain the text node feature t of the i-th behavior category. i , whose expression is: Among them, d i represents the text description of the i-th behavior category, TextEncoder(·) represents the pre-trained text encoder, represents a d-dimensional real vector space; For the i-th behavior category, N representative video frames are randomly selected from the training set {frame i1 ,frame i2 ,…,frame iN }, use the pre-trained video encoder to decode the jth video frame of the i-th behavior category ij Processing is performed to obtain the visual feature v of the jth video frame ij , whose expression is: Where ImageEncoder(·) represents the pre-trained image encoder; The visual features of all video frames of the i-th behavior category are averaged and pooled to obtain the visual node feature v of the i-th behavior category. i , whose expression is: Define edge set E = E tt ∪E vv ∪E tv , where E tt Represents the relationship set between text nodes, E vv Represents the relationship set between video nodes, E tv Represents a set of relationships between text nodes and video nodes; wherein the relationship between the text nodes is t The relationship between (i, j) and the video node sim v The expressions of (i,j) are: When sim t (p,q)>τ t When the pth text node and the qth text node establish a relationship, τ t Indicates the preset text modal connection effective threshold, sim v (x,y)>τ v When the xth video node and the yth video node establish a relationship, τ v represents the preset visual modality connection effective threshold; sim(·,·) represents the cosine similarity function.

4. The method for human behavior recognition based on multimodal knowledge graph reasoning enhancement according to claim 2 is characterized in that: Updating the preset knowledge graph includes: Adding the visual semantic anchor feature to a preset knowledge graph; The preset knowledge graph to be updated is defined as G'=(V',E'), where V'=V∪{x1,...,x K },V=V t ∪V v ; Calculate the similarity matrix S between the visual semantic anchor feature and the text node feature in the preset knowledge graph tv , whose expression is: Calculate the similarity matrix S between the visual semantic anchor feature and the visual node feature in the preset knowledge graph vv , whose expression is: Among them, X represents the visual semantic anchor feature sequence, represents transposition, C represents the behavior category of the data set, represents a vector space composed of real matrices of size K×C; According to the similarity matrix S tv And the similarity matrix S vv , construct the connection relationship between the visual semantic anchor feature and the text node feature and the visual node feature ik , whose expression is: Among them, τ represents the temperature parameter, σ represents the sigmoid activation function, when e ik >τ n , the mth visual semantic anchor feature and the text node feature t k and visual node features v k Establish a connection between n Indicates the threshold value.

5. The method for human behavior recognition based on multimodal knowledge graph reasoning enhancement according to claim 2 is characterized in that: The calculation of the importance weight of each time step includes: According to the visual semantic anchor feature x k and the knowledge-enhanced feature z k , calculate the importance weight α of each time step k , whose expression is: α k =softmax(W2·ReLU(W1h k +b1)+b2); Among them, W1, W2, b1 and b2 represent the learnable parameters of the multilayer perceptron.

6. The method for human behavior recognition based on multimodal knowledge graph reasoning enhancement according to claim 5 is characterized in that: According to the importance weight of each time step, the visual semantic anchor feature and the knowledge enhancement feature are fused and aggregated to obtain a video level representation, including: According to the importance weight α of each time step k , the visual semantic anchor feature x k and the knowledge-enhanced feature z k After fusion and aggregation, we get the video-level representation v, which is expressed as: Here, γ represents a learnable balance parameter and K represents the total number of visual semantic anchor features or knowledge-enhanced features.

7. The method for human behavior recognition based on multimodal knowledge graph reasoning enhancement according to claim 1 is characterized in that: The training process of the trained human behavior recognition network includes: Acquire a plurality of data of the preset categories, and label the data of the preset categories to obtain true labels of the data of the preset categories; Preprocessing the preset category data, sampling the preprocessed preset category data by uniform sampling to obtain video frame data; performing data enhancement on the video frame data to obtain processed video frame data to construct a training set, a validation set, and a test set; Inputting some samples in the training set into the human behavior recognition network to be trained for the rth time for training, and obtaining the prediction result output during the rth training process; According to the prediction results output in the r-th training process and the true labels of the samples of the human action recognition network to be trained for the r-th training, the loss is calculated and used as the loss of the r-th training process; Back propagation is performed according to the loss of the r-th training process to update the network parameters of the human behavior recognition network to be trained for the r-th time, and the human behavior recognition network to be trained for the r+1-th time is obtained; at the same time, at each preset verification interval value, the samples in the verification set are input into the human behavior recognition network to be trained to obtain the output prediction results, and the model selection and hyperparameter tuning are performed in combination with the real labels of the samples in the verification set, and this is iterated until the number of training times or the degree of convergence meets the preset conditions, and the trained human behavior recognition network is obtained; The samples in the test set are input into the trained human behavior recognition network to obtain the output prediction results. Combined with the true labels of the samples in the test set, the recognition accuracy of the trained human behavior recognition network is calculated to judge the performance of the trained human behavior recognition network.

8. The method for human behavior recognition based on multimodal knowledge graph reasoning enhancement according to claim 7 is characterized in that: The loss of the r-th training process Includes the loss of video node features to text node features And the loss from text node features to video node features Its expression is: Here, M represents the training batch size.

9. The method for human behavior recognition based on multimodal knowledge graph reasoning enhancement according to claim 7 is characterized in that: Preprocessing the data of the preset category, sampling the preprocessed data of the preset category in a uniform sampling manner to obtain video frame data, including: Downsampling the data of the preset category with a frame rate greater than or equal to 30fps to 30fps; keeping the data of the preset category with a frame rate less than 30fps in the original frame; For the data of the preset category with a duration of T seconds, 4 to 16 frames are evenly extracted.

10. A human behavior recognition device based on multimodal knowledge graph reasoning enhancement, comprising a processor, a memory, an input device, an output device and a system bus, characterized in that: The processor, the memory, the input device and the output device communicate with each other via the system bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 9 when executing the program stored in the memory.

Citation Information

Cited By

  • Multi-modal data set construction method and system

    CN120723918A