Figure behavior recognition method based on deep learning

Through the improved GroundedSAM model and pseudo-label generation technology, combined with optical flow information and pre-trained language model, the diversity and complexity problems in video action recognition are solved, and efficient and accurate video action recognition and natural language description are achieved.

CN120279600AActive Publication Date: 2025-07-08SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510500247.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-08
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing video action recognition methods are difficult to achieve efficient and accurate video action recognition when facing problems such as diversity, complexity, high labeling cost and large calculation overhead.

Method used

Through the improved GroundedSAM model, combining pseudo-label generation, motion feature extraction and tag-to-text generation, the pre-trained language model BERT is used for semantic expansion, an action association library is built, and combined with optical flow information and connectivity area analysis, the ROI region of interest is dynamically determined, and pseudo-labels with high semantic correlation are generated, and the target label is finally converted into natural language description.

Benefits of technology

It significantly improves the accuracy and robustness of video action recognition, reduces computing resource consumption and manual labeling costs, and improves the system's understanding of diversified and complex behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279600A_ABST
    Figure CN120279600A_ABST
Patent Text Reader

Abstract

The invention discloses a character behavior recognition method based on deep learning, and the method achieves the precise recognition of a character behavior through an improved Ground SAM model, and the improved Ground SAM model is obtained through the improvement of a pseudo tag generation module and a motion feature extraction module of an original Ground SAM model. Semantic extension is carried out on behavior tags in a labeled data set, and an action association library is constructed and used for pseudo tag input of an accurate model. And extracting an ROI (Region of Interest) of the motion feature optimization model in combination with optical flow information, and generating natural language description through a large language model. And finally, performing action prediction and classification by adopting a dual-channel classifier. According to the method, the recognition precision can be remarkably improved, and the time and calculation cost in the visual understanding process can be effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning behavior recognition, and in particular to a method for character behavior recognition based on deep learning. Background Art

[0002] With the explosive growth of video data, video content understanding and recognition technology has become one of the important research directions in the field of computer vision. Video action semantic understanding and recognition, as a core task, aims to automatically identify and understand the actions and semantic information of people from videos. This technology has broad application prospects in the fields of intelligent monitoring, video content analysis, human-computer interaction, virtual reality, etc.

[0003] Traditional video action recognition methods mainly rely on manually designed feature extraction and classifiers, such as optical flow method, spatiotemporal interest points (STIP), etc. Although these methods have achieved certain results in specific scenarios, they are difficult to cope with complex and changing video content and diverse action categories because they rely on manually designed features. In addition, traditional methods often face problems such as high computational complexity and poor generalization ability when processing large-scale video data.

[0004] In recent years, the rapid development of deep learning technology has brought new breakthroughs in video action recognition. Deep learning methods based on convolutional neural networks (CNN) and recurrent neural networks (RNN) can automatically learn the spatiotemporal features in videos, significantly improving the accuracy and robustness of action recognition. In particular, with the successful application of the Transformer model in the field of natural language processing (NLP), Transformer-based video action recognition methods have gradually become a research hotspot. These methods can better understand complex actions in videos by capturing long-range dependencies between video frames.

[0005] In the existing deep learning methods, semantic understanding and recognition of video actions still face many challenges. First, the actions in the video often show diversity and complexity, and a single feature extraction method is difficult to fully capture the semantic information of the action. Secondly, the annotation cost of video data is high, the existing annotation data set is limited in size, and it is difficult to cover all possible action categories and scenes. In addition, how to convert the action information in the video into natural language descriptions for human understanding and subsequent processing is also an urgent problem to be solved. At the same time, the problems brought about by direct recognition of video images cannot be ignored. Each frame in the video may contain a lot of details and information, but an efficient model requires complex analysis of this information, which leads to a significant increase in the cost of computing and model judgment, thereby affecting the real-time and accuracy of the model. Therefore, how to balance the comprehensiveness of information capture and the consumption of computing resources is a bottleneck that needs to be broken through in the current field of video action recognition. Summary of the Invention

[0006] The object of the present invention is to consider the problems of action diversity, complexity, high annotation cost and large computational overhead in video action semantic understanding and recognition, and propose a method for human behavior recognition based on deep learning. This method effectively solves the problems of false detection and missed detection caused by video complexity and action category diversity by introducing pseudo-label generation, motion feature extraction, and label-to-text generation, improves the accuracy of video action recognition, and effectively reduces costs.

[0007] To achieve the above object, the technical solution provided by the present invention is: a method for human behavior recognition based on deep learning, which uses an improved GroundedSAM model to achieve accurate recognition of human behavior. The improved GroundedSAM model improves the pseudo-label generation module and the motion feature extraction module of the original GroundedSAM model; the improvement of the pseudo-label generation module is: on the basis of automatically generating pseudo-labels, semantic extension is performed on the behavior labels in the labeled dataset to construct an action association library; the improvement of the motion feature extraction module is: a binary mask is initially generated by the frame difference method, and then the optical flow mask is obtained by calculating the optical flow information of the video frames, and the ROI (Region of Interest) of the model is adjusted.

[0008] The specific implementation of this human behavior recognition method includes the following steps:

[0009] 1) Through the pre-trained language model BERT, semantic extension is performed on the behavior labels in the labeled human behavior dataset, an extended label set is generated in combination with the context, and the labels in the extended label set are mapped into high-dimensional vectors by using a semantic embedding model to construct an action association library;

[0010] 2) Extract video frames from the video, generate a preliminary mask through pixel-level difference, and determine the bounding box of the motion region in combination with optical flow information and connected region analysis as the ROI of the improved GroundedSAM model to accurately determine the main motion region of the person in the video frame;

[0011] 3) Use the video frames extracted in step 2), and generate preliminary pseudo-labels through the RAM model. Combine the action association library constructed in step 1), perform semantic matching at the vector level between the preliminary pseudo-labels and the extended labels in the action association library, and use the RAG retrieval technology to refine the pseudo-labels, and finally generate a pseudo-label set with higher semantic relevance;

[0012] 4) Input the ROI region of interest obtained in step 2) and the video frame into the trained improved GroundedSAM model to generate a segmentation mask, and perform semantic matching with the pseudo-labels in the pseudo-label set to generate the target label set of the video frame;

[0013] 5) Based on the target label set generated in step 4), use the pre-trained large language model LLM and the refined Prompt template to convert the target labels in the target label set into natural language descriptions, and then generate accurate and coherent natural language text in combination with the context;

[0014] 6) In the classification stage, compare the natural language text generated in step 5) with the behavior labels in the pre-annotated human behavior dataset in step 1), and calculate the similarity and confidence scores to generate the final classification result.

[0015] Furthermore, in step 1), for the behavior labels in the pre-annotated human behavior dataset, use the pre-trained language model BERT for semantic expansion. The expansion process generates an extended label set E related to the input label T based on the context information by the large language model. The expansion formula is as follows:

[0016] E = B(T, C)

[0017] Where B represents the language model BERT, T is the input behavior label; C is the context information used to guide the relevance of generation; E is the generated extended label set {E1, E2,..., E i ,..., E n}}, where E i is the i-th extended label, i = 1, 2, 3..., n;

[0018] Map the extended label set E to a high-dimensional vector space through the pre-trained semantic embedding model F to obtain the vector representation of each label:

[0019] v i = F(E i )

[0020] Where v i is the i-th label vector

[0021] All the embedded vectors are stored in the action association library and used as the matching basis for subsequent pseudo-label generation. The action association library is defined as the vector set V:

[0022] V = {v1, v2,..., v n}}.

[0023] Further, in step 2), by using the video frames extracted from the video, and combining the pixel differences of the video frames with the changes in optical flow information, the ROI region is dynamically adjusted. The specific steps are as follows: From the original video data, continuous video frames F are extracted according to the frame extraction rule of Q frames per second t , where t represents the time index, and the frame extraction rule is expressed as:

[0024]

[0025] In the formula, V′ is the input video, Δ is the frame extraction interval, and F t is the extracted video frame;

[0026] For the extracted video frame F t , pixel-level differential processing is performed on adjacent frame images, and the frame difference value at the position (x, y) of each pixel point in the image is calculated, that is, the pixel value change at this position in the (i + 1)-th video frame; by taking the average of the frame difference values and setting the intensity threshold θ, the binary mask M(x, y) is preliminarily generated:

[0027]

[0028] In the formula, x and y respectively represent the numerical values in the x and y axis directions of the video frame, D i+1 (x, y) is the frame difference value at the position (x, y) in the (i + 1)-th video frame, I i (x, y) is the pixel value of the i-th video frame, and I i+1 (x, y) is the pixel value of the (i + 1)-th video frame;

[0029] By analyzing the motion direction and speed of each pixel point in the video frame, the optical flow information of each pixel point is obtained. Specifically: Using the optical flow constraint equation: I x (x, y, t)u(x, y, t)+I y (x, y, t)v(x, y, t)+I t (x, y, t) = 0, the optical flow vectors u(x, y, t) and v(x, y, t) of each pixel point are solved, where u(x, y, t) represents the velocity component of the optical flow in the horizontal direction, v(x, y, t) represents the velocity component of the optical flow in the vertical direction, I x (x, y, t) and I y (x, y, t) are the gradients of the video frame in the x and y directions respectively, representing the change in image intensity; I t (x, y, t) is the change in the video frame intensity over time, representing the intensity difference between consecutive frames;

[0030] Furthermore, the magnitude L of the optical flow vector is calculated as:

[0031]

[0032] The magnitude L of the optical flow vector represents the displacement magnitude of a pixel within a unit of time;

[0033] To combine pixel point information and optical flow information in order to more precisely obtain the specific range of the moving region, based on the obtained binary mask M(x, y) and the magnitude L of the optical flow vector, calculate the optical flow mask M of the key frame Flow (x, y), and the specific formula is as follows:

[0034]

[0035] In the formula, is the initial threshold;

[0036] Perform connected region analysis on the optical flow mask M Flow (x, y) and calculate the coordinates of the minimum bounding rectangle of this region. Take the coordinates of the minimum bounding rectangle as the ROI (region of interest) of the improved GroundedSAM model.

[0037] Furthermore, in step 3), input the video frame F extracted in step 2) t into the RAM model, and output the set of object and action labels corresponding to the frame, where L i is defined as the preliminary pseudo-label set;

[0038] To optimize the semantic accuracy of the preliminary pseudo-label set, perform semantic matching between it and the extended label set in the action association library; the semantic similarity is calculated by calculating the embedding vector u i of the preliminary pseudo-label set L i and the label vector v j in the action association library, evaluate the semantic correlation between the two, and the similarity formula is:

[0039]

[0040] In the formula, S ij is the cosine similarity between the two;

[0041] Use the RAG retrieval technique to retrieve the k extended labels with the highest similarity from the action association library to generate the candidate pseudo-label set W i :

[0042] W i = Top-k({S ij})

[0043] Select the label with the most relevant semantics from the candidate pseudo-label set W i as the final pseudo-label set P i :

[0044] Pi = Select(W i )

[0045] The final pseudo-label set P i is semantically consistent with the preliminary pseudo-label set L i and optimizes the accuracy of the labels.

[0046] Furthermore, in step 4), the ROI region of interest and video frames generated in step 2) are used as the image input of the improved GroundedSAM model to generate a segmentation mask. Subsequently, the GroundedSAM model fuses the generated segmentation mask with the pseudo-labels generated in step 3) for text features and image features, and performs object detection based on the fused features to generate the object label set A = {a1, a2,..., a n'}, where a n' is the n'-th object label.

[0047] Furthermore, in step 5), the LLM model and the designed Prompt template are used to achieve the conversion from object labels to natural language generation. The specific steps are as follows:

[0048] Prompt template design: Design the Prompt template O i that adapts to different scenarios and behaviors to ensure that the object label set A can be effectively converted into natural language descriptions;

[0049] Using the pre-trained large language model LLM, according to the designed Prompt template O i and the object label set A, generate the natural language text S = {s1, s2,..., s i' ,..., s m} of the human behavior in the video, satisfying:

[0050] s i' = LLM(O i' , T i' )

[0051] where s i' represents the i'-th natural language text, O i' represents the i'-th Prompt template, and T i' represents the i'-th behavior label.

[0052] Furthermore, in step 6), it is implemented through the following specific steps:

[0053] Dual-channel input processing: Input the natural language text S and the behavior label T generated in step 5) into the two channels of the dual-channel classifier respectively for independent feature extraction to obtain the feature vectors f1 and f2;

[0054] Feature vector comparison: The extracted feature vectors f1 and f2 calculate the similarity Similarity to evaluate the similarity and consistency between the text and the true label. The similarity calculation formula is as follows:

[0055]

[0056] Confidence score generation: According to the comparison results of the feature vectors, a confidence score for each behavior recognition result is generated to reflect the reliability of the prediction. The confidence score C′ is calculated by the following formula:

[0057] C′ = (Similarity)

[0058] where σ is the Sigmoid function;

[0059] Behavior category prediction: According to the confidence score and similarity results, the natural language text S is classified into the closest behavior category. The classification rules are as follows:

[0060]

[0061] where C i* represents the confidence score of the i*-th behavior category, and predictedBehavior represents the closest behavior category obtained by classification.

[0062] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0063] 1. By introducing the pre-trained language model BERT, the present invention semantically expands the behavior labels and constructs an action association library in combination with the semantic embedding model, significantly improving the semantic coverage of the labels. Combining the semantic matching between the preliminary pseudo-labels generated by the RAM model and the action association library, and using the RAG retrieval mechanism to accurately correct the pseudo-labels, effectively improves the system's semantic understanding ability of diverse and complex human behaviors, and enhances the robustness and recognition accuracy.

[0064] 2. Through pseudo-label generation and label semantic expansion, the present invention effectively reduces the manual annotation burden and significantly saves the data preparation time. The present invention combines optical flow information and connected region analysis to dynamically determine the key motion region as the ROI region of interest of GroundedSAM, enhancing the accuracy of object detection.

[0065] 3. Compared with the traditional method that performs redundant calculations on the entire video, the present invention significantly reduces the processing requirements for irrelevant regions through optical flow and ROI region localization. At the same time, the improved GroundedSAM model and segmentation mask mechanism further compress the computational overhead. The pseudo-label mechanism and natural language generation significantly reduce the dependence on large-scale manually annotated data. Finally, in the classification stage, by combining similarity and confidence scores, the computational power cost during training and inference is reduced while ensuring the accuracy of the results. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a flowchart of the method of the present invention.

[0067] Figure 2 Generate the final pseudo-label view based on the action association library and the segmentation mask of the GroundedSAM model.

[0068] Figure 3 It is a flowchart for optimizing the ROI (Region of Interest).

[0069] Figure 4 It is a flowchart for dual-channel feature matching and evaluation. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.

[0071] As Figures 1 to 4 shown, this embodiment discloses a method for human behavior recognition based on deep learning. This method uses an improved GroundedSAM model to achieve accurate recognition of human behaviors. The improved GroundedSAM model improves the pseudo-label generation module and the motion feature extraction module of the original GroundedSAM model. The improvement of the pseudo-label generation module is: on the basis of automatically generating pseudo-labels, semantic expansion is performed on the behavior labels in the labeled dataset to construct an action association library. The improvement of the motion feature extraction module is: a binary mask is initially generated by the frame difference method, and then an optical flow mask is obtained by calculating the optical flow information of the video frames to adjust the ROI (Region of Interest) of the model.

[0072] The specific implementation of this human behavior recognition method includes the following steps:

[0073] 1) Through the pre-trained language model BERT, semantic expansion is performed on the behavior labels in the labeled human behavior dataset, an extended label set is generated in combination with the context, and the labels in the extended label set are mapped into high-dimensional vectors by using a semantic embedding model to construct an action association library, specifically as follows:

[0074] For the behavior labels in the labeled human behavior dataset, semantic expansion is performed using the pre-trained language model BERT. During the expansion process, the large language model generates an extended label set E related to the input label T based on context information. The expansion formula is expressed as follows:

[0075] E = B(T, C)

[0076] In the formula, B represents the language model BERT, T is the input behavior label; C is the context information used to guide the generation of relevance; E is the generated extended label set {E1, E2,..., E i ,..., E n}}, where E i is the i-th extended label, i = 1, 2, 3..., n;

[0077] The extended label set E is mapped to a high-dimensional vector space through the pre-trained semantic embedding model F to obtain the vector representation of each label:

[0078] v i = F(E i )

[0079] In the formula, v i is the i-th label vector

[0080] All the embedded vectors are stored in the action association library and used as the matching basis for subsequent pseudo-label generation. The action association library is defined as the vector set V:

[0081] V = {v1, v2,..., v n}

[0082] 2) Extract video frames from the video, generate a preliminary mask through pixel-level difference, and combine optical flow information and connected region analysis to determine the bounding box of the motion region as the ROI (Region of Interest) of the improved GroundedSAM model, accurately determining the main motion region of the person in the video frame, as follows:

[0083] By using the video frames extracted from the video, the pixel difference of the video frames, and the change of optical flow information, the ROI region is dynamically adjusted. The specific steps are as follows: From the original video data, continuous video frames F t are extracted according to the frame extraction rule of Q frames per second, where t represents the time index, and the frame extraction rule is expressed as:

[0084]

[0085] In the formula, V' is the input video, Δ is the frame extraction interval, and F t is the extracted video frame;

[0086] For the extracted video frames Ft Perform pixel - level differential processing on adjacent frame images, calculate the frame difference value at the position (x, y) of each pixel point in the image, that is, the pixel value change of the (i + 1)-th video frame at this position; by averaging the frame difference values and setting an intensity threshold θ, initially generate a binary mask M(x, y):

[0087]

[0088] In the formula, x and y respectively represent the numerical values in the x - axis and y - axis directions of the video frame, D i+1 (x, y) is the frame difference value of the (i + 1)-th video frame at the position (x, y), I i (x, y) is the pixel value of the i - th video frame, I i+1 (x, y) is the pixel value of the (i + 1)-th video frame;

[0089] By analyzing the motion direction and speed of each pixel point in the video frame, obtain the optical flow information of each pixel point. Specifically: Use the optical flow constraint equation: I x (x, y, t)u(x, y, t)+I y (x, y, t)v(x, y, t)+I t (x, y, t) = 0, and solve for the optical flow vectors u(x, y, t) and v(x, y, t) of each pixel point. Among them, u(x, y, t) represents the velocity component of the optical flow in the horizontal direction, v(x, y, t) represents the velocity component of the optical flow in the vertical direction, I x (x, y, t) and I y (x, y, t) are the gradients of the video frame in the x - direction and y - direction respectively, representing the change in image intensity; I t (x, y, t) is the change in the video frame intensity over time, representing the intensity difference between consecutive frames;

[0090] Furthermore, calculate the magnitude L of the optical flow vector as:

[0091]

[0092] The magnitude L of the optical flow vector represents the displacement size of the pixel within a unit time;

[0093] In order to combine the pixel point information and the optical flow information to more accurately obtain the specific range of the motion area, based on the obtained binary mask M(x, y) and the magnitude L of the optical flow vector, calculate the optical flow mask M Flow (x, y) of the key frame. The specific formula is as follows:

[0094]

[0095] In the formula, is the initial threshold;

[0096] Perform connected component analysis on the optical flow mask M Flow (x,y), and calculate the coordinates of the minimum bounding rectangle of the region. Use the coordinates of the minimum bounding rectangle as the ROI (region of interest) of the improved GroundedSAM model.

[0097] 3) Utilize the video frames extracted in step 2), generate preliminary pseudo-labels through the RAM model, and combine with the action association library constructed in step 1). By performing semantic matching at the vector level between the preliminary pseudo-labels and the extended labels in the action association library, use the RAG retrieval technology to refine the pseudo-labels, and finally generate a pseudo-label set with higher semantic relevance, specifically as follows:

[0098] Input the video frames F t extracted in step 2) into the RAM model, and output the object and action label sets corresponding to the frames, where L i is defined as the preliminary pseudo-label set;

[0099] To optimize the semantic accuracy of the preliminary pseudo-label set, perform semantic matching between it and the extended label set in the action association library; the semantic similarity is calculated by computing the embedding vector u i of the preliminary pseudo-label set L i and the label vector v j in the action association library, evaluate the semantic relevance between the two, and the similarity formula is:

[0100]

[0101] In the formula, S ij is the cosine similarity between the two;

[0102] Use the RAG retrieval technology to retrieve the k extended labels with the highest similarity from the action association library to generate a candidate pseudo-label set W i :

[0103] W i = Top-k({S ij})

[0104] Select the label with the most relevant semantics from the candidate pseudo-label set W i as the final pseudo-label set P i :

[0105] P i = Select(W i )

[0106] The final pseudo-label set P i maintains semantic consistency with the preliminary pseudo-label set L i and optimizes the accuracy of the labels.

[0107] 4) Take the ROI (Region of Interest) generated in step 2) and the video frame as the image input of the improved GroundedSAM model to generate a segmentation mask. Subsequently, the GroundedSAM model fuses the text features and image features of the generated segmentation mask with the pseudo-labels generated in step 3), and performs object detection based on the fused features to generate a set of object labels \(A = \{a_1, a_2, \ldots, a\) n' \}, where \(a\) n' is the \(n'\)-th object label.

[0108] 5) Based on the set of object labels generated in step 4), use the pre-trained large language model LLM and refined Prompt templates to convert the object labels in the set of object labels into natural language descriptions, and then generate accurate and coherent natural language text in combination with the context, as follows:

[0109] Use the LLM model and designed Prompt templates to achieve the conversion from object labels to natural language generation. The specific steps are as follows:

[0110] Prompt template design: Design Prompt templates \(O\) i that adapt to different scenarios and behaviors to ensure that the set of object labels \(A\) can be effectively converted into natural language descriptions;

[0111] Use the pre-trained large language model LLM, according to the designed Prompt template \(O\) i and the set of object labels \(A\), to generate natural language text \(S=\{s_1, s_2, \ldots, s\) i' , \ldots, s\) m \} for the human behaviors in the video, satisfying:

[0112] \(s\) i' = LLM(O\) i' , T\) i' )

[0113] where \(s\) i' represents the \(i'\)-th natural language text, \(O\) i' represents the \(i'\)-th Prompt template, and \(T\) i' represents the \(i'\)-th behavior label.

[0114] 6) In the classification stage, compare the natural language text generated in step 5) with the behavior labels in the pre-annotated human behavior dataset in step 1), calculate the similarity and confidence scores to generate the final classification result, which is achieved through the following specific steps:

[0115] Dual-channel input processing: The natural language text S and the behavior label T generated in step 5) are respectively input into two channels of a dual-channel classifier for independent feature extraction to obtain feature vectors f1 and f2;

[0116] Feature vector comparison: The extracted feature vectors f1 and f2 evaluate the similarity and consistency between the text and the true label by calculating the similarity Similarity. The similarity calculation formula is as follows:

[0117]

[0118] Confidence score generation: According to the comparison results of the feature vectors, a confidence score for each behavior recognition result is generated to reflect the reliability of the prediction. The confidence score C′ is calculated by the following formula:

[0119] C′ = σ(Similarity)

[0120] In the formula, σ is the Sigmoid function;

[0121] Behavior category prediction: According to the confidence score and the similarity result, the natural language text S is classified into the closest behavior category. The classification rule is as follows:

[0122]

[0123] In the formula, C i* represents the confidence score of the i*-th behavior category, and predictedBehavior represents the closest behavior category obtained by classification.

[0124] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for human behavior recognition based on deep learning, characterized in that, This method uses an improved GroundedSAM model to achieve accurate identification of human behaviors. The improved GroundedSAM model improves the pseudo-label generation module and the motion feature extraction module of the original GroundedSAM model. The improvement of the pseudo-label generation module is as follows: on the basis of automatically generating pseudo-labels originally, a motion association library is constructed by semantically expanding the behavior labels in the annotated dataset. The improvement of the motion feature extraction module is as follows: a binary mask is initially generated by the frame difference method, and then an optical flow mask is obtained by calculating the optical flow information of the video frames, and the ROI (Region of Interest) of the model is adjusted. The specific implementation of this human behavior recognition method includes the following steps: 1) Through the pre-trained language model BERT, semantically expand the behavior labels in the annotated human behavior dataset, generate an extended label set in combination with the context, and use the semantic embedding model to map the labels in the extended label set into high-dimensional vectors to construct a motion association library; 2) Extract video frames from the video, generate a preliminary mask through pixel-level difference, and combine the optical flow information and connected region analysis to determine the bounding box of the motion region as the ROI of the improved GroundedSAM model, accurately determining the main motion region of the person in the video frame; 3) Use the video frames extracted in step 2) and generate preliminary pseudo-labels through the RAM model. Combine the motion association library constructed in step 1). By performing semantic matching at the vector level between the preliminary pseudo-labels and the extended labels in the motion association library, use the RAG retrieval technology to refine the pseudo-labels, and finally generate a pseudo-label set with higher semantic relevance; 4) Input the ROI and video frames obtained in step 2) into the trained improved GroundedSAM model to generate a segmentation mask, and perform semantic matching with the pseudo-labels in the pseudo-label set to generate a target label set for the video frame; 5) Based on the target label set generated in step 4), use the pre-trained large language model LLM and refined Prompt templates to convert the target labels in the target label set into natural language descriptions, and then generate accurate and coherent natural language texts in combination with the context; 6) In the classification stage, compare the natural language text generated in step 5) with the behavior labels in the annotated human behavior dataset in step 1), and calculate the similarity and confidence score to generate the final classification result.

2. The method for identifying human behaviors based on deep learning according to claim 1, wherein, In step 1), for the behavior labels in the annotated human behavior dataset, use the pre-trained language model BERT for semantic expansion. The expansion process generates an extended label set E related to the input label T based on the context information by the large language model. The expansion formula is expressed as follows: E = B(T, C) Wherein, B represents the language model BERT, T is the input behavior label; C is the context information used to guide the relevance of generation; E is the generated set of extended labels {E1, E2,..., E i ,..., E n}, where E i is the i-th extended label, i = 1, 2, 3..., n; Map the extended label set E into the high-dimensional vector space through the pre-trained semantic embedding model F to obtain the vector representation of each label: v i = F(E i ) where v i is the i-th label vector All the embedded vectors are stored in the motion association library and used as the matching basis for subsequent pseudo-label generation. The motion association library is defined as the vector set V: V = {v1, v2,..., v n}.

3. The method for identifying human behaviors based on deep learning according to claim 2, wherein: In step 2), the ROI region is dynamically adjusted by using the pixel difference of the video frames extracted from the video and combining the change of the optical flow information. The specific steps are as follows: From the original video data, continuous video frames F are extracted according to the frame extraction rule of Q frames per second, where t represents the time index, and the frame extraction rule is expressed as: t , Wherein, V' is the input video, Δ is the frame extraction interval, and F t is the extracted video frame; For the extracted video frame F t , perform pixel-level differential processing on adjacent frame images, calculate the frame difference at the position (x, y) of each pixel in the image, that is, the pixel value change of the (i + 1)-th video frame at this position; by taking the average of the frame differences and setting the intensity threshold θ, initially generate a binary mask M(x, y) for the differential result: Where x and y respectively represent the numerical values in the x and y axis directions of the video frame, D i+1 (x, y) is the frame difference at the position (x, y) of the (i + 1)-th video frame, I i (x, y) is the pixel value of the i-th video frame, I i+1 (x, y) is the pixel value of the (i + 1)-th video frame; By analyzing the motion direction and speed of each pixel in the video frame, the optical flow information of each pixel is obtained. Specifically: Using the optical flow constraint equation: I x (x,y,t)u(x,y,t)+I y (x,y,t)v(x,y,t)+I t (x,y,t) = 0, the optical flow vectors u(x,y,t) and v(x,y,t) of each pixel are solved, where u(x,y,t) represents the velocity component of the optical flow in the horizontal direction, and v(x,y,t) represents the velocity component of the optical flow in the vertical direction. I x (x,y,t) and I y (x,y,t) are the gradients of the video frame in the x and y directions respectively, representing the change in image intensity; I t (x,y,t) is the change in the intensity of the video frame over time, representing the intensity difference between consecutive frames; Furthermore, calculate the optical flow vector magnitude L as: The magnitude L of the optical flow vector represents the displacement magnitude of the pixel within a unit time; In order to combine pixel point information and optical flow information to more accurately obtain the specific range of the motion area, based on the obtained binary mask M(x, y) and the magnitude of the optical flow vector L, calculate the optical flow mask M Flow (x, y) of the key frame. The specific formula is as follows: In the formula, is the initial threshold; Perform connected component analysis on the optical flow mask M Flow (x, y) to calculate the coordinates of the minimum bounding rectangle of the region. Use the coordinates of the minimum bounding rectangle as the ROI (region of interest) of the improved GroundedSAM model.

4. A method for recognizing human behaviors based on deep learning according to claim 3, characterized in that: In step 3), the video frame F extracted in step 2) t is input into the RAM model to output the set of object and action labels for the corresponding frame, where L i is defined as the preliminary pseudo-label set; To optimize the semantic accuracy of the initial pseudo-label set, it is semantically matched with the extended label set in the action association library; the semantic similarity calculation is performed by calculating the embedding vector u i of the initial pseudo-label set L i and the label vector v in the action association library j to evaluate their semantic relevance. The similarity formula is as follows: where S ij is the cosine similarity between the two; Retrieve the top k extended tags with the highest similarity from the action association library using the RAG retrieval technology to generate a candidate pseudo-tag set W i : W i = Top-k({S ij [[ID=4}]}) Select the label with the most relevant semantics from the candidate pseudo-label set W i as the final pseudo-label set P i : P i = Select(W i ) The final pseudo-label set P i is semantically consistent with the preliminary pseudo-label set L i and optimizes the accuracy of the labels.

5. The method for recognizing human behaviors based on deep learning according to claim 4, characterized in that: In step 4), the ROI (Region of Interest) generated in step 2) and the video frame are used as the image input of the improved GroundedSAM model to generate a segmentation mask. Subsequently, the GroundedSAM model fuses the text features and image features of the generated segmentation mask and the pseudo-labels generated in step 3), and performs object detection based on the fused features to generate a set of object labels A = {a1, a2,..., a n'}, where a n' is the n'-th object label.

6. The method for identifying human behaviors based on deep learning according to claim 5, characterized in that: In step 5), the LLM model and the designed Prompt template are used to achieve the conversion from the target label to natural language generation. The specific steps are as follows: Prompt Template Design: Design Prompt templates that adapt to different scenarios and behaviors O i , ensuring that the target label set A can be effectively converted into natural language descriptions; Using the pre-trained large language model LLM, according to the designed Prompt template O i and the target label set A, generate the natural language text S = {s1, s2,..., s i' ,..., s m} of the human behavior in the video, satisfying: s i' = LLM(O i' , T i' ) where s i' represents the i'-th natural language text, O i' represents the i'-th Prompt template, and T i' represents the i'-th behavior label.

7. The method for identifying human behaviors based on deep learning according to claim 6, characterized in that: In step 6), it is achieved through the following specific steps: Dual-channel input processing: The natural language text S and the behavior label T generated in step 5) are respectively input into the two channels of the dual-channel classifier for independent feature extraction to obtain the feature vectors f1 and f2; Feature vector comparison: The extracted feature vectors f1 and f2 evaluate the similarity and consistency between the text and the true label by calculating the similarity Similarity. The similarity calculation formula is as follows: Confidence score generation: According to the comparison results of the feature vectors, a confidence score for each behavior recognition result is generated to reflect the reliability of the prediction. The confidence score C′ is calculated by the following formula: C′ = σ(Similarity) In the formula, σ is the Sigmoid function; Behavior category prediction: According to the confidence score and the similarity result, the natural language text S is classified into the closest behavior category. The classification rule is as follows: where C i* represents the confidence score of the i*-th behavior category, and predictedBehavior represents the closest behavior category obtained by classification.

Citation Information

Patent Citations

  • Text-to-image generation method based on fine-grained semantic reward

    CN116883530A

  • Model training method and device, image recognition method and device and readable storage medium

    CN117115585A

  • Character action recognition analysis method and system based on infrared laser and deep learning

    CN118747911A

  • Key video data extraction method based on multi-dimensional semantic information

    WO2024109308A1