Multi-modal behavior recognition method and device based on semantic augmentation and electronic equipment

By acquiring a preset sample set, generating video content text information and key frame image groups, and fine-tuning the behavior recognition model, the accuracy of deep neural networks in video behavior recognition is improved, the problem of background information interference is solved, and more accurate behavior recognition is achieved.

CN122116498APending Publication Date: 2026-05-29HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2026-04-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

When deep neural networks are used for video behavior recognition, they pay attention not only to foreground behavior but also to background information, resulting in poor recognition performance.

Method used

By employing a semantically augmented multimodal behavior recognition method, a preset sample set is obtained, video content text information is generated, key frame image groups are extracted, surveillance video to be recognized is captured by a surveillance camera, key frame image groups to be recognized are generated, and user-inputted behavioral semantic query information is received and input into the behavior recognition model for behavior recognition and solution processing, thereby improving the accuracy of behavior recognition.

Benefits of technology

By employing a semantically augmented multimodal behavior recognition method, the accuracy of behavior recognition is improved, interference from background information is reduced, and attention is enhanced to foreground behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116498A_ABST
    Figure CN122116498A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a multi-modal behavior recognition method and device based on semantic augmentation and an electronic device. A specific implementation of the method includes: obtaining a preset sample set; performing semantic augmentation processing on each label information in the preset sample set based on each person behavior video in the preset sample set to generate video content text information; determining each video content text information as a video content text information set; extracting at least one key frame image in the person behavior video to obtain a key frame image group; determining each key frame image group as a key frame image group set; fine-tuning a preset behavior recognition model to obtain a behavior recognition model; collecting a to-be-recognized monitoring video; generating a to-be-recognized monitoring key frame image group; receiving behavior semantic query information; inputting the to-be-recognized monitoring key frame image group and the behavior semantic query information into the behavior recognition model to obtain behavior semantic information. The implementation improves the effect of multi-modal behavior recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically to a method, apparatus, and electronic device for multimodal behavior recognition based on semantic augmentation. Background Technology

[0002] Multimodal behavior recognition based on semantic augmentation is a technique for identifying behaviors in videos. Currently, the common approach to identifying behaviors in videos is to automatically learn video features using deep neural networks.

[0003] However, when using the above methods to identify behaviors in videos, the following technical problems often arise: When using deep neural networks to automatically learn video features for behavior recognition, the sheer volume of background information in the video means that the deep learning model's attention is drawn not only to the foreground behavior but also to the background information when recognizing a particular behavior. For example, when recognizing the behavior of "swimming," if the training data has a swimming pool as the background, the model might misidentify the behavior when the background is a lake. This interference from background information in the video leads to poor performance in behavior recognition.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure propose a method, apparatus, electronic device, and computer-readable medium for multimodal behavior recognition based on semantic augmentation to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a multimodal behavior recognition method based on semantic augmentation. The method includes: acquiring a preset sample set, wherein each sample in the preset sample set includes a video of human behavior and tag information corresponding to the video of human behavior; performing semantic augmentation processing on each tag information included in the preset sample set based on each video of human behavior to generate video content text information; determining the generated video content text information as a set of video content text information; for each video of human behavior, extracting at least one keyframe image from the video of human behavior to obtain a set of keyframe images; and determining the obtained set of keyframe images as a key... A set of frame images; based on the above video content text information set and the above keyframe image set, the preset behavior recognition model is fine-tuned to obtain a behavior recognition model, wherein the behavior recognition model stores an alignment feature information set created during the fine-tuning process; a monitoring video to be recognized is acquired through a monitoring camera; based on the above monitoring video to be recognized, a set of monitoring keyframe images to be recognized is generated; behavioral semantic query information input by the user is received; the above monitoring keyframe image set to be recognized and the above behavioral semantic query information are input into the behavior recognition model, so that the behavior recognition model can perform behavior recognition and solution processing on the above monitoring keyframe image set to be recognized and the above behavioral semantic query information based on the above alignment feature information set to obtain behavioral semantic information.

[0008] Secondly, some embodiments of this disclosure provide a multimodal behavior recognition device based on semantic augmentation. The device includes: an acquisition unit configured to acquire a preset sample set, wherein each sample in the preset sample set includes a video of human behavior and tag information corresponding to the video of human behavior; a processing unit configured to perform semantic augmentation processing on each tag information included in the preset sample set based on each video of human behavior included in the preset sample set, so as to generate video content text information; a first determining unit configured to determine the generated video content text information as a set of video content text information; an extraction unit configured to extract at least one keyframe image from each video of human behavior to obtain a set of keyframe images; and a second determining unit configured to determine the obtained set of keyframe images as a set of keyframe images. The system comprises: a set of keyframe images; a fine-tuning unit configured to fine-tune a preset behavior recognition model based on the video content text information set and the set of keyframe images to obtain a behavior recognition model, wherein the behavior recognition model stores an alignment feature information set created during the fine-tuning process; an acquisition unit configured to acquire the surveillance video to be recognized through a surveillance camera; a generation unit configured to generate a set of keyframe images to be recognized based on the surveillance video to be recognized; a receiving unit configured to receive behavior semantic query information input by a user; and an input unit configured to input the set of keyframe images to be recognized and the behavior semantic query information into the behavior recognition model, so that the behavior recognition model can perform behavior recognition and solution processing on the set of keyframe images to be recognized and the behavior semantic query information based on the alignment feature information set to obtain behavior semantic information.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above embodiments of this disclosure have the following beneficial effects: the semantically augmented multimodal behavior recognition method of some embodiments of this disclosure improves the effect of multimodal behavior recognition. Specifically, the reason for the poor effect of multimodal behavior recognition is that when using deep neural networks to automatically learn video features for behavior recognition, due to the large amount of background information in the video, the deep learning model will pay attention not only to the foreground behavior but also to the background information when recognizing a certain behavior. For example, when recognizing the behavior of "swimming", if the background of the training data is a swimming pool, the model may make a mistake when the background is a lake. Due to the interference of background information in the video, the effect of behavior recognition in the video is poor. Based on this, the semantically augmented multimodal behavior recognition method of some embodiments of this disclosure first obtains a preset sample set, wherein each sample in the preset sample set includes a video of a person's behavior and label information corresponding to the video of the person's behavior. Thus, each video of a person's behavior and the label information corresponding to each video of a person's behavior can be obtained for generating a set of video content text information and a set of keyframe image groups. Then, based on the various person behavior videos included in the aforementioned preset sample set, semantic augmentation processing is performed on each tag information included in the preset sample set to generate video content text information. Next, the generated video content text information is defined as a video content text information set. Thus, a video content text information set can be obtained to promote cross-modal alignment of the behavior recognition model. Then, for each person behavior video, at least one keyframe image is extracted from the video to obtain a keyframe image group. Then, the obtained keyframe image groups are defined as a keyframe image group set. Thus, a keyframe image group set can be obtained for fine-tuning the preset behavior recognition model. Next, based on the aforementioned video content text information set and the aforementioned keyframe image group set, the preset behavior recognition model is fine-tuned to obtain a behavior recognition model, wherein the behavior recognition model stores an alignment feature information set created during the fine-tuning process. Thus, after fine-tuning the preset behavior recognition model, the behavior recognition model can project the feature vectors of the keyframe image groups and the feature vectors of the video content text onto the same shared feature space, thereby allowing the extracted feature vectors of the keyframe image groups to focus more on foreground behavior in the video and ignore background information. Next, surveillance video to be identified is captured via a surveillance camera. This yields the surveillance video to be identified, which is then used to generate a set of keyframe images for behavior recognition. Based on this video, the set of keyframe images is generated. This provides the set of keyframe images for behavior recognition. Finally, user-inputted behavioral semantic query information is received.Finally, the aforementioned keyframe image group to be identified and the aforementioned behavioral semantic query information are input into the aforementioned behavior recognition model. The behavior recognition model then performs behavior recognition processing on the aforementioned keyframe image group to be identified and the aforementioned behavioral semantic query information based on the aforementioned aligned feature information set, thereby obtaining behavioral semantic information. Thus, the behavioral semantic information in the surveillance video to be identified can be obtained. Furthermore, because the feature vectors of the keyframe images and the corresponding video content text feature vectors are projected into the same feature space during the fine-tuning of the preset behavior recognition model, the feature vectors of the keyframe images and the corresponding video content text feature vectors are made as close as possible to each other. This allows the extracted feature vectors of the keyframe images to focus more on the behavioral features within the keyframe images while ignoring the background features, thereby improving the recognition effect of multimodal behavior recognition. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of the semantically augmented multimodal behavior recognition method according to the present disclosure; Figure 2 This is a schematic diagram of the structure of some embodiments of the multimodal behavior recognition device based on semantic augmentation according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] Figure 1 A flow 100 of some embodiments of a semantically augmented multimodal behavior recognition method according to the present disclosure is shown. This semantically augmented multimodal behavior recognition method includes the following steps: Step 101: Obtain the preset sample set.

[0021] In some embodiments, the executing entity (e.g., a computing device) of the semantically augmented multimodal behavior recognition method can obtain a preset sample set. In practice, the executing entity can first access the official dataset website (https: / / www.kinetics-dataset.org / ) to download the sample set and obtain the preset sample set. Each sample in the preset sample set includes a video of a person's behavior and corresponding label information. The preset sample set can be the Kinetics-400 dataset. The video of the person's behavior can be a video recording a person's actions or activities (such as walking, waving, falling, etc.). The label information can be category labels representing the video of the person's behavior. For example, the label information can be swimming, running, etc.

[0022] Step 102: Based on the videos of various people's behaviors included in the preset sample set, perform semantic augmentation processing on each tag information included in the preset sample set to generate video content text information.

[0023] In some embodiments, the execution entity may perform semantic augmentation processing on each tag information included in the preset sample set based on the videos of various human behaviors included in the preset sample set, so as to generate video content text information.

[0024] In some optional implementations of certain embodiments, the execution entity may perform semantic augmentation processing on each tag information included in the preset sample set based on the individual behavior videos included in the preset sample set, in order to generate video content text information: The first step is to identify the videos of the individuals whose behavior matches the aforementioned tag information as the target individuals' behavior videos.

[0025] The second step involves inputting the aforementioned tag information and the target person's behavior video into various preset video content generation text models to obtain initial video content text information. These preset video content generation text models can be models such as Deepseek, Qwen, or GPT. The initial video content text information can be unfiltered text expanded based on the tag information. For example, the aforementioned tag information could be "robotic dance" or "running." The initial video content text information could be "The dancer simulates mechanical movement through stiff, paused movements."

[0026] The third step involves quality screening of the initial video content text information to obtain the final video content text information. In practice, firstly, the executing entity can clean the initial video content text information to obtain cleaned video content text information. Then, based on preset rules, a large model is used to screen the cleaned video content text information to obtain the final video content text information. This video content text information can be a description of the actions of characters in the video. It can include both explicit and implicit text information. Explicit text information can contain words from the tag information and describe the actions of characters in the video. Implicit text information can not contain words from the tag information and describe the actions of characters in the video. For example, the tag information could be "robot dance," and the explicit text information could be "a dance style that imitates robot movements, where dancers use stiff, mechanical movements and sudden pauses to simulate robot motion. This style is common in street dance, emphasizing body control and rhythm, and is usually performed to the beat of music." The aforementioned implicit text information could be "the performer simulating the agent's movements through stiff, rigid body movements and sudden pauses." The aforementioned large model could be the DeepSeek large model. The aforementioned preset rules include preset rules for generating display text information and preset rules for generating implicit text information. The aforementioned preset rules for generating display text information could be that the video content text information contains words from the tag information. The aforementioned preset rules for generating implicit text information could be that the video content text information does not contain words from the tag information.

[0027] Step 103: Determine the generated video content text information as a video content text information set.

[0028] In some embodiments, the executing entity may determine the generated video content text information as a video content text information set. Each piece of video content text information in the set corresponds to one tag information from the various tag information sets.

[0029] Step 104: For each video of a person's behavior, extract at least one keyframe image from the video of the person's behavior to obtain a keyframe image group.

[0030] In some embodiments, the execution entity may extract at least one keyframe image from each of the individual behavior videos to obtain a keyframe image group.

[0031] In some optional implementations of certain embodiments, the aforementioned execution entity may extract at least one keyframe image from each of the individual behavior videos to obtain a keyframe image group by performing the following steps: The first step is to generate a sequence of video frames depicting human actions based on the aforementioned video footage. In practice, the executing entity can first use video processing tools such as FFmpeg combined with automated scripts to extract the frame sequence from the video footage. Then, the obtained frame sequence is identified as the human action video frame sequence. This sequence can be a sequence of images containing human actions, arranged chronologically. Each video frame in the sequence has a sequential index. This sequential index represents the direct mapping relationship between the index and the video frames.

[0032] The second step is to determine the number of video frames representing human behavior in the above sequence of video frames as the total number of video frames representing human behavior.

[0033] The third step is to generate at least one random index based on the total number of frames in the video of the aforementioned character's behavior. In practice, the executing entity can generate random index information using a random number generation function (e.g., the numpy.random.randint() function). This random index information can be a randomly generated index value.

[0034] The fourth step involves extracting at least one keyframe image from the video frame sequence of the person's behavior based on the aforementioned at least one random index information, thus obtaining a keyframe image group. In practice, for each of the at least one random index information, the executing entity first uses the random index information to directly locate the video frame corresponding to the random index information in the aforementioned video frame sequence of the person's behavior. Finally, the video frame corresponding to the random index information is read. The resulting video frames are then identified as a keyframe image group. The keyframe image can be a single static image.

[0035] Step 105: Determine the obtained keyframe image groups as a keyframe image group set.

[0036] In some embodiments, the execution entity may determine the obtained keyframe image groups as a keyframe image group set.

[0037] Step 106: Based on the video content text information set and the key frame image set, fine-tune the preset behavior recognition model to obtain the behavior recognition model.

[0038] In some embodiments, the execution entity may fine-tune a preset behavior recognition model based on the video content text information set and the keyframe image set to obtain a behavior recognition model. The preset behavior recognition model includes a pre-trained image mapping module, a pre-trained text mapping module, a preset image feature module, a preset text feature module, and a text decoder. The text decoder may be a Transformer decoder.

[0039] In some optional implementations of certain embodiments, the aforementioned execution entity may fine-tune the preset behavior recognition model based on the aforementioned video content text information set and the aforementioned keyframe image set through the following steps to obtain the behavior recognition model: The first step involves inputting the aforementioned set of keyframe images into a pre-trained image mapping module included in the preset behavior recognition model to obtain an image feature information set. This pre-trained image mapping module can be a Convolutional Neural Network (CNN). Each image feature in the image feature information set can be a numerical feature vector representing the visual essence of the keyframe image set (such as shape, texture, contour, etc.). Each image feature in the image feature information set corresponds to one keyframe image set in the aforementioned keyframe image set. As an example, at least one sample keyframe image set and the sample image feature information set corresponding to each keyframe image set in the at least one sample keyframe image set can be obtained first. Then, using each keyframe image set in the at least one sample keyframe image set as input and the sample image feature information set corresponding to each keyframe image set in the at least one sample keyframe image set as the desired output, a trained image mapping module is obtained.

[0040] The second step involves masking the aforementioned video content text information set to obtain a masked video content text information set. In practice, firstly, for each piece of video content text information in the set, the aforementioned execution entity can segment the video content text information using a word segmenter (e.g., word segmentation based on a dictionary or statistical model) to obtain segmented video content text information. Then, part-of-speech tagging is performed on each word in the segmented video content text information to determine relevance. Finally, untagged words (Tokens) in the segmented video content text information are replaced with [MASK] for masking. As an example, the aforementioned execution entity can tag verbs as VB, action nouns as NN, and related adverbs as RB* for part-of-speech tagging. For example, the untagged words can be words not directly related to the action recognition task. These words not directly related to the action recognition task can be this, that, of, the, a, an, etc. The above video content text information could be "The man wearing a black hat is running quickly down the stairs on the east side, then he suddenly stops at the corner and looks around." The subtitled video content text information could be "[[MASK]", "wearing", "black", "hat", "[MASK]", "man", "[MASK]", "quickly", "[MASK]", "[MASK]", "east", "[MASK]", "stairs", "run", "down", "[MASK]", "[MASK]", "[MASK]", "[MASK]", "suddenly", "[MASK]", "corner", "stop", "[MASK]", "down", "[MASK]", "[MASK]", "look around", "[MASK]", "[MASK]"".

[0041] The third step involves inputting the masked video content text information set into a pre-trained text mapping module included in the preset behavior recognition model to obtain a text feature information set. Each text feature in the text feature information set corresponds to one image feature in the image feature information set. The pre-trained text mapping module can be a Recurrent Neural Network (RNN). Each text feature in the text feature information set can be a numerical feature vector representing the video content text. As an example, at least one sample masked video content text information set and the sample text feature information set corresponding to each masked video content text information set in the at least one sample masked video content text information set can be obtained first. Then, each masked video content text information set in the at least one sample masked video content text information set is used as input, and the sample text feature information set corresponding to each masked video content text information set in the at least one sample masked video content text information set is used as the expected output to train the text mapping module.

[0042] The fourth step involves fine-tuning the preset behavior recognition model based on the aforementioned image feature information set and text feature information set to obtain the behavior recognition model. This behavior recognition model can be a neural network model capable of performing behavior recognition tasks, fine-tuned from the open-source Qwen2-VL-7B multimodal large model (Qwen2-VL-7B large model).

[0043] In some optional implementations of certain embodiments, the aforementioned execution entity may fine-tune the preset behavior recognition model based on the aforementioned image feature information set and the aforementioned text feature information set through the following steps to obtain the behavior recognition model: The first step, based on the above image feature information set and the above text feature information set, is to perform the following fine-tuning steps: The first sub-step involves inputting at least one image feature from the image feature information set into a preset image feature module to obtain at least one image projection feature. Each of the at least one image feature corresponds to one of the at least one image projection feature. The preset image feature module can be an untrained neural network capable of converting image feature information into image projection feature information through an attention mechanism and linear transformation (e.g., the neural network could be a Transformer model). The image projection feature can be a digitally valued feature vector resulting from projecting the image feature information into a new feature space through a linear transformation.

[0044] The second sub-step involves inputting at least one text feature from the text feature information set into a preset text feature module to obtain at least one text projection feature. Each of the at least one text feature corresponds to one of the at least one text projection feature. The preset text feature module can be an untrained neural network capable of converting text feature information into text projection feature information through linear transformation (e.g., the neural network could be a BERT (Bidirectional Encoder Representations from Transformers) model). The text projection feature information can be a digitally valued feature vector projected onto the same feature space as the image projection feature information after linear transformation, thereby aligning the text features of the keyframe image and the video content.

[0045] The third sub-step involves generating alignment loss values ​​for at least one image projection feature and at least one text projection feature based on a preset multimodal alignment loss function, thereby obtaining various alignment loss values. The preset multimodal alignment loss function can be... . It can represent the image projection feature information of the i-th video. It can represent the j-th text projection feature in at least one text projection feature. It can represent the cosine similarity between the i-th image projection feature and the j-th text projection feature, where N is the total number of at least one text projection feature, and k can be the k-th text projection feature among the N text projection features. It can represent the k-th text projection feature among at least one text projection feature. The i-th image projection feature corresponds to the j-th text projection feature. This can be the alignment loss value.

[0046] The fourth sub-step involves comparing each alignment loss value with a preset threshold to obtain comparison information. The preset threshold can be a pre-defined hyperparameter used to determine whether training is complete. The comparison information can be the relationship between the changes in each alignment loss value and the preset threshold.

[0047] The fifth sub-step involves, in response to determining that the comparison information meets a preset condition, defining the fine-tuned preset behavior recognition model as the behavior recognition model, and performing alignment processing on the image feature information set and the text feature information set based on the behavior recognition model to generate an alignment feature information set. The preset condition can be that one of the changes in the alignment loss values ​​is less than a preset threshold.

[0048] The sixth sub-step involves adjusting the parameters of the preset image feature module and the preset text feature module in response to the determination that the comparison information does not meet the preset conditions, and using at least one unused image feature information, at least one text feature information, and the preset behavior recognition model with adjusted parameters to perform the above fine-tuning steps again.

[0049] In some optional embodiments, the execution entity may perform alignment processing on the image feature information set and the text feature information set based on the behavior recognition model through the following steps to generate an aligned feature information set: The first step involves inputting the aforementioned text feature information set into a pre-trained text feature module within the behavior recognition model to generate a target text projection feature information set. This pre-trained text feature module can be a finely tuned neural network capable of converting text feature information into text projection feature information through linear transformation. For example, this neural network could be a BERT (Bidirectional Encoder Representations from Transformers) model. Each target text projection feature in the target text projection feature information set can be a digitally valued feature vector projected onto a new feature space after a linear transformation of the text feature information.

[0050] The second step involves inputting the aforementioned image feature information set into the trained image feature module of the behavior recognition model to generate a target video projection feature information set. Each target video projection feature in this set can be a digitally valued feature vector projected into the same feature space as the target text projection feature information after processing with an attention mechanism and linear transformation. The trained image feature module can be a fine-tuned neural network capable of converting image feature information into image projection feature information through an attention mechanism and linear transformation (e.g., the neural network could be a Transformer model).

[0051] The third step involves aligning the target text projection feature information set and the target video projection feature information set to obtain an aligned feature information set. In practice, the executing entity can project the target text projection feature information set and the target video projection feature information set to the same feature space through a linear transformation to achieve alignment. Each multimodal alignment feature information in the aligned feature information set can be a feature vector obtained by aligning the target text projection feature information and the target video projection feature information in the same feature space.

[0052] Step 107: Collect the surveillance video to be identified using the surveillance camera.

[0053] In some embodiments, the aforementioned executing entity may acquire surveillance video to be identified via a surveillance camera. The surveillance video to be identified may be a real-time video stream acquired by the surveillance camera, or historical video clips pre-stored in a digital video recorder (DVR) or network video recorder (NVR). The surveillance camera includes, but is not limited to, network cameras (IPCs), high-definition PTZ cameras, and infrared night vision cameras.

[0054] Step 108: Generate a group of key frame images of the surveillance video to be identified.

[0055] In some embodiments, the aforementioned executing entity may generate a group of keyframe images of the surveillance video to be identified based on the aforementioned surveillance video to be identified.

[0056] In addressing the technical problems mentioned above, the application scenario involves surveillance cameras installed on outdoor light poles and building facades. These cameras are susceptible to vibrations from wind and passing vehicles when capturing video footage, leading to image shake and blurry video. Furthermore, multi-feature fusion requires aligning different feature sequences. If the video is blurry, the alignment and fusion process with relatively stable audio or text information fails, resulting in unreliable fused feature sequences and consequently, poor accuracy in multimodal behavior recognition. Therefore, this application scenario requires the following characteristics: applicability to behavior recognition from blurry surveillance video.

[0057] In some alternative embodiments, the execution entity may generate a group of keyframe images of the surveillance video to be identified based on the following steps: The first step is to generate a sequence of video frames to be evaluated based on the aforementioned surveillance video to be identified. In practice, the executing entity can use video processing tools such as FFmpeg combined with automated scripts to extract the video frame sequence from the surveillance video to be identified, thus obtaining the video frame sequence to be evaluated. The video frame sequence to be evaluated can be an image sequence of the surveillance video frames to be identified arranged in chronological order.

[0058] The second step involves performing quality assessment on the video frame sequence to be evaluated, obtaining quality assessment values. In practice, firstly, the execution entity can input the video frame sequence to be evaluated into a visual encoding layer to obtain a high-dimensional spatiotemporal feature map. Next, the high-dimensional spatiotemporal feature map is input into a cross-modal alignment module to obtain a token sequence. Then, the token sequence and preset prompts are concatenated to obtain an input sequence. This input sequence is then input into a pre-trained large language model to obtain a text sequence containing quality assessment descriptions. Finally, the text sequence containing quality assessment descriptions is converted into numerical quality assessment values ​​through a regression head or mapping layer. The visual encoding layer can be a Residual Network (ResNet). Specifically, the aforementioned visual coding layer can perform block embedding processing on each frame of the video frame sequence to extract spatial image features. Simultaneously, to capture the temporal features of the video frames (such as jitter and motion blur), the visual coding layer also includes a temporal attention module to calculate optical flow features or temporal difference features between consecutive frames, fusing spatial and temporal features to obtain a high-dimensional spatiotemporal feature map. The aforementioned cross-modal alignment module can be a multilayer perceptron (MLP) projection layer. This module maps the high-dimensional spatiotemporal feature map to the text semantic space of a pre-trained language model, enabling the video data features to be understood by the large language model to generate visual tokens. The large language model can be a base model based on LLaMA or Qwen architecture. The aforementioned quality assessment value can be a numerical value characterizing whether the surveillance video to be identified is blurry; a higher value indicates a clearer surveillance video. The aforementioned preset prompt word can be a predefined natural language text used to guide the model to perform a specific task or generate a specific type of output. For example, the preset prompt could be, "Please analyze the following video frame feature sequence and output the behavior categories it contains." As an example, at least one sample input sequence and a sample text sequence containing quality assessment descriptions corresponding to each of the at least one sample input sequence can be obtained. Then, using each of the at least one sample input sequence as input and the sample text sequence containing quality assessment descriptions corresponding to each of the at least one sample input sequence as the expected output, a pre-trained large language model is trained.

[0059] Third, in response to determining that the above quality assessment value is greater than or equal to a preset first threshold, the following steps are performed on the above surveillance video to be identified: The first sub-step involves generating a sequence of surveillance video frames to be identified, based on the aforementioned surveillance video to be identified. In practice, the executing entity can use video processing tools such as FFmpeg combined with automated scripts to extract the frame sequence from the surveillance video to be identified. This sequence of surveillance video frames can be an image sequence arranged chronologically. The preset first threshold can be a pre-defined threshold used to judge the video quality of the surveillance video to be identified.

[0060] The second sub-step involves performing inter-frame differencing on the aforementioned sequence of surveillance video frames to be identified, resulting in an inter-frame motion intensity sequence. This inter-frame motion intensity sequence can be used to characterize the motion intensity changing over time. Furthermore, the inter-frame motion intensity sequence can be a sequence arranged chronologically among the inter-frame motion intensities.

[0061] The third sub-step involves generating a filtered motion representation sequence based on the aforementioned inter-frame motion intensity sequence. In practice, firstly, the executing entity can suppress noise and instantaneous fluctuations in the inter-frame motion intensity sequence through filtering operations, extracting representation signals that more stably and fundamentally reflect the motion trends of the video scene. Then, the obtained representation signals are arranged in chronological order to generate a representation signal sequence. Finally, this representation signal sequence is determined as the filtered motion representation sequence. The filtered motion representation sequence can be a sequence arranged in chronological order. The representation signals can refer to low-frequency or mid-to-low-frequency components extracted from the aforementioned inter-frame motion intensity sequence through filtering operations, reflecting the motion tendency and change patterns of the main target or the global camera in the scene over a longer time scale.

[0062] The fourth sub-step involves performing event boundary detection processing on the filtered motion representation sequence to obtain various motion valley points and content abrupt change points. In practice, the execution entity can perform motion valley point and content abrupt change point detection on the filtered motion representation sequence to generate these points. The content abrupt change point refers to the moment when the visual content or behavioral pattern of the surveillance video to be identified undergoes a sudden and drastic change. The motion valley point refers to the moment when the intensity or amount of movement of a person in the surveillance video to be identified reaches a local minimum within a certain period of time.

[0063] The fifth sub-step involves segmenting the surveillance video to be identified based on the aforementioned motion valley points and content abrupt change points to obtain individual behavioral videos. In practice, the executing entity can use the aforementioned motion valley points as the starting frame of the behavioral video and the next motion valley point or the next content abrupt change point (whichever comes first) as the ending frame of the behavioral video, thereby segmenting individual behavioral videos from the surveillance video to be identified. Each of these behavioral videos contains a complete motion cycle (e.g., from the start to the end of raising a hand).

[0064] The sixth sub-step involves generating a set of keyframe images to be identified based on the aforementioned behavioral videos. In practice, the executing entity can extract at least one keyframe from each of the behavioral videos to obtain the set of keyframe images to be identified. This set of keyframe images can be a group of pictures that clearly shows a complete behavior. For example, the set of keyframe images to be identified could be six pictures that clearly show the complete behavioral chain of "approaching - starting to climb - reaching the top - landing - escaping".

[0065] Fourth, in response to determining that the above quality assessment value is less than a preset first threshold, the following steps are performed on the above-mentioned surveillance video to be identified: The first sub-step involves generating a sequence of feature information for the surveillance video frames to be identified, based on the aforementioned sequence of surveillance video frames to be identified. In practice, the executing entity first inputs each surveillance video frame in the sequence into the Vision Transformer model to generate feature information for the surveillance video frames to be identified. Then, based on the obtained feature information of each surveillance video frame to be identified, a sequence of feature information for the surveillance video frames to be identified is generated. This sequence of feature information for the surveillance video frames to be identified can be a high-dimensional feature vector sequence obtained by extracting features from the original frame sequence of the surveillance video to be identified, arranged in chronological order.

[0066] The second sub-step involves jointly aligning and fusing the aforementioned feature sequence of the surveillance video frames to be identified, resulting in a fused feature sequence of the surveillance video frames to be identified. In practice, firstly, for each feature of the surveillance video frames to be identified in the aforementioned feature sequence, the executing entity can learn the correspondence between the surveillance video frames to be identified in the sequence through a multi-head attention mechanism. Simultaneously, the features of the surveillance video frames to be identified in blurred areas are weighted and fused with the corresponding clear features of the surveillance video frames to be identified in each adjacent surveillance video frame to achieve alignment and fusion, resulting in fused features of the surveillance video frames to be identified. Finally, the obtained fused features of the surveillance video frames to be identified are arranged in chronological order to obtain the fused feature sequence of the surveillance video frames to be identified. This fused feature sequence of the surveillance video frames to be identified can be a high-dimensional feature vector sequence arranged in chronological order, obtained after jointly aligning and fusing the aforementioned feature sequence of the surveillance video frames to be identified. The feature information of each surveillance video frame to be identified in the above-mentioned fused feature sequence can be a feature vector that not only contains its own spatial features, but also deeply integrates high-quality temporal context information (dynamic motion information, multi-view complementary information and temporal dependencies) obtained from the compensation obtained from the adjacent surveillance video frames to be identified.

[0067] The third sub-step involves decoding and reconstructing the fused feature sequence of the surveillance video frames to be identified, resulting in a reconstructed video frame sequence. In practice, firstly, the executing entity can reconstruct the fused feature sequence of the surveillance video frames to be identified into a feature map. Then, through transposed convolution, the feature map is upsampled and refined to its original resolution, outputting clear images. These clear images are then identified as the reconstructed video frames. Finally, the reconstructed video frames are arranged in chronological order to obtain the reconstructed video frame sequence. This reconstructed video frame sequence can be a sequence of clear images arranged in chronological order.

[0068] The fourth sub-step involves generating a set of keyframe images to be identified based on the reconstructed video frame sequence. It should be noted that the method used to generate the set of keyframe images based on the reconstructed video frame sequence is the same as the method used to generate the set of keyframe images based on the video frame sequence itself.

[0069] The above technical solution, combined with step 1010 and related content, serves as an inventive point of this disclosure, addressing the technical problem of "poor accuracy in behavior recognition." Factors contributing to poor behavior recognition accuracy often include: surveillance cameras installed on outdoor lampposts, building facades, etc., are easily affected by vibrations from external forces such as wind and vehicle traffic when collecting video data, leading to image jitter and blurry video. Furthermore, during multi-feature fusion, different feature sequences need to be aligned. If the video is blurry, the joint alignment and fusion process with relatively stable audio or text information will fail, resulting in a less reliable fused feature sequence and consequently, poor accuracy in multimodal behavior recognition. Solving these factors improves behavior recognition accuracy. To achieve this, firstly, a sequence of video frames to be evaluated is generated based on the aforementioned video. This yields a sequence of video frames for quality assessment. Then, the sequence of video frames is subjected to quality assessment processing to obtain a quality assessment value. This yields a quality assessment value used to determine whether the video is blurry. Next, in response to determining that the aforementioned quality assessment value is greater than or equal to a preset first threshold, the following steps are performed on the aforementioned surveillance video to be identified. This determines that the surveillance video to be identified is not a blurry video. Then, based on the aforementioned surveillance video to be identified, a sequence of surveillance video frames to be identified is generated. This yields a sequence of surveillance video frames to be identified used to generate an inter-frame motion intensity sequence. Then, inter-frame differencing is performed on the aforementioned surveillance video frame sequence to obtain an inter-frame motion intensity sequence. This yields an inter-frame motion intensity sequence characterizing the change in motion intensity over time. Then, based on the aforementioned inter-frame motion intensity sequence, a filtered motion representation sequence is generated. This yields a filtered motion representation sequence used to generate each motion valley point and each content abrupt change point. Next, event boundary detection processing is performed on the aforementioned filtered motion representation sequence to obtain each motion valley point and each content abrupt change point. This yields each motion valley point and each content abrupt change point used for segmenting the surveillance video to be identified. Based on the aforementioned motion valley points and each content abrupt change point, the surveillance video to be identified is segmented to obtain individual behavioral videos. Then, based on the aforementioned video behaviors, a set of keyframe images to be identified is generated. This yields a set of keyframe images representing the complete behavioral chain. Next, in response to determining that the quality assessment value is less than a first preset threshold, the following steps are performed on the video to be identified: First, based on the sequence of video frames to be identified, a sequence of feature information for the video frames to be identified is generated. This yields a sequence of feature vectors representing the content of the video to be identified. Then, the sequence of feature information for the video frames to be identified is jointly aligned and fused to obtain a fused feature information sequence for the video frames to be identified.Therefore, the inter-frame correspondence can be learned through a self-attention mechanism. Simultaneously, the attention weights can adaptively select which frames and positions to extract information from, thereby fusing information from multiple frames to compensate for blurriness or missing details in the current frame. Then, the fused feature sequence of the surveillance video frames to be identified is decoded and the frame sequence is reconstructed to obtain the reconstructed video frame sequence. This yields a feature tensor representing a clear video. Finally, based on the reconstructed video frame sequence, a set of keyframe images of the surveillance video to be identified is generated. This results in a clear, high-quality set of keyframe images representing a complete behavioral chain. This is because a quality assessment of the surveillance video to be identified is performed during the generation of the keyframe image set. When the acquired surveillance video to be identified is blurry, the inter-frame correspondence can be learned through a self-attention mechanism. Simultaneously, the attention weights can adaptively select which frames and positions to extract information from, thereby fusing information from multiple frames to compensate for blurriness or missing details in the current frame, resulting in an aligned and fused high-quality feature sequence, thus transforming the blurry surveillance video into a clear video. Using clear video for multimodal behavior recognition avoids the failure of joint alignment and fusion with relatively stable text information due to blurry surveillance video, resulting in poor reliability of the fused feature sequence and consequently poor accuracy in multimodal behavior recognition. This approach improves the accuracy of multimodal behavior recognition.

[0070] Step 109: Receive the behavioral semantic query information input by the user.

[0071] In some embodiments, the aforementioned executing entity may receive behavioral semantic query information input by a user. This behavioral semantic query information may be natural language text querying the behavior of a person in a video. For example, the behavioral semantic query information may be "What behavior is being performed in this video?".

[0072] Step 1010: Input the monitoring keyframe image group to be identified and the behavioral semantic query information into the behavior recognition model, so that the behavior recognition model can perform behavior recognition and solution processing on the monitoring keyframe image group and the behavioral semantic query information based on the alignment feature information set, and obtain the behavioral semantic information.

[0073] In some embodiments, the execution entity may input the group of keyframe images to be identified and the behavioral semantic query information into the behavior recognition model, so that the behavior recognition model can perform behavior recognition processing on the group of keyframe images to be identified and the behavioral semantic query information based on the alignment feature information set, thereby obtaining behavioral semantic information. The behavior recognition model includes a pre-trained image mapping module, a pre-trained text mapping module, a trained image feature module, a trained preset text feature module, and a text decoder. The text decoder may be a Transformer decoder.

[0074] In addressing the technical problems mentioned above, and specifically for the application scenario of abnormal behavior detection in intelligent security monitoring systems, the following technical issues often arise: Traditional behavior recognition often relies solely on visual models, lacking common sense and scene understanding, resulting in poor accuracy in unfamiliar or complex environments. Furthermore, in traditional behavior recognition, when faced with behavior recognition results that contradict prior knowledge of the scene but possess high confidence, the lack of a confidence-grading decision-making mechanism leads to high-confidence but actually erroneous recognition results, resulting in a high false alarm rate (e.g., misclassifying normal behavior as abnormal). Therefore, this application scenario requires the following characteristics: suitable for abnormal behavior detection in high-precision intelligent security monitoring systems.

[0075] In some optional embodiments, the execution entity may input the group of monitoring keyframe images to be identified and the behavioral semantic query information into the behavior recognition model through the following steps, so that the behavior recognition model can perform behavior recognition and solution processing on the group of monitoring keyframe images to be identified and the behavioral semantic query information based on the alignment feature information set, and obtain behavioral semantic information: The first step is to process the scene information of the aforementioned keyframe image group to be identified, thereby obtaining video scene information. In practice, the executing entity can use the MobileNetV2 model to process the scene information of the aforementioned keyframe image group to be identified, thereby obtaining video scene information. This video scene information may include one or more of the following: time, location, and environmental information. The environmental information can be comprehensive information that characterizes the physical state and background elements of the video scene. For example, the environmental information could be "outdoor intersection, light rain, slippery and reflective ground."

[0076] The second step involves generating a behavioral knowledge graph based on a pre-defined knowledge graph. In practice, the aforementioned executing entity can first use pre-defined behavioral category tags (such as drinking, running) as core anchor nodes, adding them to the pre-defined knowledge graph through the following steps: First, in response to the determination that a node with the same name already exists in the pre-defined knowledge graph, this node is directly used as an anchor, and the behavioral category tag is added to the pre-defined knowledge graph. Then, in response to the determination that a node with the same name does not exist in the pre-defined knowledge graph, a new node is created and a preliminary association is established with the relevant concept nodes in the pre-defined knowledge graph. Then, through the application programming interface (API) provided by the pre-defined knowledge graph, using each behavioral anchor node as query input, a one-hop or multi-hop neighbor node query is automatically performed to obtain automatic query results. As an example, for the node "drinking", the API can return nodes directly connected to it (such as "alcohol", "social", "drunk") and connection relationships (such as "cause of", "used for", "has attributes"). The query results are returned in the form of a set of triples. Next, the automatic query results are filtered and supplemented (for example, creating the following links for drinking: used for social purposes, attribute is alcohol content, consequence is intoxication, reaction is feeling relaxed), resulting in a behavioral knowledge graph. The aforementioned pre-defined knowledge graph can be a structured semantic knowledge base, storing concepts and their semantic relationships in the form of "entity-relationship-entity" triples. For example, the aforementioned pre-defined knowledge graph could be an ATOMIC (A Theory of Mind based on Common Sense) knowledge graph.

[0077] The third step involves generating image feature information and behavioral semantic query feature information based on the aforementioned keyframe image group to be identified and the aforementioned behavioral semantic query information. In practice, firstly, the executing entity can input the keyframe image group to be identified into the pre-trained image mapping module included in the behavioral recognition model to generate image feature information. Then, the behavioral semantic query information is input into the pre-trained text mapping module included in the behavioral recognition model to generate behavioral semantic query feature information. The image feature information can be numerical features representing the visual content of the keyframe image to be identified. The behavioral semantic query feature information can be numerical features representing the behavioral semantic query. These numerical features can be numerical vectors.

[0078] The fourth step involves inputting the aforementioned image feature information to be identified into the trained image feature module included in the behavior recognition model. The trained image feature module then projects the image feature information to be identified into a feature space identical to the aligned feature information set, thereby obtaining the projected image feature information. This projected image feature information can be a feature vector that, after a linear transformation, is projected into a shared feature space identical to the aligned feature information set.

[0079] The fifth step involves inputting the aforementioned behavioral semantic query feature information into the trained text feature module of the behavior recognition model. This allows the trained text feature module to project the behavioral semantic query feature information into a feature space identical to the aligned feature information set, thereby obtaining the behavioral semantic query projected feature information. This behavioral semantic query projected feature information can be a feature vector projected onto a shared feature space identical to the aligned feature information set after a linear transformation.

[0080] The sixth step involves performing multimodal fusion processing on the above-mentioned projection feature information of the image to be identified and the above-mentioned behavioral semantic query projection feature information to obtain joint multimodal feature information. In practice, the above-mentioned executing entity can concatenate the above-mentioned projection feature information of the image to be identified and the behavioral semantic query projection feature information to obtain joint multimodal feature information.

[0081] Step 7: Based on the aforementioned alignment feature information set, perform similarity matching processing on the joint multimodal feature information to obtain the best-matching multimodal alignment feature information and behavioral confidence. In practice, firstly, the executing entity can calculate the cosine similarity value between the joint multimodal feature information and each multimodal alignment feature information in the aforementioned alignment feature information set. Then, the multimodal alignment feature information corresponding to the largest cosine similarity value is determined as the best-matching multimodal alignment feature information. Finally, the largest cosine similarity value is determined as the behavioral confidence.

[0082] Step 8: Input the aforementioned best-matching multimodal alignment feature information into the text decoder included in the behavior recognition model to obtain the semantic information of the behavior to be verified. The text decoder can be a Transformer decoder. The semantic information of the behavior to be verified can be natural language text output by the behavior recognition model describing the behavior in the surveillance video to be identified. For example, the semantic information of the behavior to be verified could be "multiple people are running in the video".

[0083] Step 9: Based on the video scene information, perform graph query processing on the behavior knowledge graph to obtain a candidate list of behavior scene adaptations. In practice, the aforementioned executing entity can use the aforementioned video scene information (such as "library") as the query node, and retrieve behavior nodes directly connected to the aforementioned scene node in the aforementioned behavior knowledge graph through specific semantic relationships to obtain a candidate list of behavior scene adaptations. The aforementioned candidate list of behavior scene adaptations can be behavior nodes arranged according to the weight or confidence level of the relationship. When constructing the aforementioned behavior knowledge graph, various semantic relationships are predefined, including HasContext (the contextual relationship between behavior and scene), AtLocation (the relationship between behavior and location), Causes (causal relationships), etc., and initial confidence scores are assigned to some relationships. The aforementioned specific semantic relationships include, but are not limited to: HasContext (having context): indicating that the behavior usually occurs in the context of this scene, for example, a HasContext relationship can be defined between the behavior "reading" and the scene "library"; AtLocation (located): indicating that the behavior is usually located in this location, for example, an AtLocation relationship can be defined between the behavior "fighting" and the scene "street". The aforementioned relationships were predefined and stored as triples of "behavior node-relationship-scene node" or "scene node-relationship-behavior node" during the construction phase of the behavioral knowledge graph. Taking the scene node "library" as an example, a query can return behavior nodes such as "reading," "studying," and "borrowing books," as well as the types of relationships connecting these behavior nodes with scene nodes. For example, behaviors strongly correlated with "library" are reading (weight: 0.909) and learning (weight: 0.891), while running is weakly or negatively correlated. The candidate list for adapting the aforementioned behavior scenarios could be [reading, learning].

[0084] Step 10: Based on the aforementioned candidate list of behavioral scenarios, perform consistency verification on the semantic information of the behavior to be verified to obtain consistency verification information. In practice, the executing entity can traverse the candidate list of behavioral scenarios to find the position of the semantic information of the behavior to be verified within the candidate list, thereby obtaining consistency verification information. The consistency verification information can represent whether the semantic information of the behavior to be verified is in the candidate list of behavioral scenarios.

[0085] Step 11: In response to determining that the above consistency verification information meets the preset verification conditions, the above-mentioned behavioral semantic information to be verified is determined as behavioral semantic information. The preset verification conditions can be that the above-mentioned behavioral semantic information to be verified is in the above-mentioned behavioral scenario adaptation candidate list.

[0086] Step 12: In response to the determination that the above consistency verification information does not meet the preset verification conditions, the following steps are executed: The first sub-step involves, in response to determining that the confidence level of the aforementioned behavior is greater than or equal to a first preset confidence level, identifying the semantic information of the behavior to be verified as behavioral semantic information and generating an early warning message. This early warning message can be an alarm notification used to indicate uncertainty in the recognition result. The first preset confidence level can be a pre-set numerical value. For example, the early warning message could be "High uncertainty, manual review required."

[0087] The second sub-step, in response to determining that the confidence level of the aforementioned behavior is less than a first preset confidence level but greater than or equal to a second preset confidence level, modifies the semantic information of the behavior to be verified based on the aforementioned candidate list of behavior scenarios, thereby obtaining the semantic information of the behavior. In practice, the executing entity can determine the behavior scenario adaptation candidate information that ranks highest in the aforementioned candidate list of behavior scenarios and has a certain degree of matching with visual features as the semantic information of the behavior. The aforementioned second preset confidence level can be a pre-set value. As an example, the semantic information of the behavior to be verified can be "running". The highest-ranking behavior scenario adaptation candidate information in the aforementioned candidate list of behavior scenarios and with a certain degree of matching with visual features can be "walking briskly". The semantic information of the behavior to be verified can be modified from "running" to "walking briskly".

[0088] The third sub-step involves generating an uncertainty warning message in response to determining that the confidence level of the aforementioned behavior is less than a second preset confidence level. This uncertainty warning message can be a notification indicating that the behavior recognition model cannot reliably determine the behavior category in the current video segment. For example, the uncertainty warning message could be "The behavior in the video cannot be determined."

[0089] The above-described technical solution and its related content, as an inventive point of this disclosure, address the technical problem of "poor accuracy in behavior recognition." Factors contributing to poor behavior recognition accuracy often include: traditional behavior recognition often relies solely on visual models, lacking common sense and scene understanding, leading to poor accuracy in unfamiliar or complex environments. Furthermore, in traditional behavior recognition, when faced with behavior recognition results that contradict prior scene knowledge but possess high confidence, the lack of a confidence-grading decision mechanism results in high-confidence but actually erroneous recognition results, leading to a high false alarm rate (e.g., misjudging normal behavior as abnormal). To achieve this technical effect, firstly, scene information acquisition processing is performed on the aforementioned monitoring keyframe image group to be recognized, obtaining video scene information. This allows the introduction of video scene information and the generation of a candidate list for behavior scene adaptation. Then, a behavior knowledge graph is generated based on a preset knowledge graph. This enables context-aware enhancement through the behavior knowledge graph and also provides a behavior knowledge graph for generating a candidate list for behavior scene adaptation. Next, based on the aforementioned keyframe image group to be identified and the aforementioned behavioral semantic query information, image feature information and behavioral semantic query feature information are generated. This yields feature vectors representing the visual features of the keyframe image to be identified and the numerical features of the behavioral semantic query. Then, the image feature information is input into a trained image feature module included in the behavior recognition model, allowing the trained image feature module to project the image feature information onto a shared feature space identical to the aligned feature information set, resulting in projected image feature information. This yields projected image feature information used to generate joint multimodal feature information. Next, the behavioral semantic query feature information is input into a trained text feature module included in the behavior recognition model, allowing the trained text feature module to project the behavioral semantic query feature information onto a shared feature space identical to the aligned feature information set, resulting in projected behavioral semantic query feature information. This yields projected behavioral semantic query feature information used to generate joint multimodal feature information. Finally, multimodal fusion processing is performed on the projected image feature information and the projected behavioral semantic query feature information to obtain joint multimodal feature information. Therefore, joint multimodal feature information in the same feature space as the alignment feature information set can be obtained, so as to generate similarity information between each alignment feature information set in the alignment feature information set and the aforementioned joint multimodal feature information. Then, based on the aforementioned alignment feature information set, similarity matching processing is performed on the aforementioned joint multimodal feature information to obtain the most matching multimodal alignment feature information and the behavior confidence score. Thus, the most matching multimodal alignment feature information used to generate the semantic information of the behavior to be verified and the behavior confidence score used to determine the accuracy of the semantic information of the behavior to be verified can be obtained.Next, the best-matching multimodal alignment feature information is input into the text decoder included in the behavior recognition model to obtain the semantic information of the behavior to be verified. Thus, the semantic information of the behavior to be verified can be obtained. Then, graph query processing is performed on the behavior knowledge graph based on video scene information to obtain a candidate list of behavior scene adaptations. This yields a candidate list of behaviors sorted by common-sense probability for consistency verification of the semantic information of the behavior to be verified, allowing for the fusion of scene semantic information and common sense to verify the semantic information of the behavior to be verified. Next, consistency verification processing is performed on the semantic information of the behavior to be verified based on the candidate list of behavior scene adaptations. This allows for common-sense verification of the semantic information of the behavior to be verified, providing a more refined evaluation of the reliability of the final output. Then, in response to determining that the semantic information of the behavior to be verified is in the candidate list of behavior scene adaptations, the semantic information of the behavior to be verified is identified as behavioral semantic information. Thus, by introducing common sense and contextual scene information, the obtained semantic information of the behavior to be verified can be verified to obtain more accurate behavioral semantic information. Then, in response to determining that the semantic information of the behavior to be verified is not in the candidate list of behavior scene adaptations, the following steps are performed. Next, in response to determining that the confidence level of the aforementioned behavior is greater than or equal to a first preset confidence level, the semantic information of the behavior to be verified is identified as behavioral semantic information and an early warning message is generated. Then, in response to determining that the confidence level of the aforementioned behavior is less than the first preset confidence level but greater than or equal to a second preset confidence level, the semantic information of the behavior to be verified is corrected based on the aforementioned behavioral scenario adaptation candidate list to obtain behavioral semantic information. Finally, in response to determining that the confidence level of the aforementioned behavior is less than the second preset confidence level, an uncertainty early warning message is generated. Therefore, when the semantic information of the behavior to be verified is not in the aforementioned behavioral scenario adaptation candidate list, a confidence-driven hierarchical decision can be introduced. Based on the magnitude of the behavior confidence level, high-confidence and early warning messages are used to retain abnormal behaviors that do not conform to prior knowledge of the scenario; corrective processing is used to optimize behavioral information with medium confidence levels; and uncertainty early warning messages are used to alert relevant personnel to low-confidence behavioral information so that relevant personnel can verify the semantic information of the behavior to be verified, thereby reducing false alarms. Because in the process of generating behavioral semantic information, contextual awareness enhancement is enhanced by introducing video scene information and behavioral knowledge graphs, scene semantic information is integrated when performing behavior recognition, and the obtained behavioral semantic information is verified to improve the accuracy of behavior recognition.

[0090] In some embodiments, the aforementioned execution entity may perform the following steps: In addressing the technical problems mentioned above by adopting technical solutions, the application scenario—security control in large transportation hubs (such as airports and train stations)—often presents the following technical challenges: After recognizing behavior in videos, due to a lack of common-sense reasoning and scene understanding inherent in humans, warnings are triggered as soon as dangerous behavior (e.g., running, crowd gathering) is detected. This leads to a proliferation of false alarms and consequently, a slow response to emergency warning information. Therefore, this application scenario requires the following characteristics: suitability for security control against rampant false alarms.

[0091] In some optional implementations of certain embodiments, the above-mentioned execution entity may perform the following steps: The first step is to perform structured extraction processing on the aforementioned behavioral semantic information to obtain structured behavioral information. In practice, firstly, the executing entity can extract keywords from the aforementioned behavioral semantic information using a pre-set behavioral keyword dictionary to obtain keyword information. Then, based on a pre-set warning behavior category library, the keyword information is matched to obtain behavior type information. The aforementioned keyword information can be "running," "waving," "falling," "multiple people," "one person," "elderly," etc. The pre-set warning behavior category library can be a structured, professional behavior classification system oriented towards security warning scenarios. It can be a dedicated category library reconstructed and expanded based on the general behavior categories output by the basic behavior recognition model (such as the 400 daily behaviors in Kinetics-400) according to the business needs of specific security scenarios, through semantic induction, risk level classification, and domain knowledge injection. The pre-set behavioral keyword dictionary can be a pre-constructed mapping table or set used to identify and extract behavior-related words from natural language text. Each entry in the aforementioned pre-defined behavioral keyword dictionary is a behavioral keyword, such as "running," "fighting," "drinking," "reading," and "climbing," corresponding to various human behaviors or actions that may appear in the surveillance video. The structured behavioral information can include behavioral type and actor. For example, structured behavioral information could be [behavior type: violent behavior, actor: single person]. As an example, the aforementioned behavioral semantic information could be a natural language sentence (e.g., "a person is quietly reading a book in a library"). The aforementioned pre-defined behavioral keyword dictionary could be [running, fighting, drinking, reading, climbing]. Therefore, for the aforementioned behavioral semantic information, the extracted keyword information is ["reading"].

[0092] The second step is to generate video scene information for the surveillance video to be identified, based on the aforementioned video to be identified. In practice, the executing entity can use the MobileNetV2 model to generate video scene information for the surveillance video to be identified. This video scene information may include one or more of the following: time, location, and environmental information.

[0093] The third step involves performing a comprehensive risk assessment on the structured behavioral information based on the aforementioned video scene information to obtain a comprehensive risk value. In practice, the executing entity can input the video scene information into a preset early warning rule engine to generate a comprehensive risk value. This preset early warning rule engine includes a series of "IF-THEN" rules. For example, IF: Behavior type is gathering behavior AND Scenario is a bank entrance AND Subject is multiple people AND Urgency level is high THEN triggers a high-level warning. The comprehensive risk value can be a numerical value representing the urgency of the risk of the structured behavioral information. For example, the video scene information could be [Location: Nursing home, Time: 8 PM]. The structured behavioral information could be [Behavior type: fall, Subject: elderly], and the comprehensive risk value could be 0.9.

[0094] Fourth, in response to determining that the overall risk value is greater than or equal to a preset second threshold, the following steps are performed: The first sub-step involves obtaining the timestamp of the surveillance video to be identified, the camera identifier of the surveillance video to be identified, and the location information of the camera. In practice, the executing entity can first obtain the timestamp of the surveillance video to be identified from the surveillance video itself or from the server system receiving the surveillance video stream. Then, based on the surveillance video to be identified, a query and match can be performed in a preset device information configuration database to find the corresponding camera identifier and camera location information. The preset device information configuration database can be a database storing camera identifiers, location information, etc. The timestamp of the surveillance video to be identified refers to metadata information used to identify the time point of acquisition of the surveillance video to be identified. The camera identifier of the surveillance video to be identified can be a number used to distinguish different surveillance camera devices. The camera location information can be a location description given in natural language to describe the location of the camera in physical space. The preset second threshold can be a pre-set value used for comprehensive risk assessment of the structured behavioral information. For example, the camera identifier can be A-1. The camera location information can be Building A - North Gate. The timestamp of the surveillance video to be identified can be a specific date and time (e.g., 2025-04-20 14:35:22.123).

[0095] The second sub-step involves generating an early warning message based on the timestamp of the surveillance video to be identified, the camera's location information, and the aforementioned behavioral semantic information. In practice, the executing entity can encapsulate the timestamp of the surveillance video to be identified, the camera's location information, and the aforementioned behavioral semantic information into an early warning message. For example, the early warning message could be: "[Emergency Warning] At {8 PM}, {an elderly person fell} at {Nursing Home Building A - North Gate}. Please take immediate action!"

[0096] The third sub-step involves generating a warning information data packet based on the aforementioned warning information, the camera identifier of the surveillance video to be identified, and the surveillance video itself. In practice, the executing entity can encapsulate the aforementioned warning information, the camera identifier of the surveillance video to be identified, and the surveillance video itself into a warning information data packet.

[0097] The fourth sub-step involves sending the aforementioned warning information data packet to a preset terminal device. This preset terminal device can be the mobile phone or computer of a designated person in charge.

[0098] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem of "slow response to emergency warning information." The reasons for this slow response are often as follows: after identifying behavior in a video, due to a lack of common-sense reasoning and scene understanding abilities inherent in humans, warnings are triggered as soon as dangerous behavior (e.g., running, crowds gathering) is detected. This leads to a proliferation of false alarms, resulting in a slow response to emergency warning information. Solving these factors can improve the response speed to emergency warning information. To achieve this technical effect, firstly, the semantic information of the behavior is extracted in a structured manner to obtain structured behavioral information. This allows human-readable, unstructured natural language descriptions to be converted into standardized, structured information that computer programs can accurately understand and process. Then, based on the surveillance video to be identified, video scene information of the surveillance video to be identified is generated. This allows the introduction of video scene information to obtain video scene information of the surveillance video to be identified for comprehensive risk assessment processing. Next, based on the video scene information, the structured behavioral information is subjected to comprehensive risk assessment processing to obtain a comprehensive risk value. Therefore, behavioral semantic information can be verified by combining video scene information to obtain a comprehensive risk value. Then, in response to determining that the comprehensive risk value is greater than or equal to a preset second threshold, the following steps are executed. Next, the timestamp of the surveillance video to be identified, the camera identifier of the surveillance video to be identified, and the camera's location information are obtained. This yields the camera identifier and camera's location information of the surveillance video to be identified, used to generate warning information. Then, based on the timestamp of the surveillance video to be identified, the camera's location information, and the behavioral semantic information, warning information is generated. This yields warning information used to generate a warning information data packet. Next, based on the warning information, the camera identifier of the surveillance video to be identified, and the surveillance video to be identified, a warning information data packet is generated. This yields a more accurate warning information packet. Finally, the warning information data packet is sent to a preset terminal device. This allows for the delivery of accurate and reliable warning information to the preset terminal devices of relevant personnel, enabling them to respond quickly. This is because, when generating early warning information, the system assesses risky behaviors based on the video scene information of the surveillance video to be identified, thereby introducing common-sense reasoning and scene understanding capabilities. This avoids the proliferation of false alarms caused by a lack of common-sense reasoning and scene understanding capabilities, thus improving the response speed to emergency early warning information.

[0099] The above embodiments of this disclosure have the following beneficial effects: the semantically augmented multimodal behavior recognition method of some embodiments of this disclosure improves the effect of multimodal behavior recognition. Specifically, the reason for the poor effect of multimodal behavior recognition is that when using deep neural networks to automatically learn video features for behavior recognition, due to the large amount of background information in the video, the deep learning model will pay attention not only to the foreground behavior but also to the background information when recognizing a certain behavior. For example, when recognizing the behavior of "swimming", if the background of the training data is a swimming pool, the model may make a mistake when the background is a lake. Due to the interference of background information in the video, the effect of behavior recognition in the video is poor. Based on this, the semantically augmented multimodal behavior recognition method of some embodiments of this disclosure first obtains a preset sample set, wherein each sample in the preset sample set includes a video of a person's behavior and label information corresponding to the video of the person's behavior. Thus, each video of a person's behavior and the label information corresponding to each video of a person's behavior can be obtained for generating a set of video content text information and a set of keyframe image groups. Then, based on the various person behavior videos included in the aforementioned preset sample set, semantic augmentation processing is performed on each tag information included in the preset sample set to generate video content text information. Next, the generated video content text information is defined as a video content text information set. Thus, a video content text information set can be obtained to promote cross-modal alignment of the behavior recognition model. Then, for each person behavior video, at least one keyframe image is extracted from the video to obtain a keyframe image group. Then, the obtained keyframe image groups are defined as a keyframe image group set. Thus, a keyframe image group set can be obtained for fine-tuning the preset behavior recognition model. Next, based on the aforementioned video content text information set and the aforementioned keyframe image group set, the preset behavior recognition model is fine-tuned to obtain a behavior recognition model, wherein the behavior recognition model stores an alignment feature information set created during the fine-tuning process. Thus, after fine-tuning the preset behavior recognition model, the behavior recognition model can project the feature vectors of the keyframe image groups and the feature vectors of the video content text onto the same shared feature space, thereby allowing the extracted feature vectors of the keyframe image groups to focus more on foreground behavior in the video and ignore background information. Next, surveillance video to be identified is captured via a surveillance camera. This yields the surveillance video to be identified, which is then used to generate a set of keyframe images for behavior recognition. Based on this video, the set of keyframe images is generated. This provides the set of keyframe images for behavior recognition. Finally, user-inputted behavioral semantic query information is received.Finally, the aforementioned keyframe image group to be identified and the aforementioned behavioral semantic query information are input into the aforementioned behavior recognition model. The behavior recognition model then performs behavior recognition processing on the aforementioned keyframe image group to be identified and the aforementioned behavioral semantic query information based on the aforementioned aligned feature information set, thereby obtaining behavioral semantic information. Thus, the behavioral semantic information in the surveillance video to be identified can be obtained. Furthermore, because the feature vectors of the keyframe images and the corresponding video content text feature vectors are projected into the same feature space during the fine-tuning of the preset behavior recognition model, the feature vectors of the keyframe images and the corresponding video content text feature vectors are made as close as possible to each other. This allows the extracted feature vectors of the keyframe images to focus more on the behavioral features within the keyframe images while ignoring the background features, thereby improving the recognition effect of multimodal behavior recognition.

[0100] Further reference Figure 2 As an implementation of the methods shown in the figures, this disclosure provides some embodiments of a multimodal behavior recognition device based on semantic augmentation. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0101] like Figure 2As shown, a multimodal behavior recognition device 200 based on semantic augmentation in some embodiments includes: an acquisition unit 201, a processing unit 202, a first determination unit 203, an extraction unit 204, a second determination unit 205, a fine-tuning unit 206, a collection unit 207, a generation unit 208, a receiving unit 209, and an input unit 2010. The acquisition unit 201 is configured to acquire a preset sample set, wherein each sample in the preset sample set includes a video of a person's behavior and tag information corresponding to the video of the person's behavior; the processing unit 202 is configured to perform semantic augmentation processing on each tag information included in the preset sample set based on each video of a person's behavior included in the preset sample set, so as to generate video content text information; the first determining unit 203 is configured to determine the generated video content text information as a set of video content text information; the extraction unit 204 is configured to extract at least one keyframe image from each video of a person's behavior in the video of a person's behavior, to obtain a keyframe image group; the second determining unit 205 is configured to determine the obtained keyframe image group as a set of keyframe image groups; the fine-tuning unit 206 is configured to... The aforementioned video content text information set and the aforementioned keyframe image set are used to fine-tune the preset behavior recognition model to obtain a behavior recognition model. The behavior recognition model stores an alignment feature information set created during the fine-tuning process. Acquisition unit 207 is configured to acquire the surveillance video to be recognized via a surveillance camera. Generation unit 208 is configured to generate a set of keyframe images to be recognized based on the surveillance video. Receiving unit 209 is configured to receive user-inputted behavior semantic query information. Input unit 2010 is configured to input the set of keyframe images to be recognized and the behavior semantic query information into the behavior recognition model, so that the behavior recognition model can perform behavior recognition processing on the set of keyframe images and the behavior semantic query information based on the alignment feature information set to obtain behavior semantic information.

[0102] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the method described above correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.

[0103] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0104] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0105] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0106] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0107] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0108] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0109] The computer-readable medium may be included in an electronic device or may exist independently, not assembled into the electronic device. The computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a preset sample set, wherein each sample in the preset sample set includes a video of human behavior and tag information corresponding to the video of human behavior; based on each video of human behavior included in the preset sample set, perform semantic augmentation processing on each tag information included in the preset sample set to generate video content text information; determine the generated video content text information as a set of video content text information; for each video of human behavior, extract at least one keyframe image from the video of human behavior to obtain a set of keyframe images; and determine the obtained set of keyframe images as a set of video content text information. A set of keyframe images is generated; based on the above video content text information set and the above keyframe image set, the preset behavior recognition model is fine-tuned to obtain the behavior recognition model, wherein the behavior recognition model stores the alignment feature information set created during the fine-tuning process; a monitoring video to be recognized is acquired through a monitoring camera; based on the above monitoring video to be recognized, a set of monitoring keyframe images to be recognized is generated; behavioral semantic query information input by the user is received; the above monitoring keyframe image set to be recognized and the above behavioral semantic query information are input into the behavior recognition model, so that the behavior recognition model can perform behavior recognition and solution processing on the above monitoring keyframe image set to be recognized and the above behavioral semantic query information based on the above alignment feature information set to obtain behavioral semantic information.

[0110] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0112] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a processing unit, a first determining unit, an extraction unit, a second determining unit, a fine-tuning unit, a acquisition unit, a generation unit, a receiving unit, and an input unit. The names of these units do not necessarily limit the specific unit; for example, the first determining unit may also be described as "a unit that determines the generated video content text information as a video content text information set."

[0113] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0114] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of technical features, but should also cover other technical solutions formed by arbitrary combinations of technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A multimodal behavior recognition method based on semantic augmentation, comprising: Obtain a preset sample set, wherein each sample in the preset sample set includes a video of a person's behavior and tag information corresponding to the video of the person's behavior; Based on the videos of various human behaviors included in the preset sample set, semantic augmentation processing is performed on each tag information included in the preset sample set to generate video content text information. The generated text information of each video content is defined as a video content text information set; For each video of a person's behavior, extract at least one keyframe image from the video of the person's behavior to obtain a keyframe image group. The obtained keyframe image groups are defined as a keyframe image group set. Based on the video content text information set and the key frame image set, the preset behavior recognition model is fine-tuned to obtain the behavior recognition model, wherein the behavior recognition model stores the alignment feature information set created during the fine-tuning process; The surveillance video to be identified is collected through surveillance cameras; Based on the surveillance video to be identified, generate a group of key frame images of the surveillance video to be identified; Receive user-input semantic query information related to behavior; The keyframe image group to be identified and the behavioral semantic query information are input into the behavior recognition model, so that the behavior recognition model can perform behavior recognition and solution processing on the keyframe image group to be identified and the behavioral semantic query information based on the alignment feature information set, and obtain behavioral semantic information.

2. The method according to claim 1, wherein, The step involves semantically augmenting each tag information in the preset sample set based on the videos of various human behaviors, to generate video content text information, including: The video of a person’s behavior that corresponds to the tag information in the video of a person’s behavior is identified as the video of the target person’s behavior. The tag information and the target person's behavior video are input into each preset video content text generation model to obtain each initial video content text information; The initial video content text information is subjected to quality screening to obtain the video content text information.

3. The method according to claim 1, wherein, The step of extracting at least one keyframe image from the video of the person's behavior to obtain a keyframe image group includes: Based on the aforementioned video of the person's behavior, a sequence of video frames of the person's behavior is generated; The number of video frames representing human behavior in the sequence of video frames representing human behavior is determined as the total number of video frames representing human behavior. Based on the total number of frames in the video of the person's behavior, at least one random index information is generated; Based on the at least one random index information, at least one keyframe image is extracted from the video frame sequence of the person's behavior to obtain a keyframe image group.

4. The method according to claim 1, wherein, The preset behavior recognition model includes a pre-trained image mapping module, a pre-trained text mapping module, a preset image feature module, and a preset text feature module. The behavior recognition model is further refined based on the video content text information set and the keyframe image set, resulting in a fine-tuned version. The set of keyframe images is input into the pre-trained image mapping module included in the preset behavior recognition model to obtain the image feature information set; The video content text information set is masked to obtain the masked video content text information set; The masked video content text information set is input into the pre-trained text mapping module included in the preset behavior recognition model to obtain the video content text feature information set; Based on the image feature information set and the video content text feature information set, the preset behavior recognition model is fine-tuned to obtain the behavior recognition model.

5. The method according to claim 4, wherein the step of fine-tuning the preset behavior recognition model based on the image feature information set and the video content text feature information set to obtain the behavior recognition model includes: Based on the image feature information set and the text feature information set, the following fine-tuning steps are performed: At least one image feature from the image feature information set is input into a preset image feature module to obtain at least one image projection feature. At least one text feature from the text feature information set is input into a preset text feature module to obtain at least one text projection feature; Based on a preset multimodal alignment loss function, alignment loss values ​​are generated by processing at least one image projection feature and at least one text projection feature to obtain various alignment loss values. Each alignment loss value is compared with a preset threshold to obtain comparison information; In response to the determination that the comparison information meets the preset conditions, the fine-tuned preset behavior recognition model is determined as the behavior recognition model, and the image feature information set and text feature information set are extracted and aligned based on the behavior recognition model to generate an aligned feature information set; In response to the determination that the comparison information does not meet the preset conditions, the parameters of the preset image feature module and the preset text feature module are adjusted, and at least one unused image feature information is used as the image feature information set, at least one unused text feature information is used as the text feature information set, and the preset behavior recognition model with adjusted parameters is executed again for fine-tuning.

6. The method according to claim 5, wherein, The behavior recognition model includes a pre-trained image mapping module, a pre-trained text mapping module, a trained image feature module, and a trained preset text feature module. The model also includes an alignment process for extracting and aligning image feature information sets and text feature information sets based on the behavior recognition model to generate an aligned feature information set, including: The text feature information set is input into the pre-trained text feature module in the behavior recognition model to generate the target text projection feature information set; The image feature information set is input into the trained image feature module in the behavior recognition model to generate the target video projection feature information set; The target text projection feature information set and the target video projection feature information set are subjected to feature alignment processing to obtain an aligned feature information set.

7. A multimodal behavior recognition device based on semantic augmentation, comprising: The acquisition unit is configured to acquire a preset sample set, wherein each sample in the preset sample set includes a video of a person's behavior and tag information corresponding to the video of the person's behavior; The processing unit is configured to perform semantic augmentation processing on each tag information included in the preset sample set based on the individual behavior videos included in the preset sample set, so as to generate video content text information. The first determining unit is configured to determine the generated video content text information as a video content text information set. The extraction unit is configured to extract at least one keyframe image from each of the individual behavior videos to obtain a keyframe image group. The second determining unit is configured to determine the obtained keyframe image groups as a keyframe image group set. The fine-tuning unit is configured to fine-tune a preset behavior recognition model based on the video content text information set and the key frame image set to obtain a behavior recognition model, wherein the behavior recognition model stores an alignment feature information set created during the fine-tuning process; The acquisition unit is configured to acquire surveillance video to be identified via a surveillance camera. The generation unit is configured to generate a group of key frame images of the surveillance video to be identified based on the video to be identified. The receiving unit is configured to receive behavioral semantic query information input by the user. The input unit is configured to input the group of monitoring keyframe images to be identified and the behavioral semantic query information into the behavior recognition model, so that the behavior recognition model can perform behavior recognition and solution processing on the group of monitoring keyframe images to be identified and the behavioral semantic query information based on the alignment feature information set, and obtain behavioral semantic information.

8. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Small sample deep learning multi-modal sign language recognition method based on key frame sampling

    CN111666845A

  • Video behavior recognition method, device and equipment based on multi-mode large model fine tuning

    CN119495127A

  • Learning state monitoring method and device, equipment and medium

    CN120783402A

  • Multi-task emotion recognition method for embedding fine-grained image blocks

    CN121074952A

  • Driver risk identification method and apparatus, terminal device and readable storage medium

    WO2025217803A1