Personnel abnormal behavior detection method and system based on visual language large model

Through the detection method of personnel abnormal behavior based on visual language big model, combined with monitoring video streams and text instructions for multimodal feature processing, the problems of low detection accuracy and insufficient generalization in the prior art are solved, and more efficient and flexible abnormal behavior detection is achieved.

CN119992641APending Publication Date: 2025-05-13TERMINUS GENERAL TECH
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202411844645.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing personnel abnormal behavior detection methods have low detection accuracy, are affected by factors such as monitoring scene complexity, lighting changes, occlusion, etc., and are insufficient generalization, and the model has poor adaptability to new abnormal behaviors or cross-scene applications.

Method used

Using a human abnormal behavior detection method based on a visual language model, the visual language model is used to continuously obtain monitoring video streams and combine text instructions to extract and process multimodal features to generate abnormal behavior judgment information.

Benefits of technology

It improves the accuracy of personnel abnormal behavior detection, enhances the generalization performance of the model, enables it to adapt to behavior recognition in different scenarios, and has strong scalability, and can be seamlessly integrated with other security technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992641A_ABST
    Figure CN119992641A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a personnel abnormal behavior detection method and system based on a visual language large model. The method is applied to the technical field of data processing, and comprises the following steps: continuously obtaining monitoring video frames; if a text instruction is received, determining a corresponding first text feature vector; determining a monitoring video data set, and inputting the monitoring video data set into a preset visual language large model to obtain a corresponding first visual feature vector; and connecting the first text feature vector with the first visual feature vector to obtain a first multi-modal feature vector, and processing the first multi-modal feature vector through a language model of a preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction. According to the scheme, the accuracy of personnel abnormal behavior detection is improved, and the method has strong generalization performance and can adapt to behavior recognition in different scenes; and the system has strong expandability and can be seamlessly integrated with other security technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a method and system for detecting abnormal behavior of personnel based on a large visual language model. Background Art

[0002] In public places and specific areas, people's behavior patterns not only reflect the diversity of social activities, but also play a vital role in public safety and order. These places are usually crowded with people. Once abnormal behaviors occur, such as violent conflicts, illegal gatherings or smoking, they will quickly lead to safety risks and disorder. Therefore, effectively monitoring and identifying people's behavior, and timely discovering and responding to potential threats are of great significance to ensuring public safety and maintaining social order.

[0003] Existing methods for detecting abnormal behavior of personnel are mainly based on computer vision technology, that is, firstly, the position of personnel in the monitoring area is detected through target detection and tracking methods, and then behavior recognition methods are used to further determine whether there are abnormal behaviors of personnel in the monitoring screen.

[0004] However, this type of method has low detection accuracy. Affected by factors such as the complexity of the monitoring scene, changes in lighting, and occlusion, target detection and tracking are prone to errors, resulting in misjudgment or missed detection of behavior recognition. On the other hand, the generalization is insufficient, and the model relies on a large amount of labeled data for training, which has poor adaptability to new abnormal behaviors or cross-scenario applications. Summary of the invention

[0005] In order to solve the shortcomings of the existing technology, the present invention provides a method and system for detecting abnormal behavior of personnel based on a large visual language model. The present invention solves the technical problems of the existing methods for detecting abnormal behavior of personnel, such as low detection accuracy, insufficient generalization, model reliance on a large amount of labeled data for training, and poor adaptability to new abnormal behaviors or cross-scenario applications.

[0006] According to a first aspect of the present disclosure, there is provided a method for detecting abnormal behavior of personnel based on a large visual language model, comprising: continuously acquiring a monitoring video stream, and continuously acquiring monitoring video frames from the monitoring video stream according to a preset fixed time interval;

[0007] If a text instruction is received from the control center, the current system time is obtained, the text instruction is input into a preset visual language model, and the text instruction is processed by a text lexicalizer of the preset visual language model to obtain a first text feature vector corresponding to the text instruction;

[0008] Determine a monitoring video data set according to the current system time, a preset video data set combination mode, and the monitoring video frame, input the monitoring video data set into a preset visual language large model, process the monitoring video data set through a visual encoder and a feature adapter of the preset visual language large model, and obtain a first visual feature vector corresponding to the monitoring video data set;

[0009] The first text feature vector is connected with the first visual feature vector through a preset visual language large model to obtain a first multimodal feature vector, and the first multimodal feature vector is processed through a language model of the preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to the control center.

[0010] According to a second aspect of the present disclosure, a method system for detecting abnormal behavior of personnel based on a large visual language model is provided for executing the method as described in the first aspect, comprising: a monitoring video frame acquisition module, for continuously acquiring a monitoring video stream, and continuously acquiring monitoring video frames from the monitoring video stream according to a preset fixed time interval;

[0011] A text feature vector determination module is used to obtain the current system time when a text instruction sent by the control center is received, input the text instruction into a preset visual language model, process the text instruction through a text lexicalizer of the preset visual language model, and obtain a first text feature vector corresponding to the text instruction;

[0012] A visual feature vector determination module is used to determine a monitoring video data set according to the current system time, a preset video data set combination mode and the monitoring video frame, input the monitoring video data set into a preset visual language large model, and process the monitoring video data set through a visual encoder and a feature adapter of the preset visual language large model to obtain a first visual feature vector corresponding to the monitoring video data set;

[0013] The abnormal behavior judgment information determination module is used to connect the first text feature vector with the first visual feature vector through a preset visual language large model to obtain a first multimodal feature vector, and process the first multimodal feature vector through a language model of the preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and send the abnormal behavior judgment information to the control center.

[0014] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method described in the first aspect is implemented.

[0015] In the above-provided method, system and device for detecting abnormal behavior of personnel based on a visual language large model, the disclosed embodiment improves the accuracy of detecting abnormal behavior of personnel by integrating visual and language information. The multimodal large model constructed using deep learning has strong generalization performance and can adapt to behavior recognition in different scenarios. In addition, the system has strong scalability and can be seamlessly integrated with other security technologies to further improve the overall security effectiveness. This makes the system not only efficient and flexible, but also suitable for various complex security needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 A schematic diagram of a method for detecting abnormal behavior of personnel based on a large visual language model according to an embodiment of the present disclosure is shown;

[0018] Figure 2 A schematic diagram of a method for detecting abnormal behavior of personnel based on a large visual language model according to an embodiment of the present disclosure is shown:

[0019] Figure 3 A schematic block diagram of a method system for detecting abnormal behavior of personnel based on a large visual language model according to an embodiment of the present disclosure is shown;

[0020] Figure 4 A block diagram of an exemplary electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0021] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.

[0022] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor represent the necessary logical order between them. It should also be understood that in the embodiments of the present disclosure, "multiple" can refer to two or more, and "at least one" can refer to one, two or more. It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, in the absence of explicit limitation or contrary revelation given in the context, it can generally be understood as one or more. In addition, the term "and / or" in the present disclosure is only a description of the association relationship of the associated objects, indicating that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present disclosure generally indicates that the associated objects before and after are an "or" relationship. It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and the same or similar parts can refer to each other. For the sake of brevity, they will not be repeated one by one.

[0023] At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present disclosure and its application or use. The techniques, methods and devices known to ordinary technicians in the relevant fields may not be discussed in detail, but where appropriate, the techniques, methods and devices should be considered part of the specification. It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0024] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0025] Figure 1 The following is a flow chart of a method for detecting abnormal behavior of personnel based on a large visual language model provided by an embodiment of the present disclosure. Figure 1 As shown, the method includes:

[0026] S101, continuously acquiring a surveillance video stream, and continuously acquiring surveillance video frames from the surveillance video stream according to a preset fixed time interval.

[0027] A surveillance video stream can be video data collected and transmitted in real time by a surveillance camera or other video acquisition device, which presents the video content in the form of a continuous frame sequence.

[0028] The preset fixed time interval may be a time interval parameter for extracting video frames from a video stream, and may be in seconds or milliseconds.

[0029] A surveillance video frame can be a single static image in a video stream and is the basic unit of video.

[0030] A video streaming protocol (such as RTSP, RTMP) can be used to connect to the monitoring device to obtain the monitoring video stream, extract frames from the monitoring video stream at a preset fixed time interval, and adjust the resolution of each image to 640×640, thereby obtaining a series of monitoring video frames [I 0 , I 1 , ..., I t ], where I t Represents the image acquired at time t.

[0031] S102, if a text instruction is received from the control center, the current system time is obtained, the text instruction is input into a preset visual language model, and the text instruction is processed by a text tokenizer of the preset visual language model to obtain a first text feature vector corresponding to the text instruction.

[0032] The control center can be the core device or system used to manage and schedule the entire monitoring system, responsible for issuing text instructions to guide the behavior of the model.

[0033] The text instruction can be a natural language task description sent by the control center to the visual language model. For example: "Please carefully analyze the content in the video to check whether there are any abnormal behaviors such as people fighting or smoking."

[0034] The current system time may be a timestamp when the instruction is received.

[0035] The preset visual language model can be a model trained based on Qwen2-VL, which has been loaded into the system and is ready to perform reasoning tasks. It combines vision and language capabilities and can understand and answer questions related to text instructions and video / image content.

[0036] A text tokenizer can be part of a model that breaks down text instructions in natural language into feature vectors that the model can understand.

[0037] The first text feature vector may be a numerical representation of a text instruction after being processed by a text tokenizer, and is usually a high-dimensional vector.

[0038] It can receive instructions from the control center and record the current system time. After receiving the text instruction, it will be input into the preset visual language model and processed by the text tokenizer of the model. The text tokenizer performs word segmentation and vectorization operations on the natural language text to generate the corresponding first text feature vector, providing input feature support for subsequent multimodal reasoning tasks.

[0039] The preset large visual language model (such as Qwen2-VL) is trained in two stages. The first stage is pre-training, which uses large-scale image / video data (such as ImageNet, COCO) and text corpus (such as Wikipedia) for multimodal alignment learning. The tasks include image-text matching, image-text generation and multimodal classification, and self-supervised learning and contrastive learning strategies are used to semantically align visual and text features. The second stage is fine-tuning. For specific tasks (such as abnormal behavior detection), the model is fine-tuned for classification using annotated abnormal behavior datasets, and the task performance is optimized by freezing the base layer parameters and adding classification heads. During the model training process, the model's multimodal understanding and classification capabilities are enhanced through loss functions (such as multimodal contrast loss and classification loss), ultimately achieving efficient abnormal behavior detection.

[0040] S103, determining a monitoring video data set according to the current system time, a preset video data set combination method and the monitoring video frame, inputting the monitoring video data set into a preset visual language large model, processing the monitoring video data set through a visual encoder and a feature adapter of the preset visual language large model, and obtaining a first visual feature vector corresponding to the monitoring video data set.

[0041] The preset video data set combination method may refer to a strategy for the system to select and combine available video data according to current task requirements, such as extracting the T-frame image closest to the current time in the surveillance image sequence.

[0042] A surveillance video dataset can be a set of video frame data obtained from a surveillance system, which is used for model reasoning after preprocessing, such as a continuous T-frame image sequence extracted by time filtering.

[0043] The visual encoder can be a deep learning model module (such as the ViT architecture) that is responsible for extracting high-dimensional visual features from the input image sequence. A feature adapter (such as MLP or Q-Former) can be used to map the visual features generated by the visual encoder to the text feature space, thereby achieving multimodal alignment.

[0044] The first visual feature vector may refer to a feature representation generated by the visual language large model after processing the surveillance video data set through a visual encoder and a feature adapter, and is used to describe the content of the input image sequence.

[0045] According to the current system time and the preset combination strategy, the T-frame video images closest to the current time can be extracted from the monitoring system to form a monitoring video dataset. The monitoring video dataset is input into the visual language model, and high-dimensional visual features are extracted through the visual encoder in turn. These features are then mapped to the text feature space through the feature adapter to generate the first visual feature vector.

[0046] S104, connecting the first text feature vector with the first visual feature vector through a preset visual language model to obtain a first multimodal feature vector, and processing the first multimodal feature vector through a language model of the preset visual language model to obtain abnormal behavior judgment information corresponding to the text instruction.

[0047] The first multimodal feature vector may be a joint representation obtained by connecting the first text feature vector (features generated by the text instruction) and the first visual feature vector (features generated by the surveillance video dataset). It integrates information from both language and vision modalities to describe the semantic relationship between the text instruction and the video content.

[0048] Abnormal behavior judgment information can be the output result of model reasoning, which is used to indicate whether there is abnormal behavior related to the text instruction (such as fighting, smoking, gathering, etc.) in the surveillance video or whether there is no abnormal behavior. The output form can include behavior category, probability distribution or specific description text.

[0049] The first text feature vector and the first visual feature vector can be connected by simple concatenation or weighting to form a first multimodal feature vector. The concatenation result maintains the independence of each modality while retaining the order relationship between the modalities. The connected multimodal feature vector combines the information of the text and visual modalities and has a shape of R (N+T)×d ;

[0050] Where R is a multimodal feature vector, which contains the concatenation of text and visual information; N is the number of word segments in the text; T is the number of input video or image frames; and d is the dimension of each feature vector.

[0051] The first multimodal feature vector is then input into a preset large language model (based on the Transformer-Decoder architecture). The language model uses a multi-layer self-attention mechanism to fuse and infer multimodal features and analyze the association between input instructions and video content. Through the reasoning calculation of the language model, the result corresponding to the text instruction is output, that is, abnormal behavior judgment information, which may include: whether there is abnormal behavior (such as "abnormal behavior exists" or "no abnormal behavior is detected"). The category of abnormal behavior (such as "fighting", "smoking", "running", etc.). Probability distribution or confidence (such as "fighting probability: 90%"). The abnormal behavior judgment information is then sent to the control center via wireless communication technology.

[0052] In the embodiment of the present application, the monitoring video stream is continuously acquired, and the monitoring video frames are continuously acquired from the monitoring video stream according to the preset fixed time interval; if a text instruction issued by the control center is received, the current system time is acquired, the text instruction is input into the preset visual language model, and the text instruction is processed by the text lexicalizer of the preset visual language model to obtain the first text feature vector corresponding to the text instruction; the monitoring video data set is determined according to the current system time, the preset video data set combination mode and the monitoring video frame, and the monitoring video data set is input into the preset visual language model, and the monitoring video data set is processed by the visual encoder and feature adapter of the preset visual language model to obtain the first visual feature vector corresponding to the monitoring video data set; the first text feature vector is connected with the first visual feature vector by the preset visual language model to obtain the first multimodal feature vector, and the first multimodal feature vector is processed by the language model of the preset visual language model to obtain the abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to the control center. Through the above-mentioned personnel abnormal behavior detection method based on the visual language model, the accuracy of personnel abnormal behavior detection is improved by integrating visual and language information. The multimodal large model built using deep learning has strong generalization performance and can adapt to behavior recognition in different scenarios. In addition, the system has strong scalability and can be seamlessly integrated with other security technologies to further improve the overall security effectiveness. This makes the system not only efficient and flexible, but also suitable for a variety of complex security needs.

[0053] Based on the above technical solution, optionally, after sending the abnormal behavior judgment information to the control center, the method further includes:

[0054] If the judgment error information sent by the control center is received, the actual behavior information is determined according to the judgment error information, and the preset visual language large model is adjusted according to the actual behavior information.

[0055] In this solution, the judgment error information can be that the model output result does not match the actual situation. When the model identifies certain behaviors, it may misjudge an event that should be identified as abnormal behavior as normal, or vice versa, misjudge a normal behavior as abnormal. For example, the model may misjudge the actual behavior information of "smoking" as normal. The model may misjudge "people discussing in a conference room" as "fighting", which is actually a normal behavior.

[0056] Actual behavior information can be the behavior that actually occurred in a specific surveillance video, confirmed by manual review or other mechanisms. It is verified data that conforms to the actual scenario.

[0057] When the control center or manual monitoring system receives the judgment error information, it means that the model has made a misjudgment in certain situations. For example, the model mistakenly marks smoking behavior as "normal behavior". Confirm the actual behavior information through manual review or in combination with other auxiliary tools (such as intelligent auxiliary annotation system, historical data feedback, etc.). For example, the control center staff confirms that the behavior is indeed "smoking" behavior based on video playback, sensor data or other clues, and the behavior occurs in a non-smoking area. Once the actual behavior information is confirmed, the system will adjust the preset visual language model based on the information. This adjustment may include analyzing the reasons for the model's misjudgment: finding out the specific reasons for the model's misjudgment, whether it is due to unclear visual features, or the fusion problem of the language model and visual data. Optimize model parameters: adjust the parameters related to the behavior in the model according to the actual behavior information. For example, enhance the model's ability to identify "smoking" behavior in non-smoking areas, or adjust the ability to understand scene semantics. Fine-tune the model: Use the actual behavior information with annotations for training and fine-tune the model so that it can better identify similar abnormal behaviors. The system then trains based on the new actual behavior information and the adjusted model to ensure that the accuracy of the model continues to improve. After that, the system can start using the adjusted model to handle future abnormal behavior detection tasks. After each update, the new model will be fed back to the control center for verification to determine whether the original misjudgment has been corrected.

[0058] In this solution, the model can correct misjudgments and improve its recognition accuracy in abnormal behavior detection through feedback of actual behavior information. The model is continuously adjusted and optimized so that it can more accurately identify and classify abnormal behaviors in different scenarios.

[0059] Figure 2 The following is a flow chart of a method for detecting abnormal behavior of personnel based on a large visual language model provided by an embodiment of the present disclosure. Figure 2 As shown, the method includes:

[0060] S201, continuously acquiring a surveillance video stream, and continuously acquiring surveillance video frames from the surveillance video stream according to a preset fixed time interval.

[0061] S202, if a text instruction is received from the control center, the current system time is obtained, the text instruction is input into a preset visual language model, the text instruction is processed by a text tokenizer of the preset visual language model, and a first text feature vector corresponding to the text instruction is obtained.

[0062] S203, determining a monitoring video data set according to the current system time, a preset video data set combination method and the monitoring video frame, inputting the monitoring video data set into a preset visual language large model, processing the monitoring video data set through a visual encoder and a feature adapter of the preset visual language large model, and obtaining a first visual feature vector corresponding to the monitoring video data set.

[0063] S204, processing the surveillance video data set through the scene recognition model of the preset visual language large model to obtain scene semantic labels, and determining the scene weights corresponding to the scene semantic labels according to the scene semantic labels and a preset scene weight allocation method.

[0064] The scene recognition model can be part of a preset visual language model, which is used to extract and identify the semantic information of the scene from the surveillance video dataset. Usually, this model analyzes the input surveillance video data through a deep learning architecture (such as a convolutional neural network CNN or a transformer Transformer) to identify and classify different scenes. The model combines the visual features in the video with the corresponding text information to better understand and describe the video content.

[0065] Scene semantic labels can be labels obtained by analyzing surveillance video datasets through scene recognition models, indicating the type or state of specific scenes appearing in the video. For example, a no-smoking area in a surveillance video may be identified as a "no-smoking area" label. This label is a concise summary of the surveillance video content, helping the system to identify whether the video involves scenes in a no-smoking area, and then determine whether abnormal behavior (such as smoking) occurs in the area.

[0066] The preset scene weight assignment method can be a scene-based rule or strategy for assigning different importance weights to different scene semantic labels. These weights reflect the priority or sensitivity of different scenes in abnormal behavior detection. For example, in a no-smoking area scene, since smoking behavior may violate regulations in this scene, a higher weight can be assigned to the "no-smoking area" scene. In contrast, in scenes such as offices, although smoking behavior may be regarded as abnormal behavior, the severity of its violation or violation is usually low, so the weight assigned may be lower.

[0067] The scene weight can be the corresponding weight of each scene semantic label, which represents the influence or priority of the scene on the detection result during the processing. The weight is usually a number, which may be assigned based on historical data, actual scene requirements, or other factors. For example, for no-smoking areas, the weight of no-smoking area scenes is higher because smoking behavior poses a greater potential threat to public health and safety. For example, if the video analysis system detects a "no-smoking area" scene and someone is smoking, the system will attach more importance to this information and judge it as an abnormal behavior.

[0068] Through the scene recognition model, the video data will be decomposed into different scene categories or states. Each scene will be marked with a semantic label, which represents a specific scene of the video content. For example, the label may be "office", "conference room", "corridor", "non-smoking area", etc. The scene recognition model generates labels such as "non-smoking area", "conference room", "rest area", etc. based on the input surveillance video data set. Each scene label will be assigned a scene weight according to the preset weight distribution method. This distribution method is usually set based on historical data analysis, business needs or the importance of the scene. For example: non-smoking area: because the non-smoking area is a high-risk area and there may be illegal smoking behavior, it is assigned a higher weight (such as 0.8). Conference room: The conference room is a relatively quiet environment with fewer abnormal behaviors, and it may be assigned a medium weight (such as 0.5). Empty corridor: Due to the small flow of people, it is assigned a lower weight (such as 0.3). The weight distribution rules can be adjusted according to system requirements to ensure that the system can sensitively respond to abnormal behaviors when processing different scenes. Then, according to the scene semantic label and the preset scene weight distribution method, the scene weight is assigned to the corresponding scene label. For example, the scene "No Smoking Area" might have a weight of 0.8, while the scene "Corridor" might have a weight of 0.3. The final output is the combination of each scene semantic label and its corresponding weight.

[0069] On the basis of the above technical solution, optionally, after determining the scene weight corresponding to the scene semantic label according to the scene semantic label and a preset scene weight allocation method, the method further includes:

[0070] The first text feature vector is weight-adjusted according to the scene semantic label, the scene weight, and a preset text weight adjustment formula to obtain a second text feature vector;

[0071] Correspondingly, the first text feature vector is connected with the second visual feature vector through the preset visual language macro model according to the preset multimodal fusion formula, the dynamic visual feature vector weight and the dynamic text feature vector weight to obtain a second multimodal feature vector, including;

[0072] The second text feature vector is connected with the second visual feature vector according to a preset multimodal fusion formula, a dynamic visual feature vector weight, and a dynamic text feature vector weight through a preset visual language macro model to obtain a third multimodal feature vector; wherein the preset multimodal fusion formula is:

[0073] M=α·V 2 +β·T 2 ;

[0074] Where M is the second multimodal feature vector; α is the dynamic visual feature vector weight; V 2 is the second visual feature vector; β is the dynamic text feature vector weight; T 1 is the second text feature vector;

[0075] Accordingly, the second multimodal feature vector is processed by the language model of the preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to the control center, including:

[0076] The third multimodal feature vector is processed by a language model of a preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to a control center.

[0077] In this solution, the preset text weight adjustment formula can be a mathematical formula for weighting and adjusting the text feature vector according to the scene semantic label (such as "no smoking area"). Its purpose is to dynamically adjust the importance of text instructions according to different scenarios, so as to better match the behavior pattern or anomaly detection task in the scenario.

[0078] The second text feature vector may be a result of weighted adjustment of the first text feature vector. Through the weight coefficient of the scene semantic label, the model can adjust the importance of the text feature according to the characteristics of the current scene, so as to better match the abnormal behavior recognition task in the current scene.

[0079] The third multimodal feature vector can be obtained by fusing the second text feature vector with the second visual feature vector. This vector integrates text and visual information and is adjusted according to a dynamic weight coefficient to improve the accuracy of the final multimodal analysis.

[0080] The scene semantic label, the scene weight and the preset text weight adjustment formula may be combined to obtain a second text feature vector.

[0081] In this solution, by adjusting the weight of text features, the sensitivity to specific abnormal behaviors can be effectively improved, especially in key scenarios. Weight adjustment helps to better integrate visual information and text information and improve the performance of multimodal models. In different scenarios, adjusting the weight of text information makes the integration between text information and visual information more accurate and improves the overall anomaly detection effect.

[0082] Based on the above technical solution, an optional, preset text weight adjustment formula is:

[0083] T 2 =T 1 ·W(S)

[0084] Among them, T 2 is the second text feature vector; T 1 is the first text feature vector; S is the scene semantic label; W(S) is the scene weight.

[0085] In this scenario, for example, the scene semantic label S is "no smoking area" and the text instruction is "someone is smoking". If W(S) is calculated as 1.5, then:

[0086] T 1 =[0.2, 0.5, 0.3] (the original feature vector of the text instruction)

[0087] W(S)=1.5

[0088] Then according to the formula, T 2 =T 1 ·W(S)=[0.2×1.5, 0.5×1.5, 0.3×1.5]=[0.3, 0.75, 0.45].

[0089] S205 , weight adjustment is performed on the first visual feature vector according to the scene semantic label, the scene weight, and a preset visual weight adjustment formula to obtain a second visual feature vector.

[0090] The preset visual weight adjustment formula can be used to weight the visual feature vector according to the scene semantic label and scene weight. This formula is designed to adjust the influence of different scenes on the model judgment according to their priority, especially in the abnormal behavior detection task, some scenes may be more sensitive or more relevant than others.

[0091] The second visual feature vector may be a visual feature vector after weight adjustment. Compared with the original visual feature vector, the second visual feature vector contains additional information based on the scene weight adjustment, so that the visual feature is more in line with the actual needs or priorities of the current scene. This feature vector is part of multimodal fusion and is usually used to further combine with the text feature vector for final reasoning.

[0092] The scene semantic label, the scene weight and the preset visual weight adjustment formula may be combined to obtain a second visual feature vector.

[0093] Based on the above technical solution, an optional, preset visual weight adjustment formula is:

[0094] V 2 =V 1 ·W(S)

[0095] Among them, V 2 is the second visual feature vector; V 1 is the first visual feature vector; S is the scene semantic label; W(S) is the scene weight.

[0096] In this scenario, for example, assume the following:

[0097] V 1 : The original visual feature vector V = [0.5, 0.7, 0.8, O.6], which represents the visual features of the video frame (assumed to be a 4-dimensional vector).

[0098] S: The scene semantic label is “no smoking area”.

[0099] W(S): According to the preset rule, the scene weight of “no smoking area” is W(S)=1.5.

[0100] Then, after weight adjustment, the adjusted visual feature vector V 2 for:

[0101] V 2 =V 11.5 = [0.5, 0.7, 0.8, 0.6] × 1.5 = [0.75, 1.05, 1.2, 0.9]. Through this adjustment, the value of the visual feature vector is amplified by 1.5 times, so that abnormal behaviors related to visual features in the "no smoking area" scene (such as smoking, etc.) can receive higher attention and priority.

[0102] S206, connecting the first text feature vector and the second visual feature vector according to a preset multimodal fusion formula through a preset visual language macro model to obtain a second multimodal feature vector; wherein the preset multimodal fusion formula is:

[0103] M=α·V 2 +β·T 1 ;

[0104] Where M is the second multimodal feature vector; α is the dynamic visual feature vector weight; V 2 is the second visual feature vector; β is the dynamic text feature vector weight; T 1 is the first text feature vector.

[0105] The second multimodal feature vector may be a fusion result obtained by weighted fusion of the first text feature vector and the second visual feature vector, which not only contains text and visual information, but also integrates the importance or weight of the two types of information.

[0106] The second visual feature vector and the first text feature vector can be weighted fused by the formula. α and β are hyperparameters set according to the needs of the scene and task, which control the contribution of visual features and text features to the final multimodal feature vector M. For example, if the task requires visual information to be more important, α can be increased and β can be reduced. During calculation, after multiplying by the weighting coefficient, the visual feature vector V and the text feature vector T are added to obtain the fused second multimodal feature vector M.

[0107] On the basis of the above technical solution, optionally, after weight adjustment is performed on the first visual feature vector according to the scene semantic label, the scene weight and a preset visual weight adjustment formula to obtain the second visual feature vector, the method further includes:

[0108] The surveillance video data set, text instructions and scene semantic labels are input into a preset weight confirmation model to obtain dynamic visual feature vector weights and dynamic text feature vector weights;

[0109] Correspondingly, the first text feature vector and the second visual feature vector are connected according to a preset multimodal fusion formula through a preset visual language macro model to obtain a second multimodal feature vector, including:

[0110] The first text feature vector and the second visual feature vector are connected by a preset visual language macro model according to a preset multimodal fusion formula, a dynamic visual feature vector weight, and a dynamic text feature vector weight to obtain a second multimodal feature vector; wherein the preset multimodal fusion formula is:

[0111] M=α·V 2 +β·T 1 ;

[0112] Where M is the second multimodal feature vector; α is the dynamic visual feature vector weight; V 2 is the second visual feature vector; β is the dynamic text feature vector weight; T 1 is the first text feature vector.

[0113] In this solution, the preset weight confirmation model can be a trained deep learning model that can generate visual feature vector weights and text feature vector weights based on multimodal input data (such as surveillance video datasets, text instructions, scene semantic labels, etc.). Its role is to learn and assign dynamic weights so that feature vectors of different modalities can be weighted and adjusted according to input scenes and contexts in subsequent tasks.

[0114] The dynamic visual feature vector weight may be a weight associated with the input visual data (surveillance video features), indicating the importance of the visual modality under the current scene and instruction semantics.

[0115] The dynamic text feature vector weight may be a weight associated with the input text instruction, indicating the importance of the text modality in the current scenario.

[0116] The surveillance video dataset, text instructions, and scene semantic labels can be converted into a format suitable for the input of the weight confirmation model, and the video frames are feature extracted. The visual feature vector V is generated through a pre-trained visual encoder (such as ResNet, ViT). The text instructions are encoded into a text feature vector T through a natural language processing (NLP) model (such as BERT, GPT). Scene semantic label: The scene label is generated using a preset scene recognition model and mapped to a numerical value or embedded vector representation S. The final input form is:

[0117] Input = [V, T, S]

[0118] Input the preprocessed feature vector and label into the preset weight confirmation model for processing. The model structure usually includes the following parts:

[0119] Modal feature fusion module: It uses attention mechanisms (such as multi-head self-attention) or multimodal fusion methods (such as simple weighted formula, linear weighting or deep fusion) to integrate visual and textual features while referring to scene semantic labels.

[0120] Dynamic weight prediction module: Use a feedforward neural network (FFNN) or a Transformer-based model to output visual feature vector weights and text feature vector weights based on the fused features.

[0121] In this solution, by dynamically adjusting the weights of visual and text features, more efficient multimodal fusion can be achieved, the model's scene perception ability can be enhanced, the accuracy of behavior recognition can be improved, and the system can be ensured to have strong adaptability and scalability. This approach can not only improve the efficiency of task execution, but also continuously optimize model performance in a constantly changing environment.

[0122] Based on the above technical solution, the optional training process of the preset weight confirmation model includes:

[0123] Obtaining historical weight records, and determining historical surveillance video data sets, historical text instructions, historical scene semantic labels, historical visual feature vector weights, and historical text feature vector weights according to the historical weight records;

[0124] Creating a first data set according to the historical surveillance video data set, historical text instructions, and historical scene semantic labels, labeling the visual feature vector weight labels of the first data set according to the historical visual feature vector weights, and labeling the text feature vector weight labels of the first data set according to the historical text feature vector weights;

[0125] A weight confirmation model is constructed, and the weight confirmation model is trained according to the first data set, the visual feature vector weight label, and the text feature vector weight label until the weight confirmation model reaches a preset weight confirmation model training standard.

[0126] In this solution, the historical weight record may refer to the weight data about the visual features and text features that have been calculated and recorded during the past training process.

[0127] Historical surveillance video dataset: Past surveillance video datasets, including scenes, activities, and background information, which have been used for model training.

[0128] The historical text instructions may refer to text instructions used during training, such as “Please check whether there is abnormal behavior in the video”, etc.

[0129] Historical scene semantic labels may refer to labels generated for video data by a scene recognition model, indicating the scene type in the video (such as “no smoking area”, “office”, etc.).

[0130] The historical visual feature vector weights may be weights of visual feature vectors obtained through model training in history, and these weights are used to measure the importance of visual features in the final decision.

[0131] The historical text feature vector weight may be the weight of the text feature vector obtained through training in history, and measures the importance of the text instruction in the decision-making.

[0132] The first data set may be a new data set created based on historical data, and is mainly used for training a weight confirmation model.

[0133] Visual feature vector weight labels can be weights assigned to each visual feature vector (e.g., visual features extracted from surveillance videos). These labels are generated based on historical data and annotate the relative importance of each visual feature vector. For example, in a specific scenario, some visual features may be more important than other features, and the visual feature vector weight labels reflect this.

[0134] Text feature vector weight labels can be weights assigned to text feature vectors (such as features extracted from text instructions). These labels help the model identify the importance of different text instructions in different scenarios. For example, for a text instruction involving an emergency, the model may set its weight higher.

[0135] The preset weight confirmation model training standard can refer to the standard used to evaluate and determine whether the training process is completed and achieves the expected effect. Specifically, it can include: Loss function: used to measure the difference between the model's predicted value and the actual label. Accuracy: the model's prediction accuracy for the training data set. Overfitting / underfitting check: ensure that the model is not overfitted on the training set or performs poorly on the test set. Training stability: whether the model's training process is stable and there are no abnormal fluctuations.

[0136] By obtaining historical weight records, the system can extract historical surveillance video data sets, text instructions, scene semantic labels, and corresponding visual and text feature vector weights, and then construct a first data set. The data set provides data for training weight confirmation models by annotating weight labels of historical visual features and text feature vectors. Then a deep learning model is designed, the input of the model is the visual features of video data, the text features of text data, and the scene semantic labels, and the output is the adjusted visual and text feature vector weights. Model architectures such as multi-layer perceptron (MLP), convolutional neural network (CNN), **recurrent neural network (RNN)** can be used. Then the first data set is input into the model. The model learns how to adjust weights based on scene semantic labels, visual features, and text features by training the input data. During the training process, the model optimizes parameters by comparing the predicted weights with the actual weight labels (obtained from historical data). The training process continues until the model reaches the preset performance standard, ensuring that the model can accurately confirm and adjust the weights of visual and text features, thereby improving the accuracy of abnormal behavior detection and scene analysis.

[0137] S207, processing the second multimodal feature vector by using the language model of the preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and sending the abnormal behavior judgment information to the control center.

[0138] In this embodiment, by assigning different weights to different scenarios, the model can better focus on key scenarios where abnormal behaviors may occur and reduce excessive attention to irrelevant scenarios, thereby improving the accuracy of abnormal behavior detection.

[0139] Figure 3 A schematic block diagram of a method system for detecting abnormal behavior of personnel based on a large visual language model provided in an embodiment of the present disclosure. The system includes:

[0140] The monitoring video frame acquisition module 301 is used to continuously acquire the monitoring video stream and continuously acquire the monitoring video frames from the monitoring video stream according to a preset fixed time interval;

[0141] The text feature vector determination module 302 is used to obtain the current system time if a text instruction sent by the control center is received, input the text instruction into a preset visual language model, process the text instruction through a text lexicalizer of the preset visual language model, and obtain a first text feature vector corresponding to the text instruction;

[0142] A visual feature vector determination module 303 is used to determine a monitoring video data set according to the current system time, a preset video data set combination mode and the monitoring video frame, input the monitoring video data set into a preset visual language large model, and process the monitoring video data set through a visual encoder and a feature adapter of the preset visual language large model to obtain a first visual feature vector corresponding to the monitoring video data set;

[0143] The abnormal behavior judgment information determination module 304 is used to connect the first text feature vector with the first visual feature vector through a preset visual language model to obtain a first multimodal feature vector, and process the first multimodal feature vector through a language model of the preset visual language model to obtain abnormal behavior judgment information corresponding to the text instruction, and send the abnormal behavior judgment information to the control center.

[0144] Figure 4 A schematic block diagram of an electronic device 400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0145] The electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in the ROM 402 or a computer program loaded from the storage unit 408 into the RAM 404. In the RAM 404, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 404 are connected to each other via a bus 404. An I / O interface 405 is also connected to the bus 404.

[0146] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0147] The computing unit 401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above, such as the method for detecting abnormal behavior of personnel based on a large visual language model. For example, in some embodiments, the method for detecting abnormal behavior of personnel based on a large visual language model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 400 via ROM402 and / or the communication unit 409. When the computer program is loaded into RAM404 and executed by the computing unit 401, one or more steps of the method for detecting abnormal behavior of personnel based on a large visual language model described above may be executed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the method for detecting abnormal behavior of personnel based on a large visual language model in any other appropriate manner (for example, by means of firmware).

[0148] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0149] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0150] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0152] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0153] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0154] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0155] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for detecting abnormal behavior of personnel based on a large visual language model, characterized in that: The method comprises: Continuously obtain the monitoring video stream, and continuously obtain the monitoring video frames from the monitoring video stream according to a preset fixed time interval; If a text instruction is received from the control center, the current system time is obtained, the text instruction is input into a preset visual language model, and the text instruction is processed by a text lexicalizer of the preset visual language model to obtain a first text feature vector corresponding to the text instruction; Determine a monitoring video data set according to the current system time, a preset video data set combination mode, and the monitoring video frame, input the monitoring video data set into a preset visual language large model, process the monitoring video data set through a visual encoder and a feature adapter of the preset visual language large model, and obtain a first visual feature vector corresponding to the monitoring video data set; The first text feature vector is connected with the first visual feature vector through a preset visual language large model to obtain a first multimodal feature vector, and the first multimodal feature vector is processed through a language model of the preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to the control center.

2. The method according to claim 1, characterized in that: in, After obtaining the first visual feature vector corresponding to the surveillance video data set, the method further includes: The surveillance video data set is processed by a scene recognition model of a preset visual language large model to obtain a scene semantic label, and a scene weight corresponding to the scene semantic label is determined according to the scene semantic label and a preset scene weight allocation method; weight-adjusting the first visual feature vector according to the scene semantic label, the scene weight, and a preset visual weight adjustment formula to obtain a second visual feature vector; Correspondingly, the first text feature vector is connected with the first visual feature vector through a preset visual language macro model to obtain a first multimodal feature vector, including: The first text feature vector and the second visual feature vector are connected by a preset visual language macro model according to a preset multimodal fusion formula to obtain a second multimodal feature vector; wherein the preset multimodal fusion formula is: M = α·V2+β·T1; Wherein, M is the second multimodal feature vector; α is the dynamic visual feature vector weight; V2 is the second visual feature vector; β is the dynamic text feature vector weight; T1 is the first text feature vector; Accordingly, the first multimodal feature vector is processed by a language model of a preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to the control center, including: The second multimodal feature vector is processed by a language model of a preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to a control center.

3. The method according to claim 2, characterized in that in, The preset visual weight adjustment formula is: V2=V1·W(S) Among them, V2 is the second visual feature vector; V1 is the first visual feature vector; S is the scene semantic label; W(S) is the scene weight.

4. The method according to claim 2, characterized in that in, After weight adjustment is performed on the first visual feature vector according to the scene semantic label, the scene weight, and a preset visual weight adjustment formula to obtain a second visual feature vector, the method further includes: The surveillance video data set, text instructions and scene semantic labels are input into a preset weight confirmation model to obtain dynamic visual feature vector weights and dynamic text feature vector weights; Correspondingly, the first text feature vector and the second visual feature vector are connected according to a preset multimodal fusion formula through a preset visual language macro model to obtain a second multimodal feature vector, including: The first text feature vector and the second visual feature vector are connected by a preset visual language macro model according to a preset multimodal fusion formula, a dynamic visual feature vector weight, and a dynamic text feature vector weight to obtain a second multimodal feature vector; wherein the preset multimodal fusion formula is: M = α·V2+β·T1; Among them, M is the second multimodal feature vector; α is the dynamic visual feature vector weight; V2 is the second visual feature vector; β is the dynamic text feature vector weight; T1 is the first text feature vector.

5. The method according to claim 4, characterized in that in, The training process of the preset weight confirmation model includes: Obtaining historical weight records, and determining historical surveillance video data sets, historical text instructions, historical scene semantic labels, historical visual feature vector weights, and historical text feature vector weights according to the historical weight records; Creating a first data set according to the historical surveillance video data set, historical text instructions, and historical scene semantic labels, labeling the visual feature vector weight labels of the first data set according to the historical visual feature vector weights, and labeling the text feature vector weight labels of the first data set according to the historical text feature vector weights; A weight confirmation model is constructed, and the weight confirmation model is trained according to the first data set, the visual feature vector weight label, and the text feature vector weight label until the weight confirmation model reaches a preset weight confirmation model training standard.

6. The method according to claim 4, characterized in that in, After determining the scene weight corresponding to the scene semantic label according to the scene semantic label and a preset scene weight allocation method, the method further includes: The first text feature vector is weight-adjusted according to the scene semantic label, the scene weight, and a preset text weight adjustment formula to obtain a second text feature vector; Correspondingly, the first text feature vector is connected with the second visual feature vector through the preset visual language model according to the preset multimodal fusion formula, the dynamic visual feature vector weight and the dynamic text feature vector weight to obtain a second multimodal feature vector, including; The second text feature vector is connected with the second visual feature vector according to a preset multimodal fusion formula, a dynamic visual feature vector weight, and a dynamic text feature vector weight through a preset visual language macro model to obtain a third multimodal feature vector; wherein the preset multimodal fusion formula is: M = α·V2+β·T2; Wherein, M is the second multimodal feature vector; α is the dynamic visual feature vector weight; V2 is the second visual feature vector; β is the dynamic text feature vector weight; T1 is the second text feature vector; Accordingly, the second multimodal feature vector is processed by the language model of the preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to the control center, including: The third multimodal feature vector is processed by a language model of a preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and the abnormal behavior judgment information is sent to a control center.

7. The method according to claim 6, characterized in that in, The preset text weight adjustment formula is: T2=T1·W(S) Among them, T2 is the second text feature vector; T1 is the first text feature vector; S is the scene semantic label; W(S) is the scene weight.

8. The method according to claim 1, characterized in that: in, After sending the abnormal behavior determination information to the control center, the method further includes: If the judgment error information sent by the control center is received, the actual behavior information is determined according to the judgment error information, and the preset visual language large model is adjusted according to the actual behavior information.

9. A method system for detecting abnormal behavior of personnel based on a large visual language model, used to execute the method according to any one of claims 1 to 8, characterized in that: The system comprises: A monitoring video frame acquisition module is used to continuously acquire a monitoring video stream and continuously acquire monitoring video frames from the monitoring video stream according to a preset fixed time interval; A text feature vector determination module is used to obtain the current system time when a text instruction sent by the control center is received, input the text instruction into a preset visual language model, process the text instruction through a text lexicalizer of the preset visual language model, and obtain a first text feature vector corresponding to the text instruction; A visual feature vector determination module is used to determine a monitoring video data set according to the current system time, a preset video data set combination mode and the monitoring video frame, input the monitoring video data set into a preset visual language large model, and process the monitoring video data set through a visual encoder and a feature adapter of the preset visual language large model to obtain a first visual feature vector corresponding to the monitoring video data set; The abnormal behavior judgment information determination module is used to connect the first text feature vector with the first visual feature vector through a preset visual language large model to obtain a first multimodal feature vector, and process the first multimodal feature vector through a language model of the preset visual language large model to obtain abnormal behavior judgment information corresponding to the text instruction, and send the abnormal behavior judgment information to the control center.

10. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8.

Citation Information

Cited By

  • Visual language model adaptation method and device, computer equipment and storage medium

    CN120671768A

  • Abnormal behavior detection method and device and electronic equipment

    CN120673303A

  • Abnormal behavior detection methods, devices and electronic equipment

    CN120673303B

  • Behavior recognition method based on computer vision

    CN120913278A

  • Computer vision based behavior recognition method

    CN120913278B