Multi-mode large-model campus security video inspection early warning method and multi-mode large-model campus security video inspection early warning system
Through the combination of multimodal large model and thinking chain technology, training samples are generated and cross-modal feature extraction and feedback optimization are solved, and the time-consuming and labor-intensive problem of traditional campus security patrol is realized, and intelligent and automated hierarchical early warning and abnormal behavior recognition are realized.
Patent Information
- Application Number
- CN202510435860.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Traditional campus security inspection methods are time-consuming and labor-intensive, and it is difficult to detect safety hazards in a timely manner. It is difficult for existing video monitoring technologies to achieve accurate and generalized abnormal behavior identification and hierarchical warning.
A multimodal large model is used to combine training sample data sets and thinking chain technology to generate training samples through video information processing, cross-modal feature extraction and feedback iterative optimization, and generate thinking chain prompts to achieve hierarchical early warning of real-time video data.
It has realized intelligent, automated patrol and hierarchical warning of campus security, and has more accurate abnormal behavior recognition capabilities, which has improved the intelligence level and early warning accuracy of campus security.
Smart Images

Figure CN120260245A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of campus security technology, and in particular, to a campus security video patrol warning method and system based on a multimodal large model. Background Art
[0002] Campus security patrol is of great significance for timely discovering potential safety hazards and preventing safety accidents. The traditional method of relying on security personnel to conduct on-site patrols is not only time-consuming and laborious, but also many potential safety hazards are difficult to discover in a timely manner, and real-time security monitoring cannot be achieved. With the development of video surveillance technology, patrols can be carried out in the monitoring room with the help of the above-mentioned videos. In particular, various intelligent video monitoring and warning technologies provide powerful means for the monitoring and warning of abnormal behaviors through feature extraction and computational analysis. However, campus security hazards are closely related to the dynamic changes of specific places, sections, internal and external environments, people, and weather. The above-mentioned technologies are difficult to comprehensively extract relevant features and perform accurate calculations, analyses, and judgments with generalization ability on the behavior risks in various scenarios.
[0003] Therefore, how to achieve intelligent, automated patrol and hierarchical warning of campus security and accurately identify abnormal behaviors has become an urgent technical problem to be solved. Summary of the Invention
[0004] In order to achieve intelligent, automated patrol and hierarchical warning of campus security and accurately identify abnormal behaviors, this application provides a campus security video patrol warning method and system based on a multimodal large model.
[0005] In the first aspect, this application provides a campus security video patrol warning method based on a multimodal large model, including:
[0006] S1. Collect video information corresponding to the security patrol object through a preset sensing device in the target area;
[0007] S2. Obtain preset processing conditions and process the video information to generate a training sample data set, including:
[0008] S2.1 Label the video according to the place number, monitoring device number, timestamp, weather, and behavior type;
[0009] S2.2 Cut each video into multiple video segments at a fixed time interval and generate a structured data set including sample serial numbers, scene descriptions, and behavior type labels;
[0010] S3. Use a preset multimodal large model to perform tuning training in combination with the training sample data set and the thought chain technology, including:
[0011] S3.1 Multimodal data fusion: Taking video clips, scene description texts, and behavior type labels as inputs, cross-modal feature extraction is performed through the feature alignment layer of the multimodal large model to generate joint embedding vectors;
[0012] S3.2 Thought chain prompt generation: Based on the scene description and the experience of security personnel, a thought chain prompt template containing abnormal behavior grading rules is constructed;
[0013] S3.3 Feedback iterative optimization: Comparing the potential security risk behavior answers output by the multimodal large model with the thought chain prompts. If the answers do not cover the actual risks, the prompt template is corrected and training samples are supplemented;
[0014] S3.4 Tuning training: Adopting a contrastive learning strategy to minimize the difference between the model output and the true labels, and at the same time maximizing the response consistency of the model to the thought chain prompts;
[0015] S4. Set the patrol mode based on the tuned preset multimodal large model, analyze the real-time video data, and output early warnings according to the grading of the possibility of abnormal behaviors.
[0016] Optionally, the video information includes: normal videos and abnormal videos;
[0017] The normal videos include: no less than 30 consecutive 1-minute videos,
[0018] The abnormal videos include: no less than 6 consecutive 1-minute videos for each type of security risk behavior.
[0019] Optionally, the step of obtaining preset processing conditions and processing the video information to generate a training sample dataset includes:
[0020] Obtain preset processing conditions, and process the video information according to the cutting requirements in the preset processing conditions to generate a training sample dataset in the format of sample serial number SS, place number SN, monitoring device number MN, video serial number VN, segment serial number CN, video segment CV, date DE, time TM, season ST, weather WT, behavior type BT, and scene description SD.
[0021] Optionally, the thought chain prompts are compiled based on the scene description and the experience of security personnel, including the grading warning rules for abnormal behaviors, and integrating the potential security risk behavior answers fed back by the large model to enhance the generalization ability.
[0022] Optionally, it further includes a test dataset, which consists of the remaining one-fourth of the normal video samples and abnormal video samples. If it is determined that the test results do not meet the expectations, the thought chain prompts are adjusted or sample data is supplemented and then retrained.
[0023] Optionally, the step of grading and outputting early warnings according to the likelihood of abnormal behavior includes: the first to third level early warnings starting from the second video segment respectively correspond to the determination results with gradually increasing likelihood of abnormal behavior.
[0024] Optionally, the patrol mode includes all-round cyclic patrol, key area cyclic patrol, and real-time monitoring mode for a single area, and switches according to the received security scenarios.
[0025] In a second aspect, the present application provides a campus security video patrol early warning system based on a multimodal large model, which executes the method as described above, including:
[0026] An information acquisition module, configured to collect video information corresponding to security patrol objects in a target area through a preset sensing device;
[0027] A training module, configured to obtain preset processing conditions and process the video information to generate a training sample data set;
[0028] An optimization module, configured to perform optimization training by using a preset multimodal large model in combination with the training sample data set and the thought chain technology, generate thought chain prompts through scene descriptions and the experience of security personnel, and supplement the answers of security risk behaviors fed back by the preset multimodal large model;
[0029] An output module, configured to set a patrol mode based on the optimized preset multimodal large model, analyze real-time video data, and output early warnings by grading according to the likelihood of abnormal behavior.
[0030] In a third aspect, the present application provides a computer device, which includes: a memory and a processor, and when the processor runs the computer instructions stored in the memory, it executes the method as described above.
[0031] In a fourth aspect, the present application provides a computer-readable storage medium, including instructions, and when the instructions run on a computer, the computer is enabled to execute the method as described above.
[0032] In summary, the present application includes the following beneficial technical effects:
[0033] This application collects video information corresponding to security inspection objects within a target area through a preset sensing device; obtains preset processing conditions, and processes the video information to generate a training sample data set; uses a preset multimodal large model to perform tuning training in combination with the training sample data set and the chain-of-thought technique, generates chain-of-thought prompts through scene descriptions and the experience of security personnel, and supplements the answers of security risk behaviors fed back by the preset multimodal large model; sets a patrol mode based on the tuned preset multimodal large model, analyzes real-time video data, and outputs warnings according to the classification of the likelihood of abnormal behaviors. It realizes intelligent, automated patrol and hierarchical warning for campus security, and has a more accurate generalization and recognition ability for abnormal behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic structural diagram of a computer device in the hardware operating environment involved in the solution of the embodiment of this application;
[0035] Figure 2 is a flowchart of the first embodiment of the method for campus security video patrol warning of the multimodal large model of this application;
[0036] Figure 3 is a schematic flowchart of the first embodiment of the campus security video patrol warning system of the multimodal large model of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further details this application through the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0038] Refer to Figure 1 , Figure 1 is a schematic structural diagram of a computer device in the hardware operating environment involved in the solution of the embodiment of this application.
[0039] As Figure 1As shown in the figure, the computer device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM), or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0040] Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0041] As Figure 1 shown, in the memory 1005 as a storage medium, there may be included an operating system, a network communication module, a user interface module, and a campus security video patrol warning program of a multi-modal large model.
[0042] In Figure 1 the computer device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; in this application, the processor 1001 and the memory 1005 may be arranged in the computer device. The computer device calls the campus security video patrol warning program of the multi-modal large model stored in the memory 1005 through the processor 1001, and executes the campus security video patrol warning method provided by the embodiments of this application.
[0043] The embodiments of this application provide a campus security video patrol warning system of a multi-modal large model. Referring to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the campus security video patrol warning method of the multi-modal large model of this application.
[0044] S1: Collect video information corresponding to the security patrol object through a preset sensing device in the target area.
[0045] It should be noted that the video information includes: normal videos and abnormal videos; the normal videos include: no less than 30 consecutive 1-minute videos, and the abnormal videos include: no less than 6 consecutive 1-minute videos for each type of security risk behavior.
[0046] In specific implementation, video monitoring devices are set at fixed positions in campus locations that require security patrols, and the locations and monitoring devices are numbered. The above locations generally refer to objects that require security patrols; according to the location numbers and monitoring device numbers, normal video recordings and abnormal video recordings of security risk behaviors that have occurred are selected from their historical video materials, and the location numbers, monitoring device numbers, and the above video recordings, together with the date, time, season, weather, behavior type, and scene description information of each video recording, are stored in the video case library and marked and explained; among them, the behavior type of the normal video is "normal", and the behavior type of the abnormal video should give the specific classification of its security risk behavior. The scene description includes the description of people, facilities, items, and activity status in the normal video and the description of the occurrence process of the security risk behavior in the abnormal video; the normal videos collected by each monitoring device should include no less than 30 representative consecutive 1-minute videos, and the abnormal videos it collects should include no less than 6 consecutive 1-minute videos from before the behavior occurs to when it occurs for each type of security risk behavior; if the historical video recordings are insufficient, the method of presenting real crowd drills can be adopted to make up for it.
[0047] It should be noted that for target area coverage: cameras and sensors are deployed in key locations on campus (such as teaching building entrances and exits, playgrounds, dormitory corridors, etc.) to ensure that video collection covers all preset security patrol objects. For normal videos: no less than 30 consecutive 1-minute videos are collected, which need to cover different time periods (day / night), seasons (spring, summer, autumn, winter), weather (sunny / rainy / snowy), and crowded scenarios. For abnormal videos: for each type of security risk behavior (such as people gathering, breaking into restricted areas, items being left behind, etc.), no less than 6 consecutive 1-minute videos are collected to ensure the integrity of behavior characteristics. For data annotation: the videos need to be marked with metadata such as location number (SN), monitoring device number (MN), timestamp (TM), weather (WT), etc.
[0048] S2: Obtain preset processing conditions and process the video information to generate a training sample data set.
[0049] It should be noted that obtaining preset processing conditions and processing the video information to generate a training sample data set includes:
[0050] S2.1 Annotate the videos according to the location number, monitoring device number, timestamp, weather, and behavior type;
[0051] S2.2 Cut each video into multiple video segments at fixed time intervals, and generate a structured dataset containing sample numbers, scene descriptions, and behavior type tags.
[0052] It can be understood that the step of obtaining the preset processing conditions and processing the video information to generate the training sample dataset includes: obtaining the preset processing conditions, and processing the video information according to the cutting requirements in the preset processing conditions to generate the training sample dataset in the format of sample number SS, venue number SN, monitoring device number MN, video number VN, segment number CN, video segment CV, date DE, time TM, season ST, weather WT, behavior type BT, and scene description SD.
[0053] In a specific implementation, each video collected in step S1 is cut into 4 video segments in units of 15 seconds, and the training sample dataset is formed according to the format for the fine-tuning training of the large model.
[0054] S3: Use the preset multi-modal large model to perform fine-tuning training in combination with the training sample dataset and the chain of thought technology, generate chain of thought prompts through scene descriptions and the experience of security personnel, and supplement the answers to security risk behaviors fed back by the preset multi-modal large model.
[0055] It should be noted that using the preset multi-modal large model to perform fine-tuning training in combination with the training sample dataset and the chain of thought technology includes:
[0056] S3.1 Multi-modal data fusion: Take the video segment, scene description text, and behavior type tag as inputs, and perform cross-modal feature extraction through the feature alignment layer of the multi-modal large model to generate a joint embedding vector;
[0057] S3.2 Chain of thought prompt generation: Based on the scene description and the experience of security personnel, construct a chain of thought prompt template containing abnormal behavior grading rules;
[0058] S3.3 Feedback iteration optimization: Compare the potential security risk behavior answers output by the multi-modal large model with the chain of thought prompts. If the answers do not cover the actual risks, correct the prompt template and supplement the training samples;
[0059] S3.4 Fine-tuning training: Adopt a contrastive learning strategy to minimize the difference between the model output and the true label, and at the same time maximize the response consistency of the model to the chain of thought prompts;
[0060] S4. Set the patrol mode based on the fine-tuned preset multi-modal large model, analyze the real-time video data, and output early warnings according to the abnormal behavior possibility grading.
[0061] In a specific implementation,
[0062] Multi-modal Data Fusion (Step S3.1)
[0063] Technical details: Use the visual encoder of a multi-modal large model (such as VisualGLM-6B) to extract video frame features, and the text encoder to process scene descriptions, and achieve feature alignment through a cross-modal attention mechanism.
[0064] Example: In the DroneRFa dataset experiment, visual features are extracted from video clips by ResNet-50, and scene description texts are encoded by BERT. Both are mapped to the same embedding space through a fully connected layer to form a joint feature vector.
[0065] Dynamic Adjustment of Chain-of-Thought Prompting (Steps S3.2 - S3.3)
[0066] Technical details: Chain-of-thought prompting is divided into static rules (preset hierarchical warnings) and dynamic rules (model feedback supplementation).
[0067] Example: The initial chain-of-thought prompt only contains "crowd gathering → level 1 warning". The model feedback identifies that "gathering + carrying knives → level 3 warning", and this rule is supplemented to the prompt template and retrained.
[0068] Optimized Training Strategy (Step S3.4)
[0069] Technical details: Adopt two-stage training:
[0070] Pre-training stage: Fine-tune the model weights on a large-scale general multi-modal dataset (such as LAION-5B);
[0071] Optimization stage: Use a campus security special dataset, and conduct supervised learning in combination with chain-of-thought prompting. The loss function is cross-entropy loss + prompt consistency loss.
[0072] Example: In the optimization stage, the learning rate is set to 3e-5, the batch size is 16, and the training cycle is 50 rounds. The F1 score of the final model on the test set is improved from 0.72 to 0.89.
[0073] Feedback Mechanism and Iterative Optimization (Step S3.3)
[0074] Technical details: If the recognition accuracy of a certain type of abnormal behavior in the test set is lower than the threshold (such as <80%), then:
[0075] Analyze the misjudgment cases of the model and supplement the training samples for this type of behavior;
[0076] Adjust the logical level of the chain-of-thought prompt (such as upgrading "breaking into the night restricted area" from level 2 to level 3 warning).
[0077] Example: For the behavior of "climbing over the wall", the initial accuracy rate was 68%. After supplementing 20 new samples and optimizing the prompts, the accuracy rate increased to 92%.
[0078] It should be noted that in this embodiment, time-series video clips that can reflect the risk evolution characteristics of various places and specific visual field ranges, as well as their time, season, and weather information, are used to perform prompt tuning training on the VisualGLM-6B large model. According to the description of the occurrence process of security risk behaviors, the experience of security personnel, and the answers of other possible security risk behaviors given by the large model, a chain of thought prompt for tuning training is compiled, providing a new technical method for campus security that is intelligent, automated, and has stronger accuracy and generalization capabilities.
[0079] In the specific implementation, the chain of thought prompt is compiled based on the scenario description and the experience of security personnel, including the classification and early warning rules for abnormal behaviors, and integrates the answers of potential security risk behaviors fed back by the large model to enhance the generalization ability.
[0080] It can be understood that the test data set consists of the remaining one-fourth of normal video samples and abnormal video samples. If it is determined that the test result does not meet the expectations, the chain of thought prompt is adjusted or the sample data is supplemented and then retrained.
[0081] It should be noted that the VisualGLM-6B multi-modal large model is adopted and the chain of thought technology is introduced to generate prompt tuning instructions according to the training samples made in step S2, and the large model is tuned and trained; according to the place number SN and the monitoring device number MN, three-fourths of the normal video samples and three-fourths of the abnormal video samples are randomly selected from the video serial numbers VN belonging to the above numbers, and a training data set is formed according to the order of their segment numbers CN, and the remaining video samples are used as the test data set; after the video samples of the monitoring devices in each place are trained, the large model is asked what other security risk behaviors may occur in this place, and the answers given by the large model are incorporated into the chain of thought prompt; after all the training data sets are trained, the test data set is used to test the large model. When the recognition accuracy reaches the expected index requirements, step S4 is entered, otherwise the chain of thought is adjusted first. If the requirements are still not met, step S1 is entered to supplement the videos and related information of the corresponding place number SN and monitoring device number MN until the expected index requirements are met.
[0082] S4: Set the patrol mode based on the tuned preset multi-modal large model, analyze the real-time video data, and output early warnings according to the classification of the possibility of abnormal behaviors.
[0083] It should be noted that the step of outputting early warnings according to the classification of the possibility of abnormal behaviors includes: the first to third level early warnings starting from the second video segment, corresponding to the determination results with gradually increasing possibilities of abnormal behaviors respectively.
[0084] In specific implementations, all round-robin inspections are for daily full-area monitoring, and the method is to poll all cameras at fixed intervals; the round-robin inspections of key locations are for high-risk periods (such as at night), and the method is to preferentially scan areas such as dormitories and laboratories; the real-time monitoring of a single location is for emergencies, and the specific method is to track and lock a specific camera for continuous analysis.
[0085] It can be understood that the inspection mode includes all round-robin inspections, round-robin inspections of key locations, and real-time monitoring of a single location, and is switched according to the received security scenario requirements.
[0086] It should be noted that this embodiment provides a campus security video inspection and early warning method based on a multimodal large model. Through steps such as video information collection, data processing, model tuning and training, and real-time inspection and early warning, it realizes the intelligent detection and hierarchical early warning of abnormal behaviors on campus. This method combines the multimodal large model and the thought chain technology, and can effectively improve the intelligent level and early warning accuracy of campus security.
[0087] During the process of video information collection, it includes:
[0088] Deploy preset sensing devices (such as high-definition cameras, infrared sensors, etc.) in the target area (such as key areas on campus such as school gates, teaching buildings, dormitory areas, playgrounds, etc.), and collect video information corresponding to the security inspection objects in real time.
[0089] The specific steps are as follows:
[0090] Device deployment: Install high-definition cameras and infrared sensors in key areas of the campus to ensure that the coverage area has no dead ends.
[0091] Video stream acquisition: Obtain the video stream in real time through the sensing device, and the resolution and frame rate of the video stream are set according to actual needs (such as 1080p resolution, 30 frames per second).
[0092] Data transmission: Transmit the collected video information to the data processing module through the network to ensure the real-time and integrity of the data.
[0093] During the process of data processing and training sample generation, it includes:
[0094] Obtain preset processing conditions, and process the collected video information to generate a training sample data set.
[0095] The specific steps are as follows:
[0096] Video frame extraction: Extract key frames from the video stream at fixed time intervals (such as 1 frame per second).
[0097] Data annotation: Manually or semi-automatically annotate the extracted key frames, and the annotation content includes:
[0098] Normal behaviors (such as walking, standing, talking, etc.).
[0099] Abnormal behaviors (such as running, pushing, climbing over the wall, etc.).
[0100] Data augmentation: Perform augmentation on the labeled data, including operations such as rotation, scaling, adding noise, etc., to improve the generalization ability of the model.
[0101] Dataset division: Divide the processed data into training set, validation set and test set, usually in the ratio of 7:2:1.
[0102] During the tuning training of the multi-modal large model, it includes:
[0103] Use a preset multi-modal large model (such as multi-modal models like CLIP, DALL-E, etc.) combined with the generated training sample dataset and the chain of thought technique for tuning training. The specific steps are as follows:
[0104] Model initialization: Load the weights of the pre-trained multi-modal large model as the base model.
[0105] Chain of thought prompt generation: Combine the scene description and the experience of security personnel to generate chain of thought prompts. For example:
[0106] Scene description: Students are doing physical activities on the playground.
[0107] Chain of thought prompt: If it is found that a student is running on the playground, it may be normal physical activity; but if it is found that a student is pushing on the playground, it may be a fight.
[0108] Model training: Use the generated training sample dataset and chain of thought prompts to tune and train the model. During the training process, the model will generate answers for security risk behaviors according to the chain of thought prompts and perform iterative optimization based on the feedback.
[0109] Model evaluation: Use the validation set and test set to evaluate the tuned model to ensure the accuracy and robustness of the model. Evaluation metrics include accuracy, recall rate and F1 score.
[0110] During the real-time patrol and hierarchical warning process, it includes:
[0111] Set the patrol mode based on the tuned preset multi-modal large model, analyze the real-time video data, and output warnings according to the likelihood of abnormal behaviors at different levels. The specific steps are as follows:
[0112] Real-time video analysis: Input the real-time video stream into the tuned multi-modal large model, and the model will analyze the behaviors in the video.
[0113] Abnormal behavior detection: The model detects abnormal behaviors in videos based on the knowledge learned during the training process. Abnormal behaviors include, but are not limited to:
[0114] Low-risk behaviors: For example, students running during non-physical education classes;
[0115] Medium-risk behaviors: For example, students shoving on the playground;
[0116] High-risk behaviors: For example, outsiders breaking into restricted areas.
[0117] Graded warning output: Output warnings according to the likelihood of abnormal behaviors. Warning levels include:
[0118] Low-risk warning: Notify campus security personnel to pay attention and observe;
[0119] Medium-risk warning: Notify campus security personnel to go to the scene for handling;
[0120] High-risk warning: Immediately notify campus security personnel and activate the emergency plan.
[0121] It should be noted that the campus security video patrol warning method in this embodiment can be applied to the following scenarios:
[0122] Daily patrol: Conduct daily video patrols on campus to promptly detect and handle abnormal behaviors.
[0123] Emergency handling: During emergencies (such as fires, earthquakes, etc.), quickly locate dangerous areas through video patrols and guide personnel evacuation.
[0124] Personnel management: Conduct real-time monitoring of outsiders to prevent unauthorized personnel from entering the campus.
[0125] In this embodiment, video information corresponding to security patrol objects is collected through preset sensing devices within the target area; preset processing conditions are obtained, and the video information is processed to generate a training sample dataset; a preset multimodal large model is used to perform tuning training in combination with the training sample dataset and the thought chain technology, generate thought chain prompts through scenario descriptions and the experience of security personnel, and supplement the answers of security risk behaviors fed back by the preset multimodal large model; based on the tuned preset multimodal large model, a patrol mode is set, the real-time video data is analyzed, and warnings are output according to the likelihood of abnormal behaviors. It realizes intelligent and automated patrols and graded warnings for campus security, and has a more accurate generalization and recognition ability for abnormal behaviors.
[0126] In addition, an embodiment of the present application also proposes a computer-readable storage medium, on which a program for campus security video patrol warning of a multi-modal large model is stored. When the program for campus security video patrol warning of the multi-modal large model is executed by a processor, the steps of the method for campus security video patrol warning of the multi-modal large model as described above are implemented.
[0127] Referring Figure 3 , Figure 3 is a system block diagram of the first embodiment of the campus security video patrol warning system of the multi-modal large model of the present application.
[0128] As Figure 3 shown, the campus security video patrol warning system proposed by the embodiment of the present application includes.
[0129] An information acquisition module 10, configured to collect video information corresponding to security patrol objects through a preset sensing device within a target area;
[0130] A training module 20, configured to obtain preset processing conditions and process the video information to generate a training sample data set;
[0131] An optimization module 30, configured to perform optimization training by using a preset multi-modal large model in combination with the training sample data set and the thought chain technology, generate thought chain prompts through scene descriptions and security personnel experience, and supplement the answers of security risk behaviors fed back by the preset multi-modal large model;
[0132] An output module 40, configured to set a patrol mode based on the optimized preset multi-modal large model, analyze real-time video data, and output a warning by grading according to the possibility of abnormal behaviors.
[0133] By collecting video information corresponding to security patrol objects through a preset sensing device within a target area; obtaining preset processing conditions and processing the video information to generate a training sample data set; performing optimization training by using a preset multi-modal large model in combination with the training sample data set and the thought chain technology, generating thought chain prompts through scene descriptions and security personnel experience, and supplementing the answers of security risk behaviors fed back by the preset multi-modal large model; setting a patrol mode based on the optimized preset multi-modal large model, analyzing real-time video data, and outputting a warning by grading according to the possibility of abnormal behaviors. The intelligent, automated patrol and graded warning of campus security are realized, and the multi-modal large model has a more accurate generalization and recognition ability for abnormal behaviors.
[0134] It should be understood that the above is only for illustration and does not constitute any limitation to the technical solution of the present application. In specific applications, those skilled in the art can set according to needs, and the present application does not make any restrictions on this.
[0135] It should be noted that the workflow described above is only illustrative and does not limit the protection scope of this application. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is made here.
[0136] In addition, for the technical details not described in detail in this embodiment, reference can be made to the method for campus security video patrol warning of the multimodal large model provided in any embodiment of this application, and details will not be repeated here.
[0137] In addition, it should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0138] The serial numbers of the embodiments of this application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as Read Only Memory (ROM) / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0140] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A campus security video patrol and warning method for a multi-modal large model, characterized in that, Including: S1. Collect video information corresponding to the security inspection object through a preset sensing device within the target area; S2. Obtain preset processing conditions and process the video information to generate a training sample dataset, including: S2.1 Annotate the video according to the venue number, monitoring device number, timestamp, weather, and behavior type; S2.2 Cut each video into multiple video segments at fixed time intervals and generate a structured dataset containing sample serial numbers, scene descriptions, and behavior type labels; S3. Use a preset multi-modal large model to perform tuning training in combination with the training sample dataset and the thought chain technology, including: S3.1 Multi-modal data fusion: Take the video segment, scene description text, and behavior type label as inputs, and perform cross-modal feature extraction through the feature alignment layer of the multi-modal large model to generate a joint embedding vector; S3.2 Thought chain prompt generation: Based on the scene description and the experience of security personnel, construct a thought chain prompt template containing abnormal behavior grading rules; S3.3 Feedback iteration optimization: Compare the potential security risk behavior answers output by the multi-modal large model with the thought chain prompt. If the answer does not cover the actual risk, correct the prompt template and supplement the training samples; S3.4 Tuning training: Adopt a contrastive learning strategy to minimize the difference between the model output and the true label, and at the same time maximize the response consistency of the model to the thought chain prompt; S4. Set up an inspection mode based on the tuned preset multi-modal large model, analyze the real-time video data, and output early warnings according to the grading of the probability of abnormal behavior.
2. The campus security video patrol warning method for the multi-modal large model according to claim 1, characterized in that, The video information includes: normal videos and abnormal videos; The normal videos include: no less than 30 consecutive 1-minute videos, The abnormal videos include: no less than 6 consecutive 1-minute videos for each type of security risk behavior.
3. The campus security video patrol warning method for the multi-modal large model according to claim 1, characterized in that, The step of obtaining preset processing conditions and processing the video information to generate a training sample dataset includes: Obtain preset processing conditions, and process the video information according to the cutting requirements in the preset processing conditions to generate a training sample dataset in the format of sample serial number SS, venue number SN, monitoring device number MN, video serial number VN, segment serial number CN, video segment CV, date DE, time TM, season ST, weather WT, behavior type BT, and scene description SD.
4. The campus security video patrol warning method for the multi-modal large model according to claim 1, wherein, The thought chain prompt is compiled based on the scene description and the experience of security personnel, includes grading warning rules for abnormal behaviors, and integrates the potential security risk behavior answers fed back by the large model to enhance the generalization ability.
5. The campus security video patrol and early warning method for the multi-modal large model according to claim 1, characterized in that, It also includes a test dataset, which consists of the remaining one-quarter of normal video samples and abnormal video samples. If it is determined that the test result does not meet the expectation, adjust the thought chain prompt or supplement the sample data and then retrain.
6. The campus security video patrol warning method for the multi-modal large model according to claim 1, characterized in that, The step of outputting early warnings according to the grading of the probability of abnormal behavior includes: level 1 to 3 early warnings starting from the second video segment, corresponding to the determination results with gradually increasing probabilities of abnormal behavior.
7. The campus security video patrol and early warning method for the multi-modal large model according to claim 1, characterized in that, The inspection mode includes all-round cycle inspection, key venue cycle inspection, and single venue real-time monitoring mode, and switches according to the received security scenarios.
8. A campus security video patrol and early warning system for a multi-modal large model, characterized in that, Implement the method according to claim 1, including: An information acquisition module for collecting video information corresponding to security inspection objects within a target area through a preset sensing device; A training module for obtaining preset processing conditions and processing the video information to generate a training sample data set; An optimization module for performing optimization training by using a preset multi-modal large model in combination with the training sample data set and the thought chain technology, generating thought chain prompts through scene descriptions and the experience of security personnel, and supplementing the answers of security risk behaviors fed back by the preset multi-modal large model; An output module for setting an inspection mode based on the optimized preset multi-modal large model, analyzing real-time video data, and outputting warnings according to the likelihood of abnormal behaviors being graded.
9. A computer device, characterized in that, The device includes: a memory and a processor, and when the processor runs the computer instructions stored in the memory, it executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Including instructions that, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Construction site personnel exception aggregation detection system based on block chain and BIM
CN111523434A
Multi-modal generative large model training method and device and computer equipment
CN117011686A
Driver behavior detection method and system integrating visual large language model and inference chain
CN118486001A
Infrared thermal imaging abnormal scene monitoring method based on multi-modal large model and related device
CN119723464A
Industrial image anomaly detection method based on multi-modal large model
CN119762891A
Cited By
Multi-mode campus monitoring method and system based on target preprocessing
CN121259697A