A method for constructing a large multimodal model for security risk identification
By constructing a multimodal campus safety dataset and combining it with a generative Transformer model, and adopting a hierarchical weight exponential decay and node acceptance rate calculation mechanism, the problem of insufficient automation and intelligence of multimodal technology in campus safety is solved, and the accurate identification and efficient positioning of safety risks are achieved, thereby improving the accuracy of risk detection and management efficiency.
Patent Information
- Application Number
- CN202510985697.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing multimodal technologies lack the level of automation and intelligence in campus security application scenarios, making it difficult to comprehensively and efficiently detect risks. In addition, the security management and assessment system is not perfect and cannot effectively improve security risk management capabilities.
A multimodal campus safety dataset consisting of surveillance videos, environmental sensors, and facility maintenance data is constructed. The generative Transformer model is fine-tuned, and the hierarchical weight exponential decay and node acceptance rate calculation mechanisms are adopted to achieve accurate identification and efficient positioning of security risks through the graph attention network.
It improves the accuracy of risk detection and the efficiency of model reasoning, realizes the automated and intelligent processing of campus safety risks, can accurately judge the risk level and type, and provide efficient and accurate support for campus safety management.
Smart Images

Figure CN120493188B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method for constructing a multimodal large model for security risk identification. Background Art
[0002] Deep learning multimodal technology continues to develop in fields such as fire safety and smart emergency management. While significant progress has been made, enabling efficient recognition and understanding of diverse data types, practical applications face numerous challenges. For one thing, the model's generalization performance is poor under unrestricted conditions, limiting its effective application in real-world scenarios. Furthermore, numerous issues exist in the actual design and application of large multimodal models, such as insufficient interaction in the dual-tower network structure and a low upper limit on the matching capability of modal fusion, which makes it incapable of fine-grained cross-modal matching tasks. While BLIP combines two training modes, vision and natural language processing require separate encoding, making it an unsuitable multimodal unified model framework. While feature alignment between vision and LLM enables LLM to possess multimodal understanding capabilities, the training data format lags behind mainstream technical solutions.
[0003] Campus safety has received increasing attention from the society in recent years. Campus safety accidents are characterized by diversity and complexity, so the investigation and rectification of safety risks have brought considerable challenges to the work of managers. In order to properly solve these problems in campus safety management, it is urgently necessary to use advanced multimodal technology to achieve efficient and intelligent safety risk monitoring and management as well as scientific and comprehensive safety management level assessment. However, the current existing multimodal technology is still insufficient in campus safety application scenarios. Some technologies lack automation and intelligence when conducting safety risk monitoring and management for various facilities and venues on campus, making it difficult to comprehensively and efficiently investigate risks. At the same time, the existing campus safety management evaluation system indicators are not perfect, and it is impossible to comprehensively evaluate the campus safety management level, which makes it difficult to help schools effectively improve their safety risk management capabilities, and is not conducive to the formulation of scientific risk financial planning by campus safety insurance business units. Summary of the Invention
[0004] An embodiment of the present application provides a method for constructing a multimodal large model for security risk identification. By constructing a campus security multimodal dataset containing multi-source data such as surveillance videos and environmental sensors, and fine-tuning it in combination with a generative Transformer model, a security risk identification model is obtained. The hierarchical weight exponential decay and node acceptance rate calculation mechanism are used in the graph attention network to achieve accurate identification and efficient positioning of campus security risks, effectively improving the accuracy of risk detection and the efficiency of model reasoning.
[0005] In a first aspect, an embodiment of the present application provides a method for constructing a multimodal large model for security risk identification, the method comprising:
[0006] Constructing a multimodal campus safety dataset, wherein the multimodal campus safety dataset includes surveillance video data, environmental sensor data, facility maintenance data, and safety feedback data on campus over multiple historical time periods;
[0007] Mapping the surveillance image of the surveillance video data into an image feature sequence, and converting the environmental sensor data, facility maintenance data, and safety feedback data corresponding to the current surveillance video data into a text feature sequence;
[0008] Perform multimodal perceptual streaming processing and spatial semantic alignment on the image feature sequence and the text feature sequence to obtain the same set of multimodal input sequences;
[0009] The decoder of the generative transformer model is fine-tuned using a multimodal input sequence to obtain a security risk identification model. The security risk identification model generates a graph attention network based on the multimodal input sequence, assigns a layer weight to each layer of the graph attention network, and the layer weight decays exponentially with the increase of the layer depth. The product of the confidence of each node on the path and the corresponding layer weight is calculated as the node acceptance rate, and the product of the node acceptance rates of all nodes on the path is used as the global acceptance rate of the path. In the graph attention network, the generated information corresponding to the path with the highest global acceptance rate is selected as the security risk.
[0010] In a second aspect, an embodiment of the present application provides a security risk identification method based on a multimodal large model, comprising:
[0011] Obtain surveillance video data to be identified;
[0012] Inputting the surveillance video data to be identified and the text prompt into the multimodal large model for security risk identification constructed by the multimodal large model construction method for security risk identification constructed in the first aspect, and outputting the security risk;
[0013] The surveillance video data is mapped into an image feature sequence, the text prompt is mapped into a text feature sequence, and the image feature sequence and the text feature sequence are spliced to be input into a multimodal large model for security risk identification.
[0014] The main contributions and innovations of the present invention are as follows:
[0015] The embodiment of the present application integrates data from multiple modalities such as video surveillance data, environmental sensor data, facility maintenance records, and teacher-student safety feedback information to achieve fine-tuning of the generative transformer generative model, so that the fine-tuned safety risk identification model can analyze from more comprehensive information to improve the accuracy of risk assessment and risk identification; the embodiment of the present application constructs a multimodal large model specifically for campus safety scenarios based on the Transformer architecture. The model can effectively learn the complex associations and complementary information between different modal data and generate a graph attention network, and conduct hierarchical and structured generation path exploration based on the graph attention network, combining the semantic association and dynamic weight evaluation of multimodal data to accurately obtain campus safety risks; the embodiment of the present application uses a multimodal large model for risk assessment and risk identification, realizing automated and intelligent processing from data to risk results. The model can comprehensively consider multiple factors, accurately judge the risk level and type, locate safety risks, and provide efficient and accurate support for campus safety management.
[0016] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 is a flowchart of a method for constructing a multimodal large model for security risk identification according to an embodiment of the present application;
[0019] Figure 2 is a flow chart for fine-tuning a multimodal large model according to an embodiment of the present application;
[0020] Figure 3 This is a flow chart of a method for realizing security risk identification by constructing a front-end page according to an embodiment of the present application;
[0021] Figure 4 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.
[0023] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0024] Example 1
[0025] The embodiment of the present application provides a method for constructing a multimodal large model for security risk identification. By constructing a campus security multimodal dataset containing multi-source data such as surveillance videos and environmental sensors, and fine-tuning it in combination with a generative Transformer model, a security risk identification model is obtained. In addition, a hierarchical weight exponential decay and node acceptance rate calculation mechanism are used in the graph attention network to achieve accurate identification and efficient positioning of campus security risks, effectively improving the accuracy of risk detection and model reasoning efficiency. Specifically, reference Figure 1 , the method comprising:
[0026] Constructing a multimodal campus safety dataset, wherein the multimodal campus safety dataset includes surveillance video data, environmental sensor data, facility maintenance data, and safety feedback data on campus over multiple historical time periods;
[0027] Mapping the surveillance image of the surveillance video data into an image feature sequence, and converting the environmental sensor data, facility maintenance data, and safety feedback data corresponding to the current surveillance video data into a text feature sequence;
[0028] Perform multimodal perceptual streaming processing and spatial semantic alignment on the image feature sequence and the text feature sequence to obtain the same set of multimodal input sequences;
[0029] The decoder of the generative transformer model is fine-tuned using a multimodal input sequence to obtain a security risk identification model. The security risk identification model generates a graph attention network based on the multimodal input sequence, assigns a layer weight to each layer of the graph attention network, and the layer weight decays exponentially with the increase of the layer depth. The product of the confidence of each node on the path and the corresponding layer weight is calculated as the node acceptance rate, and the product of the node acceptance rates of all nodes on the path is used as the global acceptance rate of the path. In the graph attention network, the generated information corresponding to the path with the highest global acceptance rate is selected as the security risk.
[0030] In some specific embodiments, in the step of converting the environmental sensor data, facility maintenance data, and safety feedback data corresponding to the current monitoring video data into a text feature sequence, facility test data is extracted from the environmental sensor data, and facility failure information is extracted from the facility maintenance data and safety feedback data, and the facility test data and facility failure information are converted into a unified text feature sequence, wherein the text feature sequence describes the safety risk of the current monitoring video data.
[0031] In some specific embodiments, surveillance videos are obtained by installing video surveillance cameras at various key locations on campus to ensure coverage of major areas of the campus, such as teaching buildings, playgrounds, canteens, dormitories, etc.; environmental sensors are deployed in classrooms, laboratories, distribution rooms and other places on campus where safety accidents are prone to occur to obtain environmental sensor data, including smoke sensors, temperature sensors, humidity sensors, etc.; the campus facility management and maintenance system generally records information such as the maintenance and upkeep of safety facilities, so facility maintenance data is obtained directly from the campus facility management and maintenance system; feedback channels for teachers and students are opened on campus, such as online feedback platforms and suggestion boxes, to obtain safety feedback data from the teacher and student feedback channels.
[0032] Furthermore, since the directly collected data may contain a lot of irrelevant information and different types of information are not unified feature vectors, the data in the campus safety multimodal dataset is preprocessed and used to fine-tune the generative transformer model.
[0033] Specifically, the target detection algorithm in deep learning is used to extract the features of targets such as people and objects, and the video is edited and labeled to remove irrelevant information to complete the preprocessing of the surveillance video.
[0034] Specifically, denoising is performed through methods such as sliding average filtering, and then the data is normalized to the range of [0, 1] according to the range and accuracy of the sensor to complete the preprocessing of the environmental sensor data.
[0035] Specifically, natural language processing technology is used to perform structured processing and extract key information such as facility name, fault description, feedback time, etc., thereby completing the preprocessing of facility maintenance data and safety feedback data.
[0036] In some embodiments, each layer in the generative Transformer model used in this solution captures the global dependencies of all tokens in the input sequence through multi-head self-attention, and image tokens and text tokens interact directly in the self-attention mechanism without the need for explicit alignment operations in the model.
[0037] In addition, through the multi-head self-attention mechanism, the generative transformer model can pay attention to local details and global context at the same time.
[0038] Specifically, a feedforward network layer is connected after each multi-head self-attention, and the fused features are further processed by the feedforward network layer to enhance the nonlinear expression ability.
[0039] Specifically, when the generative transformer model generates text, a cross-attention mechanism is introduced to allow the text generation process to refer to image information. In the fusion mechanism of modality-specific parameters, cross-attention is not necessary because the image information has been fused with contextual information through self-attention.
[0040] Specifically, the generative transformer model generates text token by token in an autoregressive manner based on the current context in each prediction, and simultaneously optimizes tasks such as image-text alignment and text generation through a unified objective function.
[0041] In some specific embodiments, when fine-tuning the generative transformer model, the learning rate is set to 5e-5, the number of training rounds is set to 10, the maximum gradient range is set to 1.0, the maximum number of samples is set to 1000000, the calculation type is bf16, the truncation length is 2048, the batch size is appropriately allocated according to the video memory size of the graphics card, the gradient accumulation is 8, and cosine is used as the learning rate regulator.
[0042] Furthermore, the generative transformer model uses average cross entropy loss as the loss function during fine-tuning, which is expressed as follows:
[0043]
[0044] Among them, L is the average cross loss on the validation set, E is the minimum loss that cannot be reduced by the dataset itself, A is the weight of controlling the impact of model capacity N on loss, B is the weight of the impact of data volume D on loss, ɑ is the scaling exponent of model parameters, which is used to reflect the sensitivity of loss to changes in model capacity N, and β is the scaling exponent of training data, which is used to reflect the sensitivity of loss to changes in data volume. is the parameter of the image, is the parameter amount of the text, is the number of training tokens for the image, is the number of training tokens of the text, is the image scaling index, is the scaling index of the text, is the model capacity weight coefficient of the image, is the model capacity weight coefficient of the text, is the scaling weight for the amount of image training data, Scaling weight for the amount of text training data.
[0045] Specifically, this scheme discovers the differences in the scaling behaviors of images and text in modal fusion by introducing independent scaling parameters in the loss function to capture the characteristics of different modalities, allowing the model to allocate parameters and data resources more flexibly and improve the modality-specific expression.
[0046] Furthermore, this solution introduces sparse hybrid experts when fine-tuning the generative transformer model, significantly improving performance by dynamically allocating modality-specific expert parameters. The formula is as follows:
[0047]
[0048] Among them, S is the sparsity rate, λ is the penalty term for controlling the expert capacity to the loss, C is the computational budget, and the relationship between C and N and D is: , A is the weight that controls the impact of model capacity N on loss, B is the weight that controls the impact of data volume D on loss, α is the scaling exponent of model parameters, β is the scaling exponent of training data, and λ is the loss weight hyperparameter.
[0049] Specifically, by introducing sparse mixture experts to quantify the performance boundaries of generative transformer models, we suggest the efficiency advantages of the fusion mechanism of modality-specific parameters, the gain mechanism of MoE, and provide a predictable framework for model scaling.
[0050] In some embodiments, the current campus surveillance image is segmented into image blocks of fixed size, and each image block is mapped to a text embedding space through a linear projection layer to obtain an image feature sequence, the user text prompt is segmented, and the segmentation results are converted into a word embedding form to obtain a text description feature sequence. In some embodiments, in the step of performing multimodal perception streaming processing on the image feature sequence and the text feature sequence, a dynamic block function is defined, the image feature sequence and the text description feature sequence are input into the dynamic block function for dynamic block division to output a block sequence consisting of image block results, text block results, and cross-modal alignment block results, and a streaming block attention calculation is performed on the block sequence to obtain a streaming generated sequence, wherein the cross-modal block result is used to align the image local features of the image block result with the paragraph boundaries of the text block result.
[0051] Furthermore, the dynamic block function divides the image feature sequence into blocks based on image salient regions and divides the text description feature sequence into blocks based on keyword importance. The formula of the dynamic block function is as follows:
[0052]
[0053] in, is the image segmentation result, is the input of the image feature sequence at time step t, The result of text segmentation is: is the input of the text description feature sequence at time step t, To align the segmentation results across modalities, it is used to align the image blocks with the boundaries of text paragraphs. is the image segmentation threshold, which is used to determine whether the image input needs to be segmented independently. The text segmentation threshold is used to determine whether the text input needs to be segmented independently. and According to training learning or adaptive adjustment, specifically, the dynamic blocking function in this scheme adaptively adjusts the size of the block by content. It is essentially a hierarchical computing method that aims to reduce the computational cost of complex systems while maintaining sensitivity to key information.
[0054] In some embodiments, the formula for calculating the streaming block attention mechanism for the block sequence is expressed as follows:
[0055]
[0056] in, Generate a sequence for streaming, is the query matrix of the previous block, is the key matrix of the current block, V is the value matrix of the current block, is the position similarity function used to calculate the cross-position alignment score, is a hyperparameter that controls the strength of position interaction.
[0057] It should be noted that this solution performs streaming block attention calculation on the text block results and image block results in the block sequence respectively to obtain adjusted text block results and image block results as a streaming generation sequence.
[0058] In some embodiments, in the step of performing spatial semantic alignment on the image feature sequence and the text feature sequence, the streaming generated sequence is input into a trained cross-modal alignment model to obtain a multimodal input sequence, wherein the cross-modal alignment model into which the streaming generated sequence is input performs position cross-modal attention and semantic cross-modal attention calculations to obtain semantic-position image space coding, the semantic-position image space coding is multi-scale decomposition obtained image features of different spatial scales, the text position of the text sequence is encoded to obtain text position coding features, and the image features and text position coding features of different spatial scales are fused layer by layer to obtain a multimodal input sequence.
[0059] Furthermore, the text positions of the text segmentation results in the streaming generation sequence are encoded and mapped to the image space dimension to obtain text position encoding, the image coordinates of the image segmentation results in the streaming generation sequence are encoded to obtain image space encoding, the image space encoding and the text position encoding are mixed to obtain position alignment encoding, the position similarity item is calculated based on the position alignment encoding, and the image space encoding is subjected to position cross-modal attention calculation to obtain position image space encoding, wherein the position similarity item is introduced into the position cross-modal attention calculation as a position alignment constraint.
[0060] Specifically, the text position of the text segmentation result is encoded using the PoPE encoding method to obtain the text position encoding, and the image coordinates of the image segmentation result are encoded using the Coord encoding method to obtain the image space encoding. In addition, a position alignment matrix is constructed and used to encode the image space encoding and the text position encoding. The formula is expressed as follows:
[0061]
[0062] in, For position alignment encoding, Encode the text position, For image space encoding, is the position alignment matrix, and .
[0063] Specifically, the calculation formula of the position similarity item is expressed as:
[0064]
[0065] in, is the position similarity term, For position alignment encoding, 、 For the corresponding query and key.
[0066] Specifically, the formula for calculating the position cross-modal attention is as follows:
[0067]
[0068] in, is the calculated position image space encoding, is the position similarity term constructed by text position coding and image space coding, V is the corresponding image space coding, and α is the scaling exponent of the model parameters.
[0069] Furthermore, the text segmentation results in the streaming generated sequence are semantically encoded to generate a dynamic alignment matrix, the text positions of the text sequence are encoded and mapped to the image space dimension based on the dynamic alignment matrix to obtain text semantic encoding, the image space encoding and text semantic encoding are mixed to obtain semantic alignment encoding, the semantic similarity items are calculated based on the semantic alignment encoding, and semantic cross-modal attention calculation is performed on the position image space encoding to obtain semantic-position image space encoding, wherein the semantic similarity items are introduced into the semantic cross-modal attention calculation as semantic alignment constraints.
[0070] Specifically, this solution introduces a learnable projection network and a CLIP semantic encoder into the position alignment matrix to obtain a dynamic alignment matrix. The formula of the dynamic alignment matrix is as follows:
[0071]
[0072] in, is the dynamic alignment matrix, is a learnable projection network, and CLIP is a CLIP semantic encoder. That is, this scheme uses the CLIP semantic encoder to understand the semantics of the text sequence and uses the semantic information of the text sequence to generate a dynamic alignment matrix related to the text semantics, thereby mixing image space encoding and text semantic encoding to obtain semantic alignment encoding.
[0073] Use the dynamic alignment matrix to encode the image space and text position. The formula is as follows:
[0074]
[0075] in, For semantic alignment encoding, Encode the text position, For image space encoding, is the dynamic alignment matrix.
[0076] Specifically, the calculation formula of the semantic similarity item is expressed as:
[0077]
[0078] in, are semantically similar terms, For semantic alignment encoding, 、 For the corresponding query and key.
[0079] Specifically, the formula for calculating the semantic cross-modal attention is as follows:
[0080]
[0081] in, is the calculated semantic-position image space encoding, is the semantic similarity term, V is the corresponding image space encoding, and α is the scaling exponent of the model parameters.
[0082] Specifically, since the text semantic coding and the image spatial coding are both obtained based on a matrix of the same dimension, after obtaining the image spatial coding and the text semantic coding, the image spatial coding and the text semantic coding can be directly mixed to obtain the semantic alignment coding.
[0083] Furthermore, the semantic-position image spatial coding is decomposed into multiple scales through multi-layer coding to obtain image features of different spatial scales. Similarly, the text position coding is decomposed into multiple scales through multi-layer coding to obtain text position coding of different scales, and then fused using the following formula:
[0084]
[0085] in, Indicates the l The fusion result of the layer, for l The image features of the layer, for l The text position code of the layer, Represents residual connection, used to retain multi-scale information, is the upsampling operation.
[0086] In some specific embodiments, the loss function of the cross-modal alignment model is:
[0087]
[0088] in, is the loss function of the cross-modal alignment model, To represent the expectation of a randomly sampled image-text pair (I, T) in the data distribution D, is the semantic similarity score between image I and text T, is the temperature coefficient, is the negative sample text set of image I, is the loss weight hyperparameter, is the position alignment loss.
[0089] In some specific embodiments, a generation priority function is defined to adjust the generation direction of the security risk identification model, and the formula is as follows:
[0090]
[0091] in, To generate the priority function, is a Sigmoid function that outputs a priority score between 0 and 1. is the parameter matrix, The content generated by the model, is the attention weight of the image.
[0092] Specifically, this solution uses a priority function to guide the model to generate content according to a specific logic, thereby adjusting the model's generation order. If the image area is prioritized, the model will give priority to generating content with high similarity to the image. If semantic coherence is prioritized, the model will prioritize ensuring that the generated content has high contextual semantic coherence.
[0093] In some specific embodiments, a cross-modal caching mechanism is used to generate security risks, and the formula is expressed as follows:
[0094]
[0095] in, is the cross-modal cache state at the current time t, is the gate weight matrix, which is used to control the degree of historical information retention. is the cross-modal cache status at the previous moment, is the dynamic block function, It is a cross attention mechanism.
[0096] Specifically, the core function of the cross-modal caching mechanism is to solve the problems of computational efficiency and information coherence in long sequence generation tasks by storing and reusing the interaction information between historical modalities. The cross-modal caching mechanism avoids repeated calculations by storing historical interaction states and dynamically updating strategies, reduces video memory usage, and reduces logical errors in generated content by enforcing cross-modal consistency. It supports dynamic adjustment of cache granularity to adapt to the differences in importance of different modalities.
[0097] In some specific embodiments, the security risk identification model is an end-to-end model calculation, which is a single transformer architecture. That is, the security risk identification model is constructed by only one transformer decoder, which includes multiple layers of self-attention and feedforward network modules.
[0098] Specifically, each layer of the Transformer captures the global dependencies of all tokens in the input sequence through multi-head self-attention. Image tokens and text tokens interact directly in the self-attention without the need for explicit alignment operations. Through the multi-head mechanism, the model can simultaneously focus on local details (such as image blocks) and global context (such as text paragraphs).
[0099] In some embodiments, the level weight satisfy , where j is the level, As a hyperparameter, the global acceptance rate of different nodes in this scheme is calculated as follows:
[0100]
[0101] in, represents the global acceptance rate up to the current node, is the level weight, is the confidence of the current node, From the root node to the current node Intermediate nodes on the path.
[0102] Specifically, this solution can reduce the contribution of deep nodes by setting hierarchical weights that decay exponentially with the increase of hierarchical depth, thereby avoiding an artificially high global acceptance rate caused by an excessively long path. In other words, by setting hierarchical weights, it can avoid the answers given by the security risk identification model being too lengthy.
[0103] In some embodiments, in the step of selecting the path with the highest global acceptance rate as the security risk, if there are multiple paths with the same global acceptance rate, the path with the shallowest level is preferentially selected as the security risk. The formula for selecting the path with the highest global acceptance rate is expressed as follows:
[0104]
[0105] in, represents the global acceptance rate of the path, From the root node to the current node Intermediate nodes on the path.
[0106] In some specific embodiments, in the process of selecting the path with the highest global acceptance rate, contextual features are introduced to ensure a balance between generation quality and efficiency. The formula is as follows:
[0107]
[0108] in, represents the global acceptance rate up to the current node, is the context feature function, For nodes Contextual features, From the root node to The path, is the confidence of the current node.
[0109] In some specific embodiments, the flowchart for fine-tuning the security risk identification model is as follows: Figure 2 As shown in the figure, model fine-tuning mainly includes steps such as environment construction, data preparation, model loading and configuration, data organization, training, and model saving. Each step has a significant impact on the final performance of the model, as follows:
[0110] 1. Environment Setup: Prepare a GPU with at least 18GB of video memory. Create a Python virtual environment and activate it. For example, run 'python -m venv venv' to create it and 'source venv / bin / activate' to activate it.
[0111] Install related dependency libraries. You can also refer to other resources, such as installing 'torch transformers datasetspeft accelerate qwen-vl-utils swanlab'. The Python version must be greater than or equal to 3.9.
[0112] 2. Data Preparation
[0113] Data Structure Organization: Organize the data in the data directory, divided into train (training set) and val (validation set) subdirectories. Each subdirectory contains the annotations.jsonl file and the corresponding image files. The annotation files must describe the image information in a specific format, such as the image name, prefix, and suffix (including task-specific information).
[0114] Data loading and processing: Define the jsonl_dataset class to handle data loading and processing, by reading the jsonl annotation file, loading the corresponding image, and formatting the data into the required conversation structure (such as the three-turn conversation format: system message, user message, assistant message).
[0115] 3. Model loading and configuration
[0116] Choose a fine-tuning method: Use LoRA (Low Rank Adaptation) or QLoRA (Quantized Low Rank Adaptation). LoRA reduces memory usage and training time by training only a small number of adapter parameters; QLoRA further quantizes the base model to 4 bits, reducing memory requirements while maintaining performance.
[0117] Load the model and configure parameters: Import related libraries, such as torch, peft, transformers, etc. Specify the model ID, such as "qwen / qwen2.5-vl-3b-instruct", and determine the device (cuda or cpu).
[0118] Configure LoRA parameters, such as lora_alpha, lora_dropout, r, bias, target_modules, and task_type. If using QLoRA, also configure bitsandbytes_config, including parameters such as load_in_4bit, bnb_4bit_use_double_quant, bnb_4bit_quant_type, and bnb_4bit_compute_type.
[0119] Load the model and apply the configuration, such as
[0120] model=qwen2_5_vlforconditionalgeneration.from_pretrained(model_id,device_map="auto",quantization_config=bnb_config if use_qlora else None,torch_dtype=torch.bfloat16), model=get_peft_model(model,lora_config).
[0121] 4. Define data sorting functions
[0122] Training data collation: The 'train_collate_fn' function prepares the training data by applying chat templates to the text, processing images into the correct format, creating attention masks, managing special tags (such as padding and image tags), and preparing labels for loss calculation.
[0123] Evaluation data wrangling: The evaluation wrangling function differs from the training wrangling function in that it does not require labels, preserves the target suffix for comparison, and strips assistant responses.
[0124] 5. Train the model: Set training parameters such as the number of training rounds, batch size, and learning rate. Use an appropriate training framework (such as Lightning) to train the model. During training, you can visualize the loss to observe the model training progress and determine whether the model has converged or is experiencing overfitting.
[0125] 6. Save the fine-tuned model: After training is complete, save the fine-tuned model for subsequent use in inference or further optimization. The saved model can be loaded and used in new tasks or scenarios to achieve efficient processing of specific tasks.
[0126] In some specific embodiments, security risks are divided into three levels: low, medium, and high, and corresponding early warning strategies are formulated for different levels. For the identification results of security risks, information such as the location, type, and severity of the risk are listed in detail.
[0127] Specifically, when a safety risk is high or medium, the system automatically sends an early warning to the safety management department via text message, email, system notification, and other means. Simultaneously, based on the safety risk identification results, a facility maintenance work order is generated, and maintenance personnel are assigned to promptly address the safety risk. Based on the safety risk identification results, the safety management department can develop long-term safety management plans, such as strengthening safety education and optimizing facility layout.
[0128] In some specific embodiments, a front-end page is constructed to enable users to identify security risks through the security risk identification model constructed by this solution. The specific identification process is as follows: Figure 3 As shown, the specific contents of the flowchart include: 1. Start: The user opens the "AI Risk Identification" application, and the process starts here. 2. Interface display: The application interface is divided into the top area, the core function area, and the bottom area. The top area displays the time, signal, Wi-Fi, and battery status; the core function area contains a construction site picture preview window, recognition result display, and rectification suggestion label; the bottom area has a "Save Risk" button. 3. User operation branch: The user can choose three operations: Re-identify: Click to return to the picture upload step, re-upload or analyze the picture. Save identified risks: Click to save the recognition results and end the process. Risk basis: Click to view a detailed explanation, which may be displayed in the form of a pop-up window or jump page. 4. End: The process ends after the user saves the results or views the risk basis.
[0129] Example 2
[0130] Based on the same concept, this application also proposes a security risk identification method based on a multimodal large model, including:
[0131] Obtain surveillance video data to be identified;
[0132] Input the surveillance video data to be identified and the text prompt into the multimodal large model for security risk identification constructed by the method for constructing a multimodal large model for security risk identification constructed in Example 1, and output the security risk;
[0133] The surveillance video data is mapped into an image feature sequence, the text prompt is mapped into a text feature sequence, and the image feature sequence and the text feature sequence are spliced to be input into a multimodal large model for security risk identification.
[0134] In some embodiments, since the security risk identification model in this solution is a generative model, the text prompt is the prompt language input by the user into the security risk identification model when using it, for example: Please help me judge the safety risk of this picture, or please help me judge whether there are exposed wires in the image, etc. The security risk identification model in this solution maps the current campus surveillance image and the text prompt into a sequence representation, and completes the splicing through semantic alignment, so that the security risk identification model can better understand the image and text. The spliced multimodal input sequence can be expressed in the format of: [image block 1, image block 2, ..., text token1, text token2].
[0135] Example 3
[0136] This embodiment also provides an electronic device, referring to Figure 4 , includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.
[0137] Specifically, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0138] Memory 404 may include a large-capacity memory 404 for data or instructions. By way of example, and not limitation, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to the data processing device. In certain embodiments, memory 404 is non-volatile memory. In certain embodiments, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. In appropriate circumstances, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM may be a fast page mode dynamic random access memory 404 (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0139] The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402 .
[0140] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any one of the security risk identification methods based on the multimodal large model in the above embodiments.
[0141] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .
[0142] Transmission device 406 can be used to receive or transmit data via a network. Specific examples of such networks may include wired or wireless networks provided by the electronic device's communications provider. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0143] The input and output devices 408 are used to input or output information. In this embodiment, the input information may be surveillance video data, environmental sensor data, facility maintenance data, and safety feedback data from multiple historical periods on campus, and the output information may be safety risks.
[0144] Optionally, in this embodiment, the processor 402 may be configured to execute the following steps through a computer program:
[0145] Constructing a multimodal campus safety dataset, wherein the multimodal campus safety dataset includes surveillance video data, environmental sensor data, facility maintenance data, and safety feedback data on campus over multiple historical time periods;
[0146] Mapping the surveillance image of the surveillance video data into an image feature sequence, and converting the environmental sensor data, facility maintenance data, and safety feedback data corresponding to the current surveillance video data into a text feature sequence;
[0147] Perform multimodal perceptual streaming processing and spatial semantic alignment on the image feature sequence and the text feature sequence to obtain the same set of multimodal input sequences;
[0148] The decoder of the generative transformer model is fine-tuned using a multimodal input sequence to obtain a security risk identification model. The security risk identification model generates a graph attention network based on the multimodal input sequence, assigns a layer weight to each layer of the graph attention network, and the layer weight decays exponentially with the increase of the layer depth. The product of the confidence of each node on the path and the corresponding layer weight is calculated as the node acceptance rate, and the product of the node acceptance rates of all nodes on the path is used as the global acceptance rate of the path. In the graph attention network, the generated information corresponding to the path with the highest global acceptance rate is selected as the security risk.
[0149] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.
[0150] In general, various embodiments may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0151] The embodiments of the present invention may be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros may be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer executable components that are configured to perform an embodiment when the program is run. One or more computer executable components may be at least one software code or a portion thereof. In addition, it should be noted at this point that, for example, Figure 4 Any block of the logic flow in the program may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. Physical media are non-transitory media.
[0152] Those skilled in the art should understand that the various technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0153] The above embodiments merely illustrate several embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for constructing a multimodal large model for security risk identification, characterized in that: The following steps are involved: Constructing a multimodal campus safety dataset, wherein the multimodal campus safety dataset includes surveillance video data, environmental sensor data, facility maintenance data, and safety feedback data on campus over multiple historical time periods; Mapping the surveillance image of the surveillance video data into an image feature sequence, and converting the environmental sensor data, facility maintenance data, and safety feedback data corresponding to the current surveillance video data into a text feature sequence; Performing multimodal perception streaming processing and spatial semantic alignment on the image feature sequence and the text feature sequence to obtain the same set of multimodal input sequences, defining a dynamic blocking function in the step of performing multimodal perception streaming processing on the image feature sequence and the text feature sequence, inputting the image feature sequence and the text description feature sequence into the dynamic blocking function for dynamic blocking to output a blocking sequence consisting of image blocking results, text blocking results, and cross-modal alignment blocking results, performing streaming blocking attention calculation on the blocking sequence to obtain a streaming generated sequence, wherein the cross-modal alignment blocking result is used to align the image local features of the image blocking result with the paragraph boundaries of the text blocking result; The decoder of the generative transformer model is fine-tuned using a multimodal input sequence to obtain a security risk identification model. The security risk identification model generates a graph attention network based on the multimodal input sequence, assigns a layer weight to each layer of the graph attention network, and the layer weight decays exponentially with the increase of the layer depth. The product of the confidence of each node on the path and the corresponding layer weight is calculated as the node acceptance rate, and the product of the node acceptance rates of all nodes on the path is used as the global acceptance rate of the path. In the graph attention network, the generated information corresponding to the path with the highest global acceptance rate is selected as the security risk.
2. The method for constructing a multimodal large model for security risk identification according to claim 1, characterized in that: In the step of converting the environmental sensor data, facility maintenance data, and safety feedback data corresponding to the current monitoring video data into a text feature sequence, facility test data is extracted from the environmental sensor data, and facility failure information is extracted from the facility maintenance data and safety feedback data, and the facility test data and facility failure information are converted into a unified text feature sequence, wherein the text feature sequence describes the safety risk of the current monitoring video data.
3. The method for constructing a multimodal large model for security risk identification according to claim 1, characterized in that: In the step of spatial semantic alignment of image feature sequences and text feature sequences, the streaming generated sequence is input into the trained cross-modal alignment model to obtain a multimodal input sequence, wherein the streaming generated sequence is input into the cross-modal alignment model to perform position cross-modal attention and semantic cross-modal attention calculation to obtain semantic-position image space coding, the semantic-position image space coding is multi-scale decomposition obtained image features of different spatial scales, the text position of the text sequence is encoded to obtain text position coding features, and the image features and text position coding features of different spatial scales are fused layer by layer to obtain a multimodal input sequence.
4. The method for constructing a multimodal large model for security risk identification according to claim 3, characterized in that: In the step of obtaining the position image space encoding, the text position of the text segmentation results in the streaming generation sequence is encoded and mapped to the image space dimension to obtain the text position encoding, the image coordinates of the image segmentation results in the streaming generation sequence are encoded to obtain the image space encoding, the image space encoding and the text position encoding are mixed to obtain the position alignment encoding, the position similarity item is calculated based on the position alignment encoding, and the image space encoding is subjected to position cross-modal attention calculation to obtain the position image space encoding, wherein the position similarity item is introduced into the position cross-modal attention calculation as a position alignment constraint.
5. The method for constructing a multimodal large model for security risk identification according to claim 3, characterized in that: In the step of inputting the streaming generated sequence into the cross-modal alignment model to perform position cross-modal attention and semantic cross-modal attention calculations to obtain the semantic-position image space encoding, the text segmentation results in the streaming generated sequence are semantically encoded to generate a dynamic alignment matrix, the text positions of the text sequence are encoded and mapped to the image space dimension based on the dynamic alignment matrix to obtain the text semantic encoding, the image space encoding and the text semantic encoding are mixed to obtain the semantic alignment encoding, the semantic similarity items are calculated based on the semantic alignment encoding, and the semantic cross-modal attention is calculated on the position image space encoding to obtain the semantic-position image space encoding, wherein the semantic similarity items are introduced into the semantic cross-modal attention calculation as semantic alignment constraints.
6. The method for constructing a multimodal large model for security risk identification according to claim 3, characterized in that: The loss function of the cross-modal alignment model is: in, is the loss function of the cross-modal alignment model, To represent the expectation of a randomly sampled image-text pair (I, T) in the data distribution D, is the semantic similarity score between image I and text T, is the temperature coefficient, is the negative sample text set of image I, is the loss weight hyperparameter, is the position alignment loss.
7. The method for constructing a multimodal large model for security risk identification according to claim 1, characterized in that: A generation priority function is defined to adjust the generation direction of the security risk identification model. The formula is as follows: in, To generate the priority function, is a Sigmoid function that outputs a priority score between 0 and 1. is the parameter matrix, The content generated by the model, is the attention weight of the image.
8. The method for constructing a multimodal large model for security risk identification according to claim 1, characterized in that: The formula for exponential decay of level weight as level depth increases is as follows: ; in, is the level weight, j is the level, is a hyperparameter.
9. A security risk identification method based on a multimodal large model, characterized in that: Obtain surveillance video data to be identified; Inputting the surveillance video data to be identified and the text prompt into the multimodal large model for security risk identification obtained by the construction method according to any one of claims 1 to 8, and outputting the security risk; The surveillance video data is mapped into an image feature sequence, the text prompt is mapped into a text feature sequence, and the image feature sequence and the text feature sequence are spliced to be input into a multimodal large model for security risk identification.
Citation Information
Patent Citations
Multi-modal representation learning method based on text guide image block screening
CN117421591A
Multi-modal emotion recognition method based on double-flow encoder and attention mechanism
CN118296548A