Engineering safety risk early warning method and system based on visual language large model

Through the visual language model, the visual content summary of the project site was extracted and the risk assessment coefficient was generated, which solved the problem of high false alarm rate of the existing monitoring system, and achieved efficient and accurate early warning of safety risks on the project site.

CN120373836APending Publication Date: 2025-07-25JINQIANMAO TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510232025.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing vision-based monitoring system has problems such as high false alarm rate in engineering site safety risk warning and limited fine-grained attributes in different workplace scenarios.

Method used

Using an engineering security risk warning method based on visual language big model, by constructing and training the visual language big model, extracting visual content summary of scene data, generating security factors and object feature matrix, calculating risk assessment coefficients, and performing time series analysis to reduce false positive rate.

Benefits of technology

It improves the accuracy and reliability of safety risk warnings on the project site, reduces false alarms, and can promptly detect potential hazards and trigger alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373836A_ABST
    Figure CN120373836A_ABST
Patent Text Reader

Abstract

The invention discloses an engineering safety risk early warning method and system based on a visual language large model. The method comprises the steps that scene data are extracted on a monitoring platform, a scene analysis model is used for obtaining a visual content abstract, the type of a current scene is judged, and a detection object is extracted; a visual prompt function module is used, and a large language model is used to generate necessary safety detection items according to a scene; detecting safety factors required by the scene and objects needing to be subjected to safety item detection in the scene, and generating a safety risk assessment coefficient; and performing risk assessment statistical analysis on the scene data within a period of time, and if the risk exceeds a security risk threshold, sending security alarm information to the platform. According to the invention, the monitoring camera is utilized to collect the picture of the whole engineering site in real time, the scene content is analyzed and judged, whether the safety risk exists is judged, the alarm is given out in time, and the reliability of safety rule inspection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence monitoring, and in particular to an engineering safety early warning method and system based on a vision-language large model. Background Art

[0002] To ensure that engineering projects can be strictly implemented in accordance with specifications and processes, a scientific and reliable safety management plan and an accurate risk early warning system are required to maintain the normal operation of the project. To conduct safety risk early warning, construction safety needs to be monitored. Existing vision-based monitoring systems need to classify each frame to identify safe or unsafe scenarios, which usually trigger false alarms due to object misdetection or false detection, thus reducing the performance of the entire monitoring system. Moreover, most existing object detection models such as Yolo, SSD, etc. have limitations in verifying the fine-grained attributes of target objects in different workplace scenarios. Therefore, how to efficiently and accurately monitor the engineering site and conduct safety early warning is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0003] In view of the above problems, the present application provides an engineering safety early warning method and system based on a vision-language large model to solve the technical problems existing in the above engineering risk early warning.

[0004] To achieve the above object, the present application provides an engineering safety risk early warning method based on a vision-language large model, including the following steps:

[0005] Construct a vision-language large model and train the vision-language large model;

[0006] Extract the scene data of the project from the monitoring platform of the project, use the trained vision-language large model to obtain the visual content summary in the scene data, and judge the current scene type according to the visual content summary, and extract the objects that need to be detected for safety items in the current scene type; the detected objects include the scene area and construction workers, and the detection of construction workers includes personnel equipment and personnel behavior;

[0007] Detect the safety factors required for the scene and the objects that need to be detected for safety items in the scene, including:

[0008] Analyze the visual content summary to obtain scene information, use the scene information to query the vision-language large model to obtain the items that construction workers should wear and prohibited behaviors in the current scene type;

[0009] Generate visual features corresponding to the items that should be worn and prohibited behaviors in the current scene type;

[0010] Generate bounding boxes for all objects in the scene, extract image patches containing the detected objects from the scene image according to the coordinate information of the bounding boxes, and use the pre-trained VLM model feature extraction module to perform feature extraction to generate the feature matrix F of multiple detected objects object and the scene feature F image ;

[0011] Generate text features according to the predefined safety detection items and the safety detection items of the visual language large model based on the current scene type;

[0012] The text feature of the wearing part is F item , the text feature of the behavior part is F behave , and the text feature of the scene safety factor is F risk ;

[0013] Generate a scene factor matrix, an object wearing matrix, and an object behavior matrix based on each of the text features, and calculate the probability P of danger based on the scene factor matrix, the object wearing matrix, and the object behavior matrix;

[0014] Generate the risk assessment coefficient of the current scene according to the probability P and the set danger coefficient.

[0015] Furthermore, the scene factor matrix is: S Image = F image T × F risk ;

[0016] The object wearing matrix is: S wear = F object T × F item ;

[0017] The object behavior matrix is: S behavet = F object T × F behave ;

[0018] The probability P = softmax(S);

[0019] where S = {S Image , S wear , S behavet}, when the probability P of S wear is greater than the threshold, it is considered that the object wears the relevant equipment and the text description in the current scene type match, otherwise it means not wearing or wearing inconsistently; when S Image and S behavet are greater than the threshold, it is considered that there are relevant scene danger factors and dangerous behaviors.

[0020] Further, it also includes performing time analysis on the scenario data, including:

[0021] Counting the number of detected frames in the current time period in the scenario data;

[0022] Marking the frames with a safety risk coefficient belonging to the unsafe level;

[0023] Counting the number of frames belonging to the unsafe level and calculating the overall proportion. When the proportion is greater than the threshold θ, it is determined that a safety warning is required currently.

[0024] Further, the training of the visual language large model includes:

[0025] Obtaining scenario data with safety risks and performing manual analysis and description on the scenario data to construct a dataset for training;

[0026] Dividing the safety risk objects in the scenario data, where the safety risk objects include the scenario area and construction workers;

[0027] Defining the safety detection items for the scenario area and construction workers and setting the risk coefficient;

[0028] Using the dataset to train and fine-tune the visual language large model.

[0029] Further, the using the dataset to train and fine-tune the visual language large model includes:

[0030] For the visual language large model, given a pre-trained weight matrix W0 ∈ R m×n , select a low rank r, r << d, where d is the original dimension of the matrix;

[0031] Perform singular value decomposition on the pre-trained matrix to obtain a low-rank matrix of m singular values, where m is the dimension of the pre-trained weight W0. The decomposition formula is as follows:

[0032] SVD(W0) = U∑V T ;

[0033] U is an m×m orthogonal matrix, Σ is an m×n diagonal matrix, where the elements on the diagonal are singular values and the other elements are 0, and V T is the transpose of an n×n orthogonal matrix, and S = [S1, S2, S3... S m is the singular value matrix of the pre-trained matrix W0;

[0034] Initialize two low-rank matrices A and B according to the weights of the pre-trained model. Compared with the original method where A is randomly initialized with Gaussian and B is initialized with zeros, the initialization of A and B is as follows;

[0035]

[0036]

[0037] Among them, U r 、 represents the sub-matrices of U and V corresponding to the selected r singular values, S T is a diagonal matrix with r singular values on the diagonal, r represents the square root of the S r matrix;

[0038] During the training process, W0 is frozen and does not receive gradient updates. A and B contain the fine-tuned weights, representing the differences to be added to the original weights. During the inference process, the fine-tuned weights are combined with the original pre-trained weights. For the input x, the modified forward pass can be expressed as:

[0039]

[0040]

[0041]

[0042] To solve the above technical problems, the present application also provides another technical solution:

[0043] An engineering safety risk early warning system based on a vision-language large model, comprising:

[0044] A data acquisition module for extracting the scene data of the project from the monitoring platform of the project;

[0045] A scene analysis module for obtaining the visual content summary in the scene data with the trained vision-language large model, judging the current scene type according to the visual content summary, and extracting the objects that need to be detected for safety items of the current scene type; the objects to be detected include the scene area and construction workers, and the detection of construction workers includes personnel equipment and personnel behavior;

[0046] A visual prompt module for detecting the safety factors required for the scene and the objects that need to be detected for safety items in the scene;

[0047] A safety detection module for generating a scene factor matrix, an object wearing matrix, and an object behavior matrix based on each of the text features, and calculating the probability P of danger based on the scene factor matrix, the object wearing matrix, and the object behavior matrix;

[0047] A risk assessment module, which is used to generate a risk assessment coefficient for the current scenario according to the probability P and the set danger coefficient.

[0048] Further, it further includes a time analysis module, which is used to count the number of detected frames in the current time period in the scenario data;

[0049] Mark the frames whose safety risk coefficients belong to the unsafe level;

[0050] Count the number of frames belonging to the unsafe level, and calculate the overall proportion. When the proportion is greater than the threshold θ, it is determined that safety warning is required currently.

[0051] Further, the scenario factor matrix is: S Image = F image T × F risk ;

[0052] The object wearing matrix is: S wear = F object T × F item ;

[0053] The object behavior matrix is: S behavet = F object T × F behave ;

[0054] The probability P = softmax(S);

[0055] where S = {S Image , S wear , S behavet}, when the probability P of S wear is greater than the threshold, it is considered that the object in the current scenario type wears relevant equipment and matches the text description, otherwise it means not wearing or wearing inconsistently; when S Image and S behavet are greater than the threshold, it is considered that there are relevant scenario danger factors and dangerous behaviors.

[0056] Further, the training of the visual language large model includes:

[0057] Obtain scenario data with safety risks, and conduct manual analysis and description of the scenario data to construct a dataset for training;

[0058] Divide the safety risk objects in the scenario data, and the safety risk objects include the scenario area and construction workers;

[0059] Define the safety detection items for the scenario area and construction workers, and set the risk coefficient;

[0060] Train and fine-tune the vision-language large model using the said dataset.

[0061] Different from the prior art, aiming at the limitations of the existing vision-based monitoring systems and object detection models, this technical solution accurately interprets scene information through a vision-language large model and bridges the semantic gap between vision and text data through a large language model, so that various safety problems in the engineering site can be detected more accurately. And in this technical solution, during the fine-tuning stage of training the vision-language large model, an integrated adapter initialized with a non-singular value matrix is introduced, reducing the uncertainty of model prediction. In the safety project detection stage, image patches of different detection items and their corresponding prompt words are used for feature extraction and matrix association to determine whether there are various risk factors and calculate the risk assessment coefficient; finally, time-series-based result statistical analysis of the scene data is carried out to further reduce the inaccuracy of prediction.

[0062] The above relevant descriptions of the invention content are only an overview of the technical solution of this application. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of this application, and thus can be implemented according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features and advantages of this application more easily understood, the following is described in conjunction with the specific embodiments and drawings of this application. Description of the Drawings

[0063] The drawings are only used to illustrate the principles, implementation methods, applications, features and effects of the specific embodiments of the present invention and other related contents, and should not be regarded as a limitation to this application.

[0064] In the drawings of the specification:

[0065] Figure 1 It is a flowchart of the engineering safety risk warning method based on the vision-language large model described in the specific embodiment;

[0066] Figure 2 It is a block diagram of the modules of the engineering safety risk warning system based on the vision-language large model described in the specific embodiment;

[0067] The descriptions of the reference numerals involved in the above drawings are as follows:

[0068] 200, Engineering safety risk warning system based on vision-language large model; 201, Data acquisition module; 202, Scene analysis module; 203, Vision prompt module; 204, Safety detection module; 205, Risk assessment module; 206, Time analysis module. Specific Embodiments

[0069] To elaborate in detail on the possible application scenarios, technical principles, specific implementable solutions, achievable objectives and effects of this application, etc., the following will be described in detail with reference to the specific examples listed and in conjunction with the accompanying drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of this application, and thus are only examples and cannot be used to limit the protection scope of this application.

[0070] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments may be included in at least one embodiment of this application. The term "embodiment" appearing at various positions in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in this application, as long as there is no technical contradiction or conflict, the various technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.

[0071] Unless otherwise defined, the meanings of the technical terms used herein are the same as those generally understood by those skilled in the technical field to which this application belongs; the use of the relevant terms herein is only for describing specific embodiments and is not intended to limit this application.

[0072] In the description of this application, the phrase "and / or" is an expression used to describe the logical relationship between objects, indicating that three relationships may exist. For example, A and / or B means: there is A, there is B, and there is both A and B at the same time. In addition, the character " / " herein generally represents an "or" logical relationship between the associated objects before and after.

[0073] In this application, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantity, primary-secondary or order relationship, etc. between these entities or operations.

[0074] Without more limitations, in this application, the open expressions such as "including", "comprising", "having" or other similar ones used in the statements are intended to cover non-exclusive inclusion. These expressions do not exclude that there may be other elements in the process, method or product including the said elements, so that the process, method or product including a series of elements may not only include those defined elements, but also include other elements not explicitly listed, or also include the elements inherent in such process, method or product.

[0075] Similar to the understanding in the "Examination Guidelines", in this application, expressions such as "greater than", "less than", "exceeding", etc. are understood not to include the present number; expressions such as "above", "below", "within", etc. are understood to include the present number. In addition, in the description of the embodiments of this application, the meaning of "a plurality of" is two or more (including two). Similar expressions related to "many" are also understood in this way, such as "multiple groups", "multiple times", etc., unless otherwise clearly and specifically defined.

[0076] In the description of the embodiments of this application, the spatially related expressions used, such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "perpendicular", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the specific embodiment or the accompanying drawings. It is only for the convenience of describing the specific embodiments of this application or for the reader to understand, rather than indicating or implying that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, it cannot be understood as a limitation to the embodiments of this application.

[0077] Unless otherwise clearly specified or limited, in the description of the embodiments of this application, the terms "installed", "connected", "joined", "fixed", "set", etc. should be understood in a broad sense. For example, the "connection" can be a fixed connection, a detachable connection, or an integral setting; it can be a mechanical connection, an electrical connection, or a communication connection; it can be directly connected, or indirectly connected through an intermediate medium; it can be the communication inside two components or the interaction relationship between two components. For those skilled in the art to which this application belongs, the specific meanings of the above terms in the embodiments of this application can be understood according to specific circumstances.

[0078] Please refer to Figure 1 , this embodiment provides an engineering safety risk early warning method based on a vision-language large model. The engineering safety risk early warning method based on the vision-language large model can be applied to the monitoring and risk early warning of various construction project sites, and is used to timely discover the risks existing in the project site and avoid the occurrence of dangerous events. In this embodiment, a vision-language model (VLM) is used. The vision-language model (VLM) provides a promising solution to the object detection limitations in engineering scene analysis by utilizing the synergy between visual and text information. This embodiment proposes a safety risk early warning method and system based on the vision-language large model and time analysis, which triggers an alarm when the comprehensive evaluation coefficient exceeds the danger threshold, thereby reducing false alarms and improving the reliability of safety rule inspections.

[0079] As Figure 1As shown in the figure, the engineering safety risk early warning method based on the vision-language large model includes the following steps:

[0080] S101. Construct a vision-language large model and train the vision-language large model.

[0081] S102. Extract the scene data of the project from the monitoring platform of the project, use the trained vision-language large model to obtain the visual content summary in the scene data, judge the current scene type according to the visual content summary, and extract the objects that need to be detected for safety items in the current scene type; the objects to be detected include the scene area and construction workers, and the detection of construction workers includes personnel equipment and personnel behavior.

[0082] S103. Detect the safety factors required for the scene and the objects that need to be detected for safety items in the scene; among them, step S103 includes:

[0083] Analyze the visual content summary to obtain scene information, use the scene information to query the vision-language large model to obtain the items that construction workers should wear and prohibited behaviors in the current scene type;

[0084] Generate visual features corresponding to the items to be worn and prohibited behaviors in the current scene type. Generating the visual features aims to guide the LLM to summarize the specific visual attributes required for the safety items, because the same item may have different fine-grained requirements in different scenes.

[0085] Generate bounding boxes for all objects in the scene, extract image patches containing the detection objects from the scene image according to the coordinate information of the bounding boxes, and use the pre-trained VLM model feature extraction module to perform feature extraction to generate the feature matrices F of multiple detection objects object and the scene feature F image ;

[0086] Generate text features according to the predefined safety detection items and the safety detection items of the vision-language large model based on the current scene type;

[0087] The text feature of the wearing part is F item and the text feature of the behavior part is F behave and the text feature of the scene safety factor is F risk .

[0088] S104. Generate a scene factor matrix, an object wearing matrix, and an object behavior matrix based on each of the text features, and calculate the probability P of danger based on the scene factor matrix, the object wearing matrix, and the object behavior matrix.

[0089] S105. Generate a risk assessment coefficient for the current scenario based on the probability P and the set risk coefficient.

[0090] In step S102, determining the scenario type and extracting the detection object includes the following:

[0091] Use the trained vision-language model to generate and analyze the scene content summary, determine the current scenario type, as well as the personnel and other objects that need to be detected for safety items in the scene, generate bounding boxes for all objects in the scene, and extract the image patches containing the detection objects from the scene image through the coordinate information of the bounding boxes.

[0092] In step S103, the scenario factor matrix is: S Image = F image T × F risk ;

[0093] The object wearing matrix is: S wear = F object T × F item ;

[0094] The object behavior matrix is: S behavet = F object T × F behave ;

[0095] The probability P = softmax(S);

[0096] where S = {S Image , S wear , S behavet}, when the probability P of S wear is greater than the threshold, it is considered that the object in the current scenario type wears relevant equipment and the text description is consistent, otherwise it means not wearing or wearing inconsistently; when S Image and S behavet are greater than the threshold, it is considered that there are relevant scene risk factors and dangerous behaviors.

[0097] In this embodiment, the engineering safety risk warning method for the vision-language large model further includes: performing time analysis on the scene data. Performing time analysis on the scene data includes:

[0098] Count the number of detection frames in the current time period in the scene data;

[0099] Mark the frames with a safety risk coefficient belonging to the unsafe level;

[0100] Count the number of frames belonging to the unsafe level and calculate the overall proportion. When the proportion is greater than the threshold θ, it is determined that a safety warning needs to be issued currently.

[0101] Schematically, the number of detected frames N in the current time period is counted and calculated using the following formula:

[0102] N = T / frame_duration;

[0103] Where T represents the length of the detection time period, and frame_duration represents the detection interval. When the safety risk coefficient of the current frame belongs to the unsafe level, it is marked as unsafe;

[0104] Count the number of frames belonging to the unsafe level and calculate the overall proportion. When it is greater than the threshold θ, it is determined that a safety warning is required at present. The calculation formula is as follows:

[0105]

[0106] In the above embodiment, the training of the visual language large model includes:

[0107] Obtain scenario data with safety risks, and perform manual analysis and description on the scenario data to construct a dataset for training;

[0108] Divide the safety risk objects in the scenario data. The safety risk objects include the scenario area and construction workers;

[0109] Define the safety detection items for the scenario area and construction workers, and set the risk coefficient;

[0110] Use the dataset to train and fine-tune the visual language large model.

[0111] Among them, the safety risk objects are divided into the scenario area (i.e., detecting the risk of the scenario area) and construction workers (i.e., detecting the risk of construction workers). Among them, the risk of construction workers is further divided into the risk of personnel equipment and the risk of personnel behavior. Define the safety detection items for the scenario area and construction workers, and set the risk coefficient; for example, in the construction area, open flames and dust are important factors affecting safety: in high-altitude operations, safety belts and safety helmets are items that must be worn, and climbing behavior is prohibited; in the scenario area where smoking is strictly prohibited, behaviors such as smoking and making phone calls are prohibited behaviors. These scenarios and factors can form specific safety detection items through definition and association.

[0112] In the above embodiment, the using the dataset to train and fine-tune the visual language large model includes:

[0113] For the visual language large model, given a pre-trained weight matrix W0 ∈ R m×n , select a low rank r, r << d, where d is the original dimension of the matrix;

[0114] Perform singular value decomposition on the pre-trained matrix to obtain a low-rank matrix of m singular values, where m is the dimension of the pre-trained weight W0. The decomposition formula is as follows:

[0115] SVD(W0) = U∑V T ;

[0116] U is an m×m orthogonal matrix, Σ is an m×n diagonal matrix, where the elements on the diagonal are singular values and the other elements are 0, and V T is the transpose of an n×n orthogonal matrix. S = [S1, S2, S3... S m is the singular value matrix of the pre-trained matrix W0;

[0117] Initialize two low-rank matrices A and B according to the weights of the pre-trained model. Compared with the original method where A is randomly initialized with Gaussian and B is initialized with zeros, this initialization has higher task performance, faster convergence during fine-tuning, and better preserves the knowledge in the original pre-trained model. A and B are initialized as follows:

[0118]

[0119] where U r , represents the sub-matrices of U and V corresponding to the selected r singular values, S T is a diagonal matrix with r singular values on the diagonal, r represents taking the square root of the S matrix; during training, W0 is frozen and does not receive gradient updates. A and B contain the fine-tuned weights, representing the differences to be added to the original weights. During inference, the fine-tuned weights are combined with the original pre-trained weights. For the input x, the modified forward pass can be expressed as: r For the input x, the modified forward pass can be expressed as:

[0120]

[0121] where r1 is the selected smaller rank, r2 is the selected larger rank, and α1 and α2 are trainable adjustment parameters; ΔW1 = B1A1 is the adapter matrix initialized with the singular values corresponding to the r1 rank, and ΔW2 = B2A2 is the adapter matrix initialized with the singular values corresponding to the r2 rank.

[0122] As Figure 2 shown, in another embodiment, an engineering safety risk warning system 200 based on a vision-language large model is provided. The engineering safety risk warning system 200 based on the vision-language large model includes: a data acquisition module 201, a scene analysis module 202, a visual prompt module 203, a safety detection module 204, and a risk assessment module 205.

[0123] The data acquisition module 201 is used to extract the scenario data of the project from the monitoring platform of the project.

[0124] The scenario analysis module 202 is used to obtain the visual content summary in the scenario data by using the trained visual language large model, judge the current scenario type according to the visual content summary, and extract the objects that need to be detected for safety items in the current scenario type; the objects to be detected include the scenario area and construction workers, and the detection of construction workers includes personnel equipment and personnel behavior.

[0125] The visual prompt module 203 is used to detect the safety factors required for the scenario and the objects that need to be detected for safety items in the scenario.

[0126] The safety detection module 204 is used to generate a scenario factor matrix, an object wearing matrix, and an object behavior matrix based on each of the text features, and calculate the probability P of danger based on the scenario factor matrix, the object wearing matrix, and the object behavior matrix.

[0127] The risk assessment module 205 is used to generate a risk assessment coefficient for the current scenario according to the probability P and the set danger coefficient.

[0128] Further, the engineering safety risk warning system 200 based on the visual language large model further includes a time analysis module 206, and the time analysis module 206 is used to: count the number of detected frames in the current time period in the scenario data;

[0129] Mark the frames with a safety risk coefficient belonging to the unsafe level;

[0130] Count the number of frames belonging to the unsafe level and calculate the overall proportion. When the proportion is greater than the threshold θ, it is determined that a safety warning needs to be issued currently.

[0131] In the above embodiment, the scenario factor matrix is: S Image = F image T × F risk ;

[0132] The object wearing matrix is: S wear = F object T × F item ;

[0133] The object behavior matrix is: S behavet = F object T × F behave ;

[0134] The probability P = softmax(S);

[0135] where S = {S Image , S wear , S behavet}, when the probability P of S wear is greater than the threshold, it is considered that the object in the current scene type wears relevant equipment and the text description matches, otherwise it means not wearing or wearing inconsistently; when S Image and S behavet are greater than the threshold, it is considered that there are relevant scene risk factors and dangerous behaviors exist.

[0136] The training of the visual language large model includes:

[0137] Obtain scene data with security risks, and perform manual analysis and description on the scene data to construct a dataset for training; divide the security risk objects in the scene data, and the security risk objects include scene areas and construction workers;

[0138] Define the security detection items for the scene area and construction workers, and set the risk coefficients;

[0139] Use the dataset to train and fine-tune the visual language large model.

[0140] In the above embodiment, aiming at the limitations of the existing vision-based monitoring system and object detection model, this patent proposes a security risk warning method and system. This patent accurately interprets scene information by integrating an image caption model, and integrates a large language model to bridge the semantic gap between visual and text data, thus solving the challenges of various security requirements; in the fine-tuning stage, an integrated adapter that is not initialized with a singular value matrix is introduced to reduce the uncertainty of model prediction; in the security project detection stage, image patches of different detection items and their corresponding prompt words are used for feature extraction and matrix association, so as to judge whether there are various risk factors and calculate the risk assessment coefficient; finally, the result statistical analysis based on time series is performed on the scene data to further reduce the inaccuracy of prediction.

[0141] Finally, it should be noted that although the above embodiments have been described in the text of the specification and drawings of this application, the patent protection scope of this application cannot be limited thereby. Any technical solutions obtained by equivalent structure or equivalent process substitution or modification using the content recorded in the text of the specification and drawings of this application based on the essential concept of this application, as well as those directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, are all included in the patent protection scope of this application.

Claims

1. An engineering safety risk warning method based on a vision-language large model, characterized in that, Including the following steps: Construct a vision-language large model and train the vision-language large model; Extract the scene data of the project from the project monitoring platform, obtain the visual content summary in the scene data with the trained vision-language large model, judge the current scene type according to the visual content summary, and extract the objects that need to be detected for safety items in the current scene type; the detected objects include the scene area and construction workers, and the detection of construction workers includes personnel equipment and personnel behavior; Detect the safety factors required for the scene and the objects that need to be detected for safety items in the scene, including: Analyze the visual content summary to obtain scene information, query the vision-language large model with the scene information to obtain the items that construction workers should wear and prohibited behaviors in the current scene type; Generate visual features corresponding to the items to be worn and prohibited behaviors in the current scene type; Generate bounding boxes for all objects in the scene, extract image blocks containing the detected objects from the scene image according to the coordinate information of the bounding boxes, perform feature extraction using the pre-trained VLM model feature extraction module, and generate feature matrices F for multiple detected objects object and scene feature F image ; Generate text features according to the predefined safety detection items and the safety detection items of the vision-language large model based on the current scene type; The text feature of the wearable part is F item , and the text feature of the behavior part is F behave , and the text feature of the scene safety factor is F risk ; Generate a scene factor matrix, an object wearing matrix and an object behavior matrix based on each of the text features, and calculate the probability P of danger based on the scene factor matrix, the object wearing matrix and the object behavior matrix; Generate a risk assessment coefficient for the current scene according to the probability P and the set danger coefficient.

2. The engineering safety risk early warning method based on the vision-language large model according to claim 1, wherein The scenario factor matrix is: S Image = F image T × F risk ; The object wearing matrix is: S wear = F object T × F item ; The object behavior matrix is: S behavet = F object T × F behave ; The probability P = softmax(S); where S = {S Image , S wear , S behavet}, when the probability P of S wear is greater than the threshold, it is considered that the object wears relevant equipment and the text description in the current scene type match, otherwise it means not wearing or wearing inconsistently; when S Image and S behavet are greater than the threshold, it is considered that there are relevant scene risk factors and dangerous behaviors exist.

3. The engineering safety risk early warning method based on the vision-language large model according to claim 2, wherein It also includes time analysis of the scene data, including: Count the number of detected frames in the current time period of the scene data; Mark the frames with a safety risk coefficient belonging to the unsafe level; Count the number of frames belonging to the unsafe level and calculate the overall proportion. When the proportion is greater than the threshold θ, it is judged that safety warning is required currently.

4. The engineering safety risk early warning method based on the vision-language large model according to claim 1, characterized in that The training of the vision-language large model includes: Obtain the scene data with safety risks, and conduct manual analysis and description of the scene data to construct a dataset for training; Divide the safety risk objects in the scene data, and the safety risk objects include the scene area and construction workers; Define the safety detection items for the scene area and construction workers and set the risk coefficient; Use the dataset to train and fine-tune the vision-language large model.

5. The engineering safety risk early warning method based on the vision-language large model according to claim 4, wherein, The use of the dataset to train and fine-tune the vision-language large model includes: For the visual language large model, given a pre-trained weight matrix W0 ∈ R m×n , a low rank r is selected, where r << d and d is the original dimension of the matrix; Perform singular value decomposition on the pre-trained matrix to obtain a low-rank matrix of m singular values, where m is the dimension of the pre-trained weight W0, and the decomposition formula is as follows: SVD(W0) = U∑V T ; U is an m×n orthogonal matrix, Σ is an m×n diagonal matrix, where the elements on the diagonal are singular values and the other elements are 0, and V T is the transpose of an n×n orthogonal matrix, and S = [S1, S2, S3......S m is the singular value matrix of the pre-trained matrix W0; Initialize two low-rank matrices A and B according to the weights of the pre-trained model. Compared with the original method where A is randomly initialized with Gaussian and B is initialized with zero, A and B are initialized as follows; Among them, U r , represent the submatrices of U and V corresponding to the selected r singular values T , S r is a diagonal matrix with r singular values on the diagonal, represents the square root of the S r matrix; During the training process, W0 is frozen and does not receive gradient updates. A and B contain the fine-tuned weights, indicating the differences to be added to the original weights. During the inference process, the fine-tuned weights are combined with the original pre-trained weights. For the input x, the modified forward pass can be expressed as: Where r1 is the smaller rank selected, r2 is the larger rank selected, and α1 and α2 are trainable adjustment parameters; ΔW1 = B1A1 is the adapter matrix initialized with the singular values corresponding to the r1 rank, and ΔW2 = B2A2 is the adapter matrix initialized with the singular values corresponding to the r2 rank.

6. An engineering safety risk early warning system based on a vision-language large model, characterized in that, Including: A data acquisition module for extracting the scene data of the project from the monitoring platform of the project; A scene analysis module for obtaining the visual content summary in the scene data using the trained visual language large model, judging the current scene type according to the visual content summary, and extracting the objects that need to be detected for safety items of the current scene type; the detected objects include the scene area and construction workers, and the detection of construction workers includes personnel equipment and personnel behavior; A visual prompt module for detecting the safety factors required for the scene and the objects that need to be detected for safety items in the scene; A safety detection module for generating a scene factor matrix, an object wearing matrix, and an object behavior matrix based on each text feature, and calculating the probability P of danger occurring based on the scene factor matrix, the object wearing matrix, and the object behavior matrix; A risk assessment module for generating a risk assessment coefficient of the current scene according to the probability P and the set danger coefficient.

7. The engineering safety risk early warning system based on the vision-language large model according to claim 6, characterized in that, It further includes a time analysis module for counting the number of detected frames in the current time period of the scene data; Marking the frames with a safety risk coefficient belonging to the unsafe level; Counting the number of frames belonging to the unsafe level and calculating the overall proportion. When the proportion is greater than the threshold θ, it is determined that a safety warning is required currently.

8. The engineering safety risk early warning system based on the vision-language large model according to claim 6, wherein, The scenario factor matrix is: S Image = F image T × F risk ; The object wearing matrix is: S wear = F object T × F item ; The object behavior matrix is: S behavet = F object T × F behave ; The probability P = softmax(S); where S = {S Image , S wear , S behavet}, when the probability P of S wear is greater than the threshold, it is considered that the object wears relevant equipment and the text description in the current scene type match; otherwise, it means not wearing or wearing inconsistently; when S Image and S behavet are greater than the threshold, it is considered that there are relevant scene risk factors and dangerous behaviors exist.

9. The engineering safety risk early warning system based on the vision-language large model according to claim 6, wherein The training of the visual language large model includes: Obtaining the scene data with safety risks, and performing manual analysis and description on the scene data to construct a dataset for training; Dividing the safety risk objects in the scene data, and the safety risk objects include the scene area and construction workers; Defining the safety detection items for the scene area and construction workers, and setting the risk coefficient; Using the dataset to train and fine-tune the visual language large model.

Citation Information

Cited By

  • Human-computer interaction method and system based on AI large model

    CN121742373A