Building risk prediction management and control method and system based on multi-modal LLM

By integrating multimodal data through multimodal LLM technology, real-time prediction and dynamic management of building risks have been achieved, solving the problems of data silos and response delays in existing technologies, improving the accuracy of risk prediction and management efficiency, and enhancing the system's adaptability.

CN120912376APending Publication Date: 2025-11-07TIANJIN UNIV

Patent Information

Application Number
CN202511342868.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing building risk prediction technologies suffer from problems such as data silos, response delays, and limited analytical capabilities. They cannot effectively integrate multimodal data, resulting in incomplete risk prediction and low management efficiency, and lack of forward-looking identification and dynamic adaptability of potential risks.

Method used

A building risk prediction and control method based on multimodal LLM is adopted. The multimodal feature extraction module parses the multimodal data stream into structured feature vectors. The multimodal LLM inference engine is used to perform cross-modal semantic fusion and risk coupling analysis to generate potential risk identifiers and level assessments. The control rules of the building safety code library are dynamically matched to drive the on-site execution equipment to implement control actions. Feedback data is collected in real time and the model parameters and rule weights are updated.

Benefits of technology

It significantly improves the accuracy and comprehensiveness of risk prediction, enables real-time identification and customized response to potential risks, enhances system compatibility and execution reliability, optimizes resource allocation, and improves the resilience and self-evolution capabilities of building safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912376A_ABST
    Figure CN120912376A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal LLM-based building risk prediction management and control method and system, and relates to the technical field of building safety management, and the method comprises the steps: analyzing sensor data, a text report, an image video and a voice instruction of a building construction site through a multi-modal feature extraction module, and generating a structured feature vector set; performing cross-modal semantic fusion and risk coupling analysis by using a multi-modal LLM inference engine to generate a potential risk identification set and a risk level assessment result; dynamically matching a management and control rule of the building safety specification library based on the risk identifier, and outputting a strategy set consisting of an equipment regulation and control instruction, a personnel early warning notification and a regional management and control suggestion; driving a field execution device to implement a control action, and collecting a multi-modal feedback data stream; and calculating a strategy execution efficiency index through a closed-loop optimization module, dynamically updating LLM model parameters and rule weights, and forming a self-adaptive optimization link. The system correspondingly comprises a multi-modal feature extraction and fusion module, an LLM inference engine module, a dynamic strategy generation module, an execution feedback module and a closed-loop optimization module. According to the method, the problems of key feature omission and risk response lag in traditional single-mode analysis are solved, and the risk prediction accuracy and the management and control real-time performance are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of safety monitoring, in particular to a building risk prediction management and control method and system based on a multi-modal LLM. BACKGROUND

[0002] With the acceleration of urbanization and the complex of building functions, complex building forms such as high-rise civil buildings, large complexes and industrial plants are emerging. Their internal structures are intertwined, equipment systems are dense, and personnel flow is frequent. The risk types they face show diversified characteristics, including sudden safety risks such as fire, gas leakage, and electrical failure, and implicit progressive risks such as structural aging, equipment wear and tear, and personnel violation. Currently, artificial intelligence technology is undergoing a revolutionary evolution from single-modal processing to multi-modal fusion, and multi-modal large language models (MM-LLMs) have become the core technology to support intelligent decision-making in complex scenarios. Building risk management and control technology has gradually developed from early manual inspection and single-sensor early warning to a collaborative management and control mode based on the Internet of Things and spatial modeling, such as the invention patent CN202410180559.0 "A modern building information management and control system and method based on the Internet of Things". Through a three-dimensional scanner to build a building space model, combined with the monitoring data of fire warning Internet of Things sensors, it locks the abnormal area and builds a risk source abnormal interference chain, and then generates a disaster control adjustment area to realize fire spread prediction and local gas pipeline control. However, there is still a lack of multi-modal risk data fusion capability in the full-dimensional risk management and control of complex buildings. Existing technology only relies on the quantitative monitoring of temperature, gas concentration, and flow data by Internet of Things sensors, and cannot effectively fuse unstructured multi-modal data in building operation and maintenance, such as text logs recorded by inspection personnel, image data captured by cameras, or real-time voice information reported by personnel. These unstructured data contain a lot of implicit risk information, but due to the lack of semantic understanding and cross-modal correlation capabilities, they are missed by existing systems, resulting in "fragmentation" blind spots in risk information collection. In addition, the existing technology lacks foresight and reasoning ability in risk prediction. The existing risk prediction relies on the "abnormal trigger-rule analysis" mode, such as CN202410180559.0, which needs to build a risk chain through a pre-set formula after the sensor detects an anomaly. This is essentially a "post-response" prediction that cannot predict potential risks in advance. The risk analysis is limited to fixed rules and cannot identify complex risks caused by the coupling of multiple factors, such as extreme high temperature, insufficient ventilation, or equipment overload. The reasoning dimension is relatively single. Secondly, the intelligence and dynamic adaptability of the control strategy are weak. The existing technology's control measures are mostly pre-set processes, lacking the ability to dynamically adjust according to real-time multi-modal data. For example, when the system receives "fire warning", "person trapped voice report", and "elevator failure image" at the same time, the existing system can only execute a single control according to priority, and cannot coordinate rescue paths, evacuation plans, and equipment shutdown sequences through multi-objective reasoning, resulting in low control efficiency.Finally, the multi-type risk collaborative management system is missing, and the existing technology focuses on the associated risks of fire and gas pipelines, but does not cover other core risk types such as structural safety, electrical failure, and personnel concentration. Moreover, the coupling relationship between different risks is not analyzed, resulting in a "fragmentation" of management and control, which cannot form a global risk coordination mechanism and is prone to "trade-off" management loopholes.

[0003] Therefore, there is an urgent need for a building risk prediction management and control method and system based on multi-modal LLM to at least solve the above problems. SUMMARY

[0004] The purpose of the present application is to provide a building risk prediction management and control method and system based on multi-modal LLM to solve the problems of data silos, response delays, and single analysis capabilities in the prior art. The specific technical solutions are as follows:

[0005] The present application provides a building risk prediction management and control method based on multi-modal LLM, comprising:

[0006] Step 1: According to the multi-modal data stream collected in real time at the construction site, the multi-modal feature extraction module is used to analyze the structured feature vector set and output it to the multi-modal fusion engine;

[0007] Step 2: According to the structured feature vector set, the multi-modal LLM inference engine is used for cross-modal semantic fusion and risk coupling analysis to generate a set of potential risk identifiers and risk level evaluation results, wherein the risk types include physical facility abnormalities, behavior specification violations, and environmental dynamic threats;

[0008] Step 3: According to the set of potential risk identifiers and risk level evaluation results, the control rules of the building safety specification library are dynamically matched to output a set of risk control strategies;

[0009] Step 4: According to the set of risk control strategies, the on-site execution equipment is driven to implement control actions, and multi-modal feedback data streams are collected;

[0010] Step 5: According to the feedback data stream, the strategy execution performance indicators are calculated, and the LLM model parameters and rule weights are dynamically updated.

[0011] Further, the step 1 includes: using a time series feature extractor to process sensor time series data, a visual feature extractor to process image and video data, a text feature extractor to process text reports, and a voice feature extractor to process voice instructions; and the multi-modal feature vectors are timestamped and device ID bound to generate a structured feature vector set.

[0012] Further, the step 2 comprises: calculating the correlation weight of different modal feature vectors through a cross-modal attention mechanism to generate a fused semantic representation; and mapping the fused semantic representation to a building risk knowledge graph based on a risk propagation model of a graph neural network to analyze risk coupling effects.

[0013] Further, the step 3 comprises: matching the safety specification library rules by using a Rete algorithm, and dynamically adjusting the execution parameters according to the risk level; and ordering the control instructions according to the priority weight through a conflict resolution algorithm.

[0014] Further, the step 5 comprises: updating the LLM attention layer parameters based on a reinforcement learning reward function; and dynamically adjusting the rule weight according to the rule historical performance score, and triggering a disabling mechanism if the performance score is continuously low.

[0015] The application also relates to a building risk prediction and control system based on the multi-modal LLM, comprising:

[0016] A multi-modal feature extraction and fusion module is used for analyzing multi-modal data streams and generating a set of structured feature vectors.

[0017] A multi-modal LLM inference engine module is used for performing cross-modal semantic fusion and risk coupling analysis, and outputting risk identification and level evaluation.

[0018] A dynamic strategy generation module is used for adapting safety specification library rules and generating a set of risk control strategies.

[0019] An execution feedback module is used for driving on-site execution equipment to implement control actions and collect feedback data.

[0020] A closed-loop optimization module is used for updating LLM model parameters and rule weights.

[0021] Further, the multi-modal feature extraction and fusion module comprises:

[0022] A time sequence feature extractor, a visual feature extractor, a text feature extractor and a speech feature extractor.

[0023] A multi-modal fusion engine receives the set of feature vectors through an API interface.

[0024] Further, the multi-modal LLM inference engine module is built-in with a building risk knowledge graph, defines node entities and risk transmission relationships, and outputs risk identification including risk type, location, confidence and timestamp.

[0025] Further, the dynamic strategy generation module calls a building safety specification library, the rule types include equipment regulation rules, personnel warning rules and area control rules; the strategy set is packaged in JSON or Protocol Buffers format; the performance evaluation model of the closed-loop optimization module is a multi-layer perception machine, and an output comprehensive performance score is output, and a strategy gradient algorithm is used to update the LLM model parameters.

[0026] The application also relates to an electronic device comprising a processor and a memory, the memory storing a computer program which, when executed by the processor, implements the multi-modal LLM-based building risk prediction and control method as described.

[0027] The method and system integrate multi-modal data processing and large language model (LLM) technology to realize real-time prediction and dynamic control of building construction site risks, and the core beneficial effects include: significantly improving the accuracy and comprehensiveness of risk prediction, effectively reducing the false alarm rate and identifying derived risks (such as the correlation between equipment inclination abnormalities and environmental threats) through cross-modal semantic fusion and knowledge graph-driven risk coupling analysis; efficiently generating customized response strategies based on dynamic rule matching and priority sorting to optimize the execution efficiency of equipment regulation, personnel warning and area control instructions; enhancing system compatibility and execution reliability, supporting multi-protocol industrial device control (such as Modbus and MQTT) and standby link switching to ensure instruction reachability; realizing intelligent allocation of resources by adaptively updating LLM model parameters and rule weights through a closed-loop optimization module to continuously improve risk control performance; ultimately building an end-to-end closed-loop link to comprehensively improve the resilience, real-time performance and self-evolution capability of building safety management, and reduce the incidence of secondary accidents.

[0028] The technical solutions of the application will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0029] The accompanying drawings are used to provide a further understanding of the application, and constitute a part of the specification, together with the embodiments of the application, to explain the application, and do not constitute a limitation on the application. In the drawings:

[0030] Figure 1 a schematic diagram of the multi-modal LLM-based building risk prediction and control method in the embodiments of the application;

[0031] Figure 2 a schematic diagram of the multi-modal LLM-based building risk prediction and control system in the embodiments of the application. DETAILED DESCRIPTION

[0032] The preferred embodiments of the present application are described below with reference to the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0033] The present embodiment provides a building risk prediction management method based on a multi-modal LLM, as shown in the figure, comprising: Figure 1

[0034] Step 1: According to the multi-modal data stream collected in real time at the construction site, the multi-modal feature extraction module is used to analyze the structured feature vector set and output to the multi-modal fusion engine;

[0035] Among them, the multi-modal data stream includes: sensor data (such as time series data collected by inclination sensors, stress and strain sensors, temperature and humidity sensors, and noise sensors), text reports (such as structured inspection reports and unstructured safety logs), image and video (such as 2D / 3D visual data collected by fixed monitoring cameras and unmanned aerial vehicle inspection equipment), and voice instructions (such as safety officer intercom voice and equipment operation instructions);

[0036] Among them, the multi-modal feature extraction module is a combination of feature encoders based on deep neural networks, which is used to map raw data of different modalities to a unified high-dimensional feature space; Specifically, it includes:

[0037] The time series feature extractor (such as one-dimensional convolutional neural network or long short-term memory network) is used to extract time domain feature vectors for sensor time series data;

[0038] The visual feature extractor (such as convolutional neural network) is used to extract spatial feature vectors for image and video data;

[0039] The text feature extractor (such as a pre-trained language model based on Transformer) is used to extract semantic feature vectors for text reports;

[0040] The speech feature extractor (such as an automatic speech recognition model combined with a text feature extractor, or an end-to-end acoustic model) is used to extract speech semantic feature vectors for voice instructions;

[0041] Among them, the structured feature vector set is a standardized data set formed after the time stamp alignment and device ID binding of different modal feature vectors, and the feature dimension is determined by the specific neural network model adopted (such as the image feature vector dimension is 2048, and the text feature vector dimension is 768);

[0042] Among them, the multi-modal fusion engine is a software module that receives the structured feature vector set and performs cross-modal feature fusion calculation, and receives the feature vector set input through the message queue or API interface.

[0043] ​Step 2: Based on the structured feature vector set, cross-modal semantic fusion and risk coupling analysis are performed by a multi-modal LLM inference engine to generate a set of potential risk identifications and risk level assessment results, where the risk types include physical facility abnormalities, behavior specification violations, and environmental dynamic threats.

[0044] wherein the multi-modal LLM inference engine is a deep learning model based on the Transformer architecture, which obtains multi-modal understanding and inference capabilities through pre-training and fine-tuning, specifically receiving the structured feature vector set as input and processing through the following technical means:

[0045] Cross-modal semantic fusion is achieved using a cross-modal attention mechanism (Cross-Modal Attention) to calculate the correlation weights between different modal feature vectors and generate unified fusion semantic representations; for example, calculating the attention weights of image features and text features to determine the consistency of text descriptions and visual scenes.

[0046] Risk coupling analysis uses a risk propagation model based on graph neural networks to map the fused semantic representations to a building risk knowledge graph and analyze the associations and transmission paths between different risk factors; specifically including calculating the coupling effects between physical facility status, personnel behavior, and environmental factors.

[0047] wherein the set of potential risk identifications is a standardized representation of specific risk instances identified by the inference engine, and each risk identification includes: risk type, risk location, risk feature vector, confidence, and timestamp.

[0048] Risk types include: physical facility abnormalities (such as scaffold inclination exceeding limits, support structure stress abnormalities), behavior specification violations (such as not wearing safety helmets, violating operation equipment), and environmental dynamic threats (such as strong wind warnings, heavy rain warnings).

[0049] wherein the risk level assessment result is a quantitative risk assessment output based on a multi-classification model, which takes the fused risk semantic representation as input and outputs risk level labels (such as no risk, low risk, medium risk, high risk) and corresponding probability distributions; the risk assessment model is supervised trained by historical accident data and expert annotations.

[0050] Step 3: Based on the set of potential risk identifications and risk level assessment results, the dynamic strategy generation module adapts the control rules in the building safety specification library to output a set of targeted risk control strategies.

[0051] wherein the dynamic strategy generation module is a software decision system based on rule engines and optimization algorithms, which receives the set of potential risk identifications and risk level assessment results as input and processes through the following technical means:

[0052] Call the building safety regulation library query interface to obtain a set of control rules that match the current risk type, risk level, and construction site context;

[0053] Use a multi-constraint rule matching algorithm (such as the Rete algorithm or a graph-based pattern matching algorithm) to perform rule adaptation and calculate the matching degree of the rule premise conditions and risk characteristics;

[0054] For rules that match successfully, generate a preliminary set of control instructions based on the confidence level in the risk level assessment results and the rule weight;

[0055] The building safety regulation library is a structured database that stores standardized safety management rules, and its rules are represented using machine-readable logical expressions. Each rule contains:

[0056] Rule ID, rule type, applicable risk type, trigger condition (e.g., risk level ≥ medium risk and confidence level ≥ 0.8), control action, execution parameter, and priority weight;

[0057] Rule types include: device control rules (e.g., turning off specific devices, adjusting operating parameters), personnel warning rules (e.g., sending voice alerts, pushing text notifications), and area control rules (e.g., setting electronic fences, restricting personnel entry);

[0058] The risk control strategy set is a dynamically generated set of standardized instructions, and each strategy contains:

[0059] Strategy ID, target device / personnel identifier, control action type, execution parameter, trigger time window, and priority;

[0060] Control action types include: device control instructions (e.g., "Tower T001: Stop running," "Pump station P005: Power reduced to 30%"), personnel warning notifications (e.g., "Area A3: Immediately evacuate" voice broadcast, "Personnel P088: Please wear a safety helmet" mobile push), and area control suggestions (e.g., "Lock B2 area," "Start emergency lighting").

[0061] Step 4: Based on the risk control strategy set, drive the on-site execution equipment (including sensor networks, mechanical controllers, and voice broadcast systems) to implement control actions through the execution feedback module, and collect real-time multi-modal feedback data streams after execution;

[0062] The execution feedback module is a software and hardware cooperative system based on industrial Internet of Things protocols and device control APIs. It receives the risk control strategy set as input and processes it through the following technical means:

[0063] The control action types and execution parameters in the strategy set are analyzed and converted into control instructions recognizable by specific execution devices;

[0064] The control instructions are sent to the field execution devices through the device driver interface to drive the devices to perform specific control actions;

[0065] The field execution devices include:

[0066] Sensor network: various sensors (such as inclination sensors, stress sensors, vision sensors) for monitoring environment and device status, which access through Modbus, OPC UA, etc. Industrial protocols;

[0067] Mechanical controller: control device (such as PLC controller, relay module, motor driver) for executing physical action, which receives instructions through digital / analog output signal or EtherCAT, Profinet, etc. Real-time Ethernet protocol;

[0068] Voice broadcast system: field device (such as IP broadcast terminal, digital power amplifier device) for issuing voice warning, which receives text-to-speech (TTS) instructions through SIP protocol or audio stream API;

[0069] The specific implementation of the control action includes:

[0070] Sending device control instructions to the mechanical controller (such as sending "stop running" instruction to the tower crane PLC, the instruction format is Modbus RTU protocol function code 06H write to the holding register);

[0071] Sending warning notification to the voice broadcast system (such as converting the text "Area A3 evacuate immediately" to G.711 encoded audio stream through TTS engine, and broadcasting to the specified area terminal through SIP protocol);

[0072] Sending collection strategy adjustment instructions to the sensor network (such as adjusting the sampling frequency of the vision sensor from 1Hz to 5Hz to focus on monitoring high-risk areas);

[0073] The multi-modal feedback data stream is the response data collected by the field execution devices in real time after executing the control action, including:

[0074] Device state feedback data (such as current inclination of tower crane, motor running current, device start / stop state);

[0075] Environmental response data (such as personnel evacuation video stream, area temperature and humidity change, voice instruction confirmation signal);

[0076] Execution efficiency raw data (such as instruction issue timestamp, device response delay, action execution completion flag).

[0077] Step 5: According to the multi-modal feedback data stream, the strategy execution performance indicators are calculated by the closed-loop optimization module, and the model parameters of the multi-modal LLM inference engine and the rule weights of the security specification library are dynamically updated to form an adaptive optimization link.

[0078] Among them, the closed-loop optimization module is a parameter optimization system based on reinforcement learning and online learning algorithm, which receives the multi-modal feedback data stream as input and processes it through the following technical means:

[0079] Extract the performance feature vector related to strategy execution from the feedback data stream, including device response delay, action execution completion rate, risk state change rate, and warning response compliance;

[0080] Calculate the strategy execution performance indicators using the performance evaluation model, which is a supervised trained multi-layer perceptron (MLP), and the output includes comprehensive performance score and each sub-indicator weight;

[0081] Among them, the strategy execution performance indicators are standardized measurement values that quantify the effectiveness of policy execution, including:

[0082] Instant performance indicators: instruction response time, action execution success rate;

[0083] Risk control indicators: risk level decline rate, secondary risk occurrence rate;

[0084] Comprehensive performance score: a percentage score calculated based on weighted summation, used to evaluate the effectiveness of the strategy as a whole;

[0085] Among them, the dynamic update of the model parameters of the multi-modal LLM inference engine is realized by using the online learning algorithm, which specifically includes:

[0086] Construct a reinforcement learning reward function based on the strategy performance indicators, which is configured to: comprehensively consider the response efficiency of instructions, the success rate of action execution, and the control effect of risk level, to generate a comprehensive reward value. Among them, the shorter the response time, the higher the action execution success rate, and the more significant the risk level decline, the higher the reward value generated;

[0087] Use the policy gradient algorithm to calculate the model parameter update gradient, for example, it can be a proximal policy optimization (PPO) or REINFORCE algorithm;

[0088] Update the attention layer parameters and output layer weights of the multi-modal LLM through backpropagation, and according to the policy gradient, use the gradient ascent algorithm to iteratively update the model parameters of the multi-modal LLM inference engine to maximize the future expected cumulative reward value;

[0089] The rule weight of the dynamic updating of the security specification library is realized by a statistical learning method, and specifically includes:

[0090] According to the effectiveness data of the statistical rule of the policy execution effect, the weight of the rule in the security specification library is dynamically adjusted according to the historical performance score of the rule. The higher the performance score is, the greater the weight adjustment is; if the performance score is continuously low, the weight will be adjusted downward accordingly;

[0091] For continuously failed rules (performance score lower than threshold for N times), a rule disable flag is triggered and an artificial review request is pushed.

[0092] The new rules are automatically extracted from successful strategy cases by a rule mining algorithm and are stored in the library after confidence screening.

[0093] The working principle and beneficial effects of the above technical solution are: the technical solution constructs an end-to-end closed-loop building risk prediction and control system, which works from real-time collection and structured feature extraction of multi-modal data, cross-modal semantic fusion and risk coupling analysis by a multi-modal LLM inference engine, accurate identification of potential risks and assessment of grades; then, a dynamic strategy generation module adapts security specification rules based on risk identification, and outputs targeted control strategies; an execution feedback module drives on-site devices to implement actions and collects multi-modal feedback; finally, a closed-loop optimization module dynamically updates model parameters and rule weights using feedback data, forming a self-adaptive optimization link. The beneficial effects include significantly improving the accuracy and comprehensiveness of risk prediction, efficiently generating customized response strategies, enhancing execution reliability and system compatibility, realizing intelligent allocation of resources and continuous self-evolution ability, thereby comprehensively improving the safety management efficiency and resilience of the construction site.

[0094] In one embodiment, step 1: according to the multi-modal data stream collected in real time from the construction site, the multi-modal feature extraction module is used to analyze the structured feature vector set and output to the multi-modal fusion engine, including:

[0095] 101. One-dimensional convolutional neural network (Conv1D) is used to extract time domain feature vectors from the time series data of the tilt sensor, and the convolution kernel size is set to 5x1 with a step of 2 to capture the device tilt mutation feature;

[0096] 102. ResNet-50 convolutional neural network is used to extract spatial feature vectors from the monitoring video data, focusing on the visual features of the scaffold connection point and the safety helmet wearing area;

[0097] 103. The multi-modal features are synchronized to a unified time axis by a timestamp alignment algorithm (such as dynamic time warping DTW), eliminating the millisecond-level time sequence deviation of sensor and video collection.

[0098] The working principle and beneficial effects of the above technical solution are: different modal data is processed by the heterogeneous neural network, solving the problem of missing key features in traditional single-modal analysis; timestamp alignment ensures the real-time nature of risk analysis, avoiding misjudgment due to asynchronous data (such as loss of relevance between personnel violation behavior and device state change).

[0099] In one embodiment, step 2: according to the structured feature vector set, cross-modal semantic fusion and risk coupling analysis are performed by a multi-modal LLM inference engine to generate a set of potential risk identifications and risk level evaluation results, wherein the risk types include physical facility abnormalities, behavior specification violations, and environmental dynamic threats, including:

[0100] 201. Input the aligned multi-modal feature vectors into a cross-modal encoder to generate context-aware fusion features;

[0101] 202. Input the fusion features into a risk classifier to identify potential risk types and generate initial risk identifications;

[0102] 203. Perform relevance verification and coupling analysis on the initial risk identifications based on a risk knowledge graph to correct false positives and supplement coupled risks;

[0103] 204. Calculate the final risk level according to the risk severity and occurrence probability, and output the standardized evaluation results.

[0104] In one embodiment, step 2: according to the structured feature vector set, cross-modal semantic fusion and risk coupling analysis are performed by a multi-modal LLM inference engine to generate a set of potential risk identifications and risk level evaluation results, wherein the risk types include physical facility abnormalities, behavior specification violations, and environmental dynamic threats, including:

[0105] 211. Build a building risk knowledge graph: nodes include entities such as "scaffold inclination > 5°" and "no safety helmet", and edges define risk transmission relationships (such as "inclination exceeding limit → collapse risk → impact radius 10m");

[0106] 212. Cross-modal attention mechanism calculates the relevance weight of image features and text reports, for example, when video detects that a high-altitude worker is not wearing a safety belt, and the safety log lacks the worker's training record, the weight is increased to 0.92;

[0107] 213. The risk propagation model based on graph neural network outputs coupled risk identifications, such as "inclination anomaly + heavy rain warning" triggering "slope landslide" secondary derivative risk.

[0108] The working principle and beneficial effects of the above technical solution are: the knowledge graph structures discrete risk factors, solving the defect of traditional methods ignoring risk correlation; the attention mechanism realizes multi-modal evidence cross-validation, reducing the false positive rate; the risk propagation model predicts derived risks and triggers control strategies in advance.

[0109] In one embodiment, step 3: according to the potential risk identification set and risk level evaluation result, the control rules in the building safety specification library are adapted by the dynamic strategy generation module, and a targeted risk control strategy set is output, including:

[0110] 301. Analyze the risk type, location and feature vector in the risk identification set, and construct a rule query request;

[0111] 302. Retrieve matching rules from the safety specification library, and calculate the semantic similarity of rule conditions and risk characteristics;

[0112] 303. Adjust the rule execution parameters according to the risk level (such as: higher risk corresponds to more stringent control actions);

[0113] 304. Eliminate instruction conflicts by a strategy optimization algorithm, and sort according to priority weight;

[0114] 305. Output the standardized strategy set to the execution interface, and the strategy format adopts JSON or Protocol Buffers serialization protocol.

[0115] In one embodiment, the dynamic strategy generation module of step 3 performs the following operations:

[0116] 311. Match the safety specification library rules using the Rete algorithm: if the risk identification is "scaffold inclination > 8° and confidence ≥ 0.9", trigger rule ID R003 (control action: block a 15m radius area + tower crane stop);

[0117] 312. Dynamically adjust parameters according to risk level: under high risk level, the frequency of personnel warning notification is increased from 1 time / minute to 3 times / minute;

[0118] 313. Conflict resolution algorithm prioritizes high-weight strategies, such as "personnel evacuation" instruction weight higher than "device power reduction".

[0119] The working principle and beneficial effects of the above technical solution are: Rete algorithm has higher efficiency than traditional SQL query; dynamic parameter adjustment realizes risk response hierarchical management, avoiding resource waste; conflict resolution ensures priority execution of key instructions, reducing secondary accidents caused by control delay.

[0120] In one embodiment, step 4: According to the risk control strategy set, the feedback module is executed to drive the on-site execution equipment (including sensor network, mechanical controller and voice broadcast system) to implement control actions, and real-time collection of multi-modal feedback data stream after execution is performed, including:

[0121] 401. Convert the standardized strategy set into device-specific protocol messages through an instruction encoder (e.g., convert JSON strategy into Modbus TCP frame);

[0122] 402. Distribute control instructions to target devices through an industrial Internet of Things platform (e.g., MQTT or Apache Kafka-based message middleware);

[0123] 403. Monitor instruction execution status, and trigger retry mechanism or degradation processing when timeout or failure occurs;

[0124] 404. Synchronously collect multi-modal sensor data after execution, and associate with uniform timestamp and strategy ID;

[0125] 405. Encapsulate the feedback data stream into a standardized format (e.g., ApacheAvro serialized data) and transmit it to the closed-loop optimization module.

[0126] In one embodiment, step 4 includes:

[0127] 411. Convert the JSON format strategy set into industrial protocol instructions: such as "area blockade" instruction encoding into Modbus TCP function code 10H, write into PLC controller register address 0x0A;

[0128] 412. Broadcast instructions through MQTT message middleware, and automatically switch to backup communication link (such as LoRaWAN) when timeout without response;

[0129] 413. Feedback data stream is marked with strategy ID and timestamp, and associated with "instruction issuance -> device execution -> risk change" full-link data.

[0130] The working principle and beneficial effects of the above technical solution are: protocol conversion can effectively compatible with industrial devices, solving the control problem of heterogeneous devices; dual-link communication ensures more accurate instructions; data association mechanism provides traceable evidence chain for closed-loop optimization.

[0131] In one embodiment, step 5: According to the multi-modal feedback data stream, the closed-loop optimization module calculates the strategy execution performance index, and dynamically updates the model parameters of the multi-modal LLM inference engine and the rule weight of the safety specification library, forming an adaptive optimization link, including:

[0132] 501. Real-time monitoring of feedback data stream, extracting performance characteristics and calculating comprehensive score;

[0133] 502、when the performance score is lower than the threshold value, triggering a parameter optimization process;

[0134] 503、using a sliding time window to sample the latest K times of policy execution data to construct a training data set;

[0135] 504、updating the LLM model parameters and rule weights synchronously through a multi-task learning framework;

[0136] 505、verifying the performance of the updated model; if the performance improvement does not meet the preset optimization target, the update is cancelled and the system is rolled back to the previous stable version;

[0137] 506、recording optimization logs and updating the version number to complete the closed-loop optimization cycle.

[0138] The working principle and beneficial effects of the above technical solution are as follows: the closed-loop optimization module monitors the multi-modal feedback data stream in real time, extracts performance features such as device response delay and risk state change rate, and calculates a comprehensive performance score. When the score is lower than the threshold value, the parameter optimization process is triggered. A sliding time window is used to construct a training data set of the latest K times of policy execution. A multi-task learning framework is used to update the attention layer parameters of the multi-modal LLM and the rule weights of the safety specification library. After updating, the optimization effect is judged by a performance verification mechanism. If the preset target is not met, the system is automatically rolled back to the previous stable version. Finally, optimization logs are recorded and the system version number is updated to complete the closed-loop optimization cycle from data feedback to model iteration.

[0139] A building risk prediction and control system based on multi-modal LLM, as shown in Figure 2 , includes:

[0140] A multi-modal feature extraction and fusion module is used to interface with the multi-modal data stream collected in real time at the construction site. The multi-modal feature extraction sub-module is used to analyze the structured feature vector set and output it to the multi-modal fusion engine sub-module.

[0141] The multi-modal feature extraction sub-module is based on a deep neural network and includes a time series feature extractor (such as a one-dimensional convolutional neural network or a long short-term memory network for processing sensor time series data), a visual feature extractor (such as a convolutional neural network for processing image and video data), a text feature extractor (such as a pre-trained language model based on Transformer for processing text reports), and a speech feature extractor (such as an automatic speech recognition model combined with a text feature extractor for processing speech instructions).

[0142] The structured feature vector set is a standardized data set formed by timestamp alignment and device ID binding of different modal feature vectors (such as an image feature vector with a dimension of 2048 and a text feature vector with a dimension of 768).

[0143] The multi-modal fusion engine sub-module receives the structured feature vector set through a message queue or an API interface, performs cross-modal feature fusion calculation, and ensures data synchronization and consistency.

[0144] The multi-modal LLM inference engine module is used to receive the structured feature vector set, perform cross-modal semantic fusion and risk coupling analysis, and generate a potential risk identification set and a risk level evaluation result.

[0145] The cross-modal semantic fusion is realized by using a cross-modal attention mechanism to calculate the correlation weight between different modal feature vectors (for example, the attention weight of image features and text features to judge consistency), and generate a unified fusion semantic representation.

[0146] The risk coupling analysis uses a risk propagation model based on a graph neural network to map the fusion semantic representation to a building risk knowledge graph (such as defining the conduction relationship between the "scaffold inclination > 5°" node and the "collapse risk" edge), and analyze the coupling effect between physical facility state, personnel behavior and environmental factors.

[0147] The potential risk identification set includes risk type (such as physical facility anomaly, behavior specification violation, and environmental dynamic threat), risk location, risk feature vector, confidence and timestamp.

[0148] The risk level evaluation result outputs a quantitative risk level label (such as no risk, low risk, medium risk, and high risk) and a probability distribution based on a multi-classification model.

[0149] The dynamic strategy generation module is used to adapt the control rules in the building safety specification library according to the potential risk identification set and the risk level evaluation result, and output a set of targeted risk control strategies.

[0150] The building safety specification library stores standardized rules (such as rule ID, rule type, applicable risk type, trigger condition, control action, execution parameter and priority weight), and the rule type includes equipment control rule, personnel warning rule and area control rule.

[0151] The module calls a rule query interface and uses a multi-constraint rule matching algorithm (such as Rete algorithm) to calculate the matching degree of rule premise condition and risk feature; for the matched rule, the execution parameter is dynamically adjusted according to the risk level evaluation result (such as increasing the personnel warning notification frequency under high risk level), and the conflict resolution algorithm is used to eliminate the instruction conflict (such as executing the high weight strategy first).

[0152] The risk control strategy set is a standardized instruction set (such as JSON or Protocol Buffers format), each strategy contains a strategy ID, target device / personnel identification, control action type (such as device regulation instruction, personnel warning notification, area control suggestion), execution parameter, trigger time window and priority.

[0153] The execution feedback module is used to drive the on-site execution device to implement control actions according to the risk control strategy set, and to collect real-time multi-modal feedback data streams after execution.

[0154] The on-site execution device includes a sensor network (such as an inclination sensor, a visual sensor, accessed through Modbus, OPC UA protocol), a mechanical controller (such as a PLC controller, receiving instructions through EtherCAT, Profinet protocol) and a voice broadcast system (such as an IP broadcast terminal, receiving TTS instructions through SIP protocol).

[0155] The module parses the strategy set and converts it into device-specific protocol messages (such as encoding JSON strategies into ModbusTCP frames), and distributes instructions through an industrial Internet of Things platform (such as a message middleware based on MQTT or Apache Kafka). At the same time, monitor the execution status of the instructions, and trigger a retry mechanism or a degradation process (such as switching to a LoRaWAN backup link) when the timeout or failure occurs.

[0156] The multi-modal feedback data stream includes device state feedback data (such as device inclination, running status), environmental response data (such as personnel evacuation video stream) and execution efficiency raw data (such as instruction response delay, action completion flag), which are encapsulated into a standard format (such as Apache Avro serialized data) after being timestamped and associated with the strategy ID.

[0157] The closed-loop optimization module is used to calculate the strategy execution performance indicators according to the multi-modal feedback data stream, and dynamically update the model parameters of the multi-modal LLM inference engine module and the rule weights of the building safety specification library.

[0158] The module extracts the performance feature vector (such as device response delay, action execution completion degree, risk state change rate), uses the performance evaluation model (such as multi-layer perception MLP) to calculate the strategy execution performance indicators, including immediate performance indicators (instruction response time, action execution success rate), risk control indicators (risk level decline rate, secondary risk occurrence rate) and comprehensive performance score (percentage weighted sum result).

[0159] The model parameters of the dynamic updating multi-modal LLM inference engine module are implemented by an online learning algorithm: a reinforcement learning reward function is constructed based on the performance indicators (the shorter the response time, the higher the success rate, and the more significant the risk reduction, the higher the reward value), and the attention layer and output layer weights are updated using a policy gradient algorithm (such as PPO or REINFORCE).

[0160] The rule weights of the dynamic updating building safety specification library are implemented by a statistical learning method: the weights are adjusted according to the historical performance scores of the rules (the scores are adjusted upwards, and the scores are adjusted downwards or disabled if they are continuously low), and new rules are extracted from successful strategy cases through a rule mining algorithm.

[0161] The system realizes the unified processing of heterogeneous data through the multi-modal feature extraction and fusion module, solves the problem of missing key features in traditional single-modal analysis; the multi-modal LLM inference engine module reduces the false positive rate by cross-verifying risks using a knowledge graph and an attention mechanism; the dynamic strategy generation module realizes millisecond-level rule matching using the Rete algorithm, improving response efficiency; the execution feedback module supports multi-protocol device control, ensuring the reachability of instructions; the closed-loop optimization module forms an adaptive optimization link through reinforcement learning and online learning, continuously improving risk prediction accuracy and control effectiveness. The entire system is strictly based on the logic of Document 2, ensuring the adaptability and real-time performance of risk prediction and control.

[0162] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

Claims

1. A method for building risk prediction management and control based on a multi-modal LLM, characterized in that, Comprising: Step 1: According to the real-time acquisition of multi-modal data flow in construction site, through multi-modal feature extraction module to analyze into structured feature vector set, and output to multi-modal fusion engine; Step 2: According to the structured feature vector set, through multi-modal LLM inference engine to carry out cross-modal semantic fusion and risk coupling analysis, generate potential risk identification set and risk level evaluation result, wherein the risk type includes physical facility anomaly, behavior specification violation and environmental dynamic threat; Step 3: According to the potential risk identification set and risk level evaluation result, dynamically match the control rules of building safety specification library, output risk control strategy set; Step 4: According to the risk control strategy set, drive the on-site execution equipment to implement control action, and collect multi-modal feedback data flow; Step 5: According to the feedback data flow, calculate the strategy execution performance index, dynamically update the LLM model parameters and rule weight.

2. The multi-modal LLM-based building risk prediction management method of claim 1, wherein, The step 1 comprises: Using time series feature extractor to process sensor time series data, visual feature extractor to process image video data, text feature extractor to process text report, and voice feature extractor to process voice instruction; the multi-modal feature vectors are timestamped and device ID bound to generate structured feature vector set.

3. The multi-modal LLM-based building risk prediction management method of claim 1, wherein, The step 2 comprises: Through cross-modal attention mechanism to calculate the correlation weight of different modal feature vectors, generate fusion semantic representation; based on the risk propagation model of graph neural network, map the fusion semantic representation to building risk knowledge graph, analyze the risk coupling effect.

4. The multi-modal LLM-based building risk prediction management method of claim 1, wherein, The step 3 comprises: using Rete algorithm to match safety specification library rules, dynamically adjusting execution parameters according to risk level; through conflict resolution algorithm, the control instructions are ordered according to priority weight.

5. The multi-modal LLM-based building risk prediction management method of claim 1, wherein, The step 5 comprises: Based on reinforcement learning reward function to update LLM attention layer parameters; according to the rule historical performance score, dynamically adjust the rule weight, and trigger the disable mechanism when the performance score is continuously too low.

6. The multi-modal LLM-based building risk prediction management system of any one of claims 1-5, wherein, Comprising: Multi-modal feature extraction and fusion module, used for analyzing multi-modal data flow and generating structured feature vector set; Multi-modal LLM inference engine module, used for carrying out cross-modal semantic fusion and risk coupling analysis, outputting risk identification and level evaluation; Dynamic strategy generation module, used for adapting safety specification library rules and generating risk control strategy set; Execution feedback module, used for driving on-site execution equipment to implement control action and collecting feedback data; Closed loop optimization module, used for updating LLM model parameters and rule weight.

7. The multi-modal LLM-based building risk prediction management system of claim 6, wherein, The multi-modal feature extraction and fusion module comprises: Time series feature extractor, visual feature extractor, text feature extractor and voice feature extractor; Multi-modal fusion engine receives feature vector set through API interface.

8. The multi-modal LLM-based building risk prediction management system of claim 6, wherein, The multi-modal LLM inference engine module has built-in building risk knowledge graph, defining node entity and risk transmission relationship; the output risk identification includes risk type, location, confidence and timestamp.

9. The multi-modal LLM-based building risk prediction management system of claim 6, wherein, The dynamic strategy generation module calls a building safety specification library, the rule types include equipment regulation rules, personnel warning rules and area control rules; the strategy set is packaged in JSON or Protocol Buffers format; the closed-loop optimization module performance evaluation model is a multi-layer perception machine, and an output comprehensive performance score is adopted to update the LLM model parameters.

10. An electronic device comprising a processor and a memory, characterized in that The memory stores a computer program, which, when executed by the processor, implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Modern building information management and control system and method based on Internet of Things

    CN118037492A

Cited By

  • Project risk prediction and dynamic early warning method and system based on multi-modal AI

    CN121094566A

  • Risk dynamic processing method and device in real-time communication scene, equipment and medium

    CN121151505A

  • LLM-based ship aided driving realization method and system

    CN121277187A

  • Multi-source data fusion type passenger station intelligent management and control method and system based on AI drive

    CN121544079A

  • Dynamic risk management and control method and system based on papermaking production safety management

    CN122243159A