A confined space operation multi-modal risk fusion management and control system and method
Patent Information
- Application Number
- CN202611042881.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]现有视频监控和气体检测数据多以独立文件或独立记录形式保存,事故发生后难以准确对应同一时刻的环境浓度、人员状态和现场画面,影响事故复盘和证据追溯
[0057]与现有技术相比,本发明具有的有益效果是:通过状态机将身份核验、通风、预采、监护确认和正式作业准入绑定为强制流程,提高作业流程执行一致性;通过预采检测合格与监护确认共同作为正式作业准入条件,降低未检测合格即进入受限空间作业的风险;通过像素级数据叠加实现气体数据、报警状态与视频画面的帧级同步显示,使气体检测数据作为画面内容与现场视频同步留存,提高事故追溯便利性并降低数据与画面分离造成的证据不一致风险;通过多模态风险融合方法综合判断气体、行为、监护和应急事件,提升风险识别完整性;通过边缘本地判断与云端复核协同,兼顾低时延告警和复杂场景分析能力。
Smart Images

Figure CN122840672A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of confined space safety operation monitoring technology, specifically to a multimodal risk fusion control system and method for confined space operations. Background Technology
[0002] Confined spaces, such as reactors, towers, storage tanks, pipelines, and wells, are typically characterized by poor ventilation, limited access, and complex internal environments, which can easily lead to safety risks such as the accumulation of toxic and harmful gases, insufficient oxygen content, personnel collapse, absenteeism of supervisors, and blind rescue attempts.
[0003] Current methods for monitoring confined space operations typically rely on independent gas detectors, video monitoring terminals, and manual recording. Gas detection data and video feeds are displayed separately, requiring monitoring personnel to switch between different devices, making it difficult to promptly make comprehensive assessments of multi-source risks.
[0004] Current process control relies primarily on manual execution by operators and supervisors. Although relevant operational standards require the implementation of procedures such as ventilation, testing, verification, and work access, there is a lack of technical means on-site to link gas detection results, personnel identity verification results, and supervisor confirmation signals as mandatory access conditions.
[0005] Existing video surveillance and gas detection data are mostly stored in the form of independent files or independent records. After an accident, it is difficult to accurately match the environmental concentration, personnel status and on-site footage at the same time, which affects accident review and evidence tracing. Summary of the Invention
[0006] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0007] Therefore, the purpose of this invention is to provide a multimodal risk fusion management and control system and method for confined space operations to solve the problems mentioned in the background art.
[0008] To address the aforementioned technical problems, according to one aspect of the present invention, the present invention provides the following technical solution:
[0009] A multimodal risk fusion management and control system for confined space operations includes an edge computing host, an all-in-one gas sampler, an explosion-proof camera, and a remote cloud platform; the edge computing host includes:
[0010] The mandatory compliance control module presets identity verification status, ventilation status, pre-sampling and testing status, formal operation status and operation completion status, and uses personnel identity verification results, guardian confirmation signals, gas detection thresholds and abnormal alarm status as status transition conditions to forcibly bind the operation access process.
[0011] The multi-source data fusion module is used to decode the video stream collected by the explosion-proof camera, obtain the video frame pixel matrix, draw the gas concentration data, alarm status and timestamp collected by the multi-in-one gas sampler onto the video frame pixel matrix according to preset coordinates, and encode the superimposed video frames to generate an integrated video stream in which gas data, alarm status and monitoring screen are displayed synchronously.
[0012] An AI inference engine is used to output a comprehensive risk level based on gas concentration data, video behavior scene recognition results, and monitoring confirmation feedback, through a multimodal risk fusion method.
[0013] The alarm linkage module is used to trigger local alarms, remote alarms, or work stoppage control according to the comprehensive risk level, and to directly trigger alarm linkage and record evidence according to the SOS button trigger signal.
[0014] The evidence report generation module is used to aggregate identity verification records, state machine transition records, gas detection data, risk assessment results, alarm records, and video slices to generate an operation report after the operation is completed.
[0015] In a preferred embodiment of the multimodal risk fusion management and control system for confined space operations described in this invention, the mandatory compliance control module is configured as follows:
[0016] During identity verification, images of on-site personnel are captured and facial features are extracted. The facial features are then compared with a pre-defined database of authorized personnel. Entry into the ventilation state is only permitted after the comparison is successful.
[0017] Receive ventilation confirmation signal from the supervisor while in ventilation mode, and receive gas concentration data while in pre-sampling and detection mode;
[0018] When the gas concentration data meets the preset safety threshold and a pre-sampling confirmation signal is received from the supervisor, the formal operation state is allowed. When the gas concentration data exceeds the preset safety threshold, an alarm is triggered and the formal operation state is prohibited.
[0019] As a preferred embodiment of the multimodal risk fusion and management system for confined space operations described in this invention, the AI inference engine, when performing multimodal risk fusion, specifically includes:
[0020] Calculate the gas risk sub-value, the behavioral scenario risk sub-value, and the monitoring risk sub-value respectively, and output the comprehensive risk level based on the gas risk sub-value, the behavioral scenario risk sub-value, and the monitoring risk sub-value;
[0021] The gas risk sub-value is determined by the ratio of the measured values of oxygen, combustible gas, and toxic and harmful gas to the corresponding safety threshold.
[0022] The risk sub-value of the behavioral scenario is determined by at least one of the following events: not wearing a safety helmet, personnel falling to the ground, personnel sleeping on duty, unauthorized use of electronic devices, smoke detection, or flame detection.
[0023] The monitoring risk sub-value is determined by at least one of the following events: guardian feedback timeout, guardian leaving the monitoring area, or lack of confirmation signal.
[0024] As a preferred embodiment of the multimodal risk fusion management and control system for confined space operations described in this invention, the AI inference engine further includes a credibility assessment unit and a dynamic weight adjustment unit.
[0025] The credibility assessment unit calculates the credibility of the video modality, the gas modality, the network modality, and the monitoring interaction modality, respectively, where:
[0026] The video modality confidence is determined based on image clarity, occlusion ratio, target detection confidence, and continuous frame stability.
[0027] The reliability of the gas mode is determined based on the sensor sampling volatility, drift, response time, and calibration status.
[0028] The dynamic weight adjustment unit dynamically adjusts the fusion weights of the gas risk sub-value, behavioral scenario risk sub-value, and monitoring risk sub-value in the comprehensive risk value based on the confidence level of each modality.
[0029] As a preferred embodiment of the multimodal risk fusion management and control system for confined space operations described in this invention, the AI inference engine further includes a cross-modal consistency judgment unit.
[0030] The cross-modal consistency judgment unit calculates the cross-modal consistency index based on the correspondence between gas modes, behavioral modes, and scene modes;
[0031] When the cross-modal consistency index is lower than the preset consistency threshold, the AI inference engine control system enters the cloud review mode;
[0032] Wherein, the cross-modal consistency index C = α·Sgb + β·Sgs + γ·Sbs, where Sgb represents the consistency score between the gas mode and the behavioral mode, Sgs represents the consistency score between the gas mode and the scene mode, Sbs represents the consistency score between the behavioral mode and the scene mode, and α, β, and γ are the corresponding weight coefficients; Sgb, Sgs, and Sbs are obtained by looking up the current recognition result or matching the rules through a preset modal correspondence table, and their values range from 0 to 1.
[0033] As a preferred embodiment of the multimodal risk fusion management and control system for confined space operations described in this invention, the AI inference engine further includes a time-series risk trend prediction unit.
[0034] The time-series risk trend prediction unit predicts the comprehensive risk level within a preset time period based on the gas concentration sequence, personnel behavior event sequence, monitoring feedback sequence, comprehensive risk value sequence, and cross-modal consistency index sequence within a preset sliding time window.
[0035] When the predicted comprehensive risk level reaches the warning level, alarm level, or work stoppage level within a preset time in the future, the system will trigger a risk warning in advance or enter the cloud review mode.
[0036] As a preferred embodiment of the multimodal risk fusion management and control system for confined space operations described in this invention, the multi-source data fusion module includes a data synchronization unit, a character layer generation unit, and a hardware encoding unit.
[0037] The data synchronization unit associates gas concentration data, alarm status and corresponding video frames based on timestamps.
[0038] The character layer generation unit converts gas concentration data, alarm status, and time information into character layers and draws them onto a preset area of the video frame pixel matrix.
[0039] The hardware encoding unit performs secondary encoding on the superimposed video frames, and the resulting integrated video stream is used for local display on the host, video recording and storage, and cloud viewing.
[0040] As a preferred embodiment of the multimodal risk fusion management and control system for confined space operations described in this invention, the system further includes a monitoring mechanism for guardian absence.
[0041] In the formal operation state, the edge computing host pushes an interactive confirmation request to the monitoring terminal according to a preset period T;
[0042] If the guardian does not provide a confirmation signal within the preset response time M, and does not receive a valid feedback signal within N consecutive confirmation cycles, a guardian absence alarm will be generated and the alarm linkage module will be triggered.
[0043] The all-in-one gas sampler is also equipped with an SOS physical alarm button. When the SOS physical alarm button is triggered, the alarm linkage module simultaneously triggers the host-side audible and visual alarm and the remote monitoring-side alarm, and writes the SOS trigger time into the evidence report.
[0044] In a preferred embodiment of the multimodal risk fusion management and control system for confined space operations described in this invention, the remote cloud platform is configured as follows:
[0045] Receive structured risk events, key video slices, and sensor data uploaded by the edge computing host;
[0046] When the local comprehensive risk level is at the preset review level, the cross-modal consistency index is lower than the preset consistency threshold, or the predicted comprehensive risk level reaches the preset review level, the multimodal model or preset rule engine is invoked to perform risk review.
[0047] The risk review results are returned to the edge computing host, including maintaining the original level, upgrading the level, downgrading the level, or marking it as a suspected false alarm.
[0048] A multimodal risk fusion management method for confined space operations includes the following steps:
[0049] S1. Collect images of on-site personnel and verify their identities. Once the identity verification is successful, proceed to ventilation mode.
[0050] S2. Perform the ventilation process while the ventilation is in progress, and enter the pre-sampling and testing state after receiving the ventilation confirmation signal from the supervisor.
[0051] S3. Collect gas concentration data in the pre-collection detection state. When the gas concentration data meets the preset safety threshold and a confirmation signal is received from the supervisor, enter the formal operation state.
[0052] S4. In the formal operation state, synchronously collect video stream, gas concentration data, personnel behavior scene recognition results, monitoring confirmation feedback and SOS button trigger signal;
[0053] S5. Calculate the comprehensive risk level based on gas concentration data, personnel behavior scene recognition results, and monitoring confirmation feedback, and use the SOS button trigger signal as an independent emergency trigger signal;
[0054] S6. Superimpose gas concentration data and alarm status onto the video frame pixel matrix and encode to generate an integrated video stream;
[0055] S7. Trigger local alarms, remote alarms, or work stoppage control based on the comprehensive risk level, and directly trigger alarm linkage and record storage based on the SOS button trigger signal;
[0056] S8. After the operation is completed, an operation report is generated that includes identity verification records, state machine transition records, gas detection data, risk assessment results, alarm records, and video clips.
[0057] Compared with existing technologies, the beneficial effects of this invention are as follows: It binds identity verification, ventilation, pre-sampling, monitoring confirmation, and formal operation access into a mandatory process through a state machine, improving the consistency of operation process execution; it uses pre-sampling detection qualification and monitoring confirmation as joint conditions for formal operation access, reducing the risk of entering confined spaces without passing the detection; it achieves frame-level synchronous display of gas data, alarm status, and video footage through pixel-level data overlay, allowing gas detection data to be stored synchronously with the on-site video as screen content, improving the convenience of accident tracing and reducing the risk of inconsistency in evidence caused by data and video separation; it comprehensively judges gas, behavior, monitoring, and emergency events through a multimodal risk fusion method, improving the completeness of risk identification; and it combines low-latency alarms and complex scenario analysis capabilities through edge local judgment and cloud verification collaboration. Attached Figure Description
[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0059] Figure 1 This is a flowchart illustrating the overall architecture and data interaction process of a multimodal risk fusion management and control system for confined space operations according to the present invention.
[0060] Figure 2 This is a schematic diagram illustrating the logical flow of identity verification, ventilation, pre-sampling detection, formal operation, and operation completion status in the operation mode of a multimodal risk fusion control system for confined space operations according to the present invention.
[0061] Figure 3 This is a flowchart illustrating pixel-level data overlay and frame-level synchronization of a multimodal risk fusion management and control system for confined space operations according to the present invention. Detailed Implementation
[0062] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0063] Figures 1-3 The diagram shown illustrates the framework and flowchart of a multimodal risk fusion management and control system for confined space operations according to the present invention. Please refer to [link / reference]. Figures 1-3This embodiment of a confined space operation multimodal risk fusion management and control system includes an edge computing host, an all-in-one gas sampler, an explosion-proof camera, and a remote cloud platform; the edge computing host includes a mandatory compliance control module, a multi-source data fusion module, an AI inference engine, an alarm linkage module, and an evidence storage report generation module.
[0064] The mandatory compliance control module presets identity verification status, ventilation status, pre-sampling and testing status, formal operation status and operation completion status, and uses personnel identity verification results, guardian confirmation signals, gas detection thresholds and abnormal alarm status as status transition conditions to forcibly bind the operation access process;
[0065] The multi-source data fusion module is used to decode the video stream collected by the explosion-proof camera, obtain the video frame pixel matrix, draw the gas concentration data, alarm status and timestamp collected by the multi-in-one gas sampler onto the video frame pixel matrix according to preset coordinates, and encode the superimposed video frames to generate an integrated video stream in which gas data, alarm status and monitoring screen are displayed synchronously.
[0066] The AI inference engine is used to output a comprehensive risk level based on gas concentration data, video behavior scene recognition results, and monitoring confirmation feedback, through a multimodal risk fusion method.
[0067] The alarm linkage module is used to trigger local alarms, remote alarms, or work stoppage control according to the comprehensive risk level, and to directly trigger alarm linkage and record evidence according to the SOS button trigger signal;
[0068] The evidence report generation module is used to aggregate identity verification records, state machine transition records, gas detection data, risk assessment results, alarm records, and video slices to generate an operation report after the operation is completed.
[0069] The mandatory compliance control module is configured to: acquire images of on-site personnel and extract facial features during identity verification; compare the facial features with a preset authorized personnel database; and allow entry into ventilation mode only after successful comparison; receive ventilation confirmation signals from supervisors during ventilation mode and receive gas concentration data during pre-sampling detection mode; allow entry into formal operation mode when gas concentration data meets a preset safety threshold and a pre-sampling confirmation signal from supervisors is received; and trigger an alarm and prohibit entry into formal operation mode when gas concentration data exceeds the preset safety threshold.
[0070] When performing multimodal risk fusion, the AI inference engine specifically includes: calculating gas risk sub-values, behavioral scenario risk sub-values, and monitoring risk sub-values respectively, and outputting a comprehensive risk level based on the gas risk sub-values, behavioral scenario risk sub-values, and monitoring risk sub-values; the gas risk sub-value is determined by the ratio of the measured value of oxygen, combustible gas, and toxic and harmful gas to the corresponding safety threshold; the behavioral scenario risk sub-value is determined by at least one of the following events: not wearing a safety helmet, personnel falling to the ground, personnel sleeping on duty, unauthorized use of electronic devices, smoke detection, or flame detection; the monitoring risk sub-value is determined by at least one of the following events: monitoring person feedback timeout, monitoring person leaving the monitoring area, or lack of confirmation signal.
[0071] The AI inference engine further includes a credibility assessment unit and a dynamic weight adjustment unit. The credibility assessment unit calculates the credibility of the video modality, the gas modality, the network modality, and the monitoring interaction modality, respectively. The video modality credibility is determined based on image clarity, occlusion ratio, target detection confidence, and continuous frame stability. The gas modality credibility is determined based on sensor sampling volatility, drift, response time, and calibration status. The dynamic weight adjustment unit dynamically adjusts the fusion weights of the gas risk sub-value, the behavioral scenario risk sub-value, and the monitoring risk sub-value in the comprehensive risk value based on the credibility of each modality.
[0072] The AI inference engine also includes a cross-modal consistency judgment unit; the cross-modal consistency judgment unit calculates a cross-modal consistency index based on the correspondence between gas mode, behavioral mode and scene mode; when the cross-modal consistency index is lower than a preset consistency threshold, the AI inference engine control system enters cloud review mode;
[0073] Wherein, the cross-modal consistency index C = α·Sgb + β·Sgs + γ·Sbs, where Sgb represents the consistency score between the gas mode and the behavioral mode, Sgs represents the consistency score between the gas mode and the scene mode, Sbs represents the consistency score between the behavioral mode and the scene mode, and α, β, and γ are the corresponding weight coefficients; Sgb, Sgs, and Sbs are obtained by looking up the current recognition result or matching the rules through a preset modal correspondence table, and their values range from 0 to 1.
[0074] When a person is detected falling to the ground and the gas modality is within a safe range, and the scene modality does not detect fire smoke, the event is marked as a personnel anomaly event; when the gas modality shows that the hazardous gas exceeds the standard and the behavioral modality does not detect personnel anomalies, and the scene modality does not detect smoke or fire, the event is marked as a gas anomaly event; when the scene modality detects smoke and the gas modality and temperature status are not abnormal, the event is marked as a visual suspected false alarm event.
[0075] The AI inference engine also includes a time-series risk trend prediction unit; the time-series risk trend prediction unit predicts the comprehensive risk level within a preset time period based on the gas concentration sequence, personnel behavior event sequence, monitoring feedback sequence, comprehensive risk value sequence and cross-modal consistency index sequence within a preset sliding time window; when the predicted comprehensive risk level reaches the warning level, alarm level or work stoppage level within the preset time period, the system triggers a risk warning in advance or enters the cloud review mode.
[0076] The multi-source data fusion module includes a data synchronization unit, a character layer generation unit, and a hardware encoding unit. The data synchronization unit associates gas concentration data, alarm status, and corresponding video frames based on timestamps. The character layer generation unit converts gas concentration data, alarm status, and time information into character layers and draws them onto a preset area of the video frame pixel matrix. The hardware encoding unit performs secondary encoding on the superimposed video frames, and the generated integrated video stream is used for local display on the host, recording storage, and cloud viewing.
[0077] The system also includes a guardian absence monitoring mechanism. During formal operation, the edge computing host pushes an interactive confirmation request to the monitoring terminal according to a preset cycle T. If the guardian does not provide a confirmation signal within a preset response time M, and does not receive a valid feedback signal within N consecutive confirmation cycles, a guardian absence alarm is generated and the alarm linkage module is triggered. The all-in-one gas sampler is also equipped with an SOS physical alarm button. When the SOS physical alarm button is triggered, the alarm linkage module simultaneously triggers the host-side audible and visual alarm and the remote monitoring terminal alarm, and writes the SOS trigger time into the evidence report. After the operation is completed, the evidence report generation module aggregates personnel identity verification records, state machine transition records, gas detection data, comprehensive risk level, alarm records, SOS trigger records, and video clips to generate an operation report, which is then distributed via QR code, local storage, email, or cloud interface.
[0078] The remote cloud platform is configured to: receive structured risk events, key video slices, and sensor data uploaded by the edge computing host; when the local comprehensive risk level is at a preset review level, the cross-modal consistency index is lower than a preset consistency threshold, or the predicted comprehensive risk level reaches a preset review level, invoke a multimodal model or a preset rule engine to perform risk review; and return the risk review result to the edge computing host, wherein the review result includes maintaining the original level, raising the level, lowering the level, or marking it as a suspected false alarm.
[0079] Combination Figures 1-3 The specific steps of the multimodal risk fusion management and control system for confined space operations according to this embodiment are as follows:
[0080] S1. Collect images of on-site personnel and verify their identities. Once the identity verification is successful, proceed to ventilation mode.
[0081] S2. Perform the ventilation process while the ventilation is in progress, and enter the pre-sampling and testing state after receiving the ventilation confirmation signal from the supervisor.
[0082] S3. Collect gas concentration data in the pre-collection detection state. When the gas concentration data meets the preset safety threshold and a confirmation signal is received from the supervisor, enter the formal operation state.
[0083] S4. In the formal operation state, synchronously collect video stream, gas concentration data, personnel behavior scene recognition results, monitoring confirmation feedback and SOS button trigger signal;
[0084] S5. Calculate the comprehensive risk level based on gas concentration data, personnel behavior scene recognition results, and monitoring confirmation feedback, and use the SOS button trigger signal as an independent emergency trigger signal;
[0085] S6. Superimpose gas concentration data and alarm status onto the video frame pixel matrix and encode to generate an integrated video stream;
[0086] S7. Trigger local alarms, remote alarms, or work stoppage control based on the comprehensive risk level, and directly trigger alarm linkage and record storage based on the SOS button trigger signal;
[0087] S8. After the operation is completed, an operation report is generated that includes identity verification records, state machine transition records, gas detection data, risk assessment results, alarm records, and video clips.
[0088] Preferably, step S5 further includes: calculating the video modal credibility, gas modal credibility, network modal credibility, and monitoring interaction modal credibility respectively, and adjusting the fusion weight of the gas risk sub-value, behavior scenario risk sub-value, monitoring risk sub-value, or cloud review result in the comprehensive risk value according to the credibility of each modality; calculating the cross-modal consistency index based on the correspondence between the gas modality, behavior modality, and scenario modality; entering the cloud review mode when the cross-modal consistency index is lower than the preset consistency threshold; predicting the comprehensive risk level within a preset time period based on the gas concentration sequence, personnel behavior event sequence, monitoring feedback sequence, comprehensive risk value sequence, and cross-modal consistency index sequence within a preset sliding time window; and triggering a risk warning or entering the cloud review mode in advance when the predicted comprehensive risk level reaches the warning level, alarm level, or work stoppage level.
[0089] Preferably, step S6 includes sequential decoding of the original video stream to obtain a video frame pixel matrix; associating gas concentration data and alarm status with the corresponding video frames based on timestamps; generating a character layer containing gas concentration data, alarm status, and time information; drawing the character layer onto a preset area of the video frame pixel matrix; and hardware encoding the superimposed video frames to obtain an integrated video stream.
[0090] In one specific embodiment, the system administrator pre-enters the identity data of operators, monitors and managers in the background, and configures the gas detection threshold, state machine transition conditions, alarm threshold and report distribution method.
[0091] Before the operation begins, the guardian initiates the operation process through the edge computing host. The edge computing host starts the identity verification program, collects images of on-site personnel and extracts facial features, and compares the extracted facial features with the authorized personnel database. If the comparison fails, the system remains in identity verification state and outputs an alarm message; if the comparison is successful, the system enters ventilation state.
[0092] During ventilation, the system executes the ventilation procedure. After ventilation is complete, the supervisor inputs a ventilation confirmation signal; when the confirmation signal is valid, the system enters the pre-sampling detection state.
[0093] In the pre-sampling detection state, the operator connects the sampler to the pre-sampling tube and places it inside the confined space. If any gas concentration data exceeds the preset safety threshold, the system triggers an alarm and returns to ventilation mode; if all gas concentration data meet the preset safety threshold and the supervisor confirms that pre-sampling is complete, the system enters the formal operation state.
[0094] In formal operation mode, construction workers enter confined spaces carrying samplers and cameras. Explosion-proof cameras collect video footage of the work site, and an AI inference engine identifies events such as wearing safety helmets, workers falling to the ground, workers sleeping on duty, unauthorized use of electronic devices, and smoke or fire. The edge computing host simultaneously receives gas concentration data, monitoring confirmation feedback, and SOS button signals, and outputs a comprehensive risk level.
[0095] In a preferred embodiment, the AI inference engine further calculates the credibility of each modality and adjusts the corresponding fusion weights when the video is obscured by smoke, gas sensor fluctuations occur, or network latency occurs. The AI inference engine can also calculate a cross-modal consistency index, triggering cloud-based review when there are conflicts between gas, behavior, and scene modalities. The AI inference engine can also predict the risk level within a preset time period based on the risk sequence within a sliding time window, thereby triggering risk warnings in advance or entering cloud-based review mode.
[0096] The multi-source data fusion module decodes the camera video stream to obtain a video frame pixel matrix; the data synchronization unit matches the gas concentration data, alarm status, and system time at the corresponding time according to the timestamp; the character layer generation unit converts the above data into a character layer and draws it onto the preset area of the video frame pixel matrix; the hardware encoding unit re-encodes the superimposed pixel matrix into an integrated video stream.
[0097] As a specific implementation of cross-modal consistency scoring, the system pre-configures a modal correspondence table. Taking the gas modality-behavioral modality score Sgb as an example, when the gas concentration exceeds the limit and the video simultaneously identifies a person falling to the ground, the Sgb value is 1.0; when the gas concentration is normal and the video does not identify any abnormal behavior, the Sgb value is 0.9; when the gas concentration exceeds the limit but the video does not identify any abnormal behavior, the Sgb value is 0.3; when the gas concentration is normal but the video identifies a person falling to the ground or fainting, the Sgb value is 0.4; other combinations are determined according to a preset rule table. Sgs and Sbs are configured in a similar manner. In a specific embodiment, α, β, and γ are set to 0.4, 0.3, and 0.3, respectively. When C is lower than 0.5, a modal conflict is determined, and the system enters the cloud review mode. The above values are only examples, and those skilled in the art can adjust them according to the actual scenario.
[0098] As a specific implementation of cloud-based verification, the remote cloud platform receives structured risk events, key video slices, gas concentration data, personnel behavior scene recognition results, cross-modal consistency index, and timestamp information uploaded by the edge computing host. The remote cloud platform calls a multimodal model or a preset rule engine to verify the consistency between gas modality, behavior modality, and scene modality, and generates a risk verification result. The risk verification result includes at least one of the following: maintaining the original comprehensive risk level, increasing the comprehensive risk level, decreasing the comprehensive risk level, or marking it as a suspected false alarm.
[0099] As a specific implementation of the monitoring of guardian absence, in formal operation, the edge computing host pushes an interactive confirmation request to the monitoring terminal according to a preset period T, where T is 300 seconds; the guardian needs to provide a confirmation signal via button or touch within a preset response time M, where M is 30 seconds; if no valid feedback signal is received within N consecutive confirmation periods, the guardian is determined to be absent from their post and a guardian absence alarm is generated, where the minimum value of N is 1; the alarm simultaneously triggers an audible and visual alarm on the host side and an alarm on the remote monitoring terminal, and is written into the evidence report. The specific values of T, M, and N can be adjusted according to the operational risk level.
[0100] As a typical example of cross-modal consistency judgment triggering events: when the gas modality shows that the hydrogen sulfide concentration exceeds the alarm threshold, the video behavior modality does not identify any abnormal personnel, and the scene modality does not identify any smoke or fire, the Sgb is low, the system judges it as a gas abnormality event and enters the cloud for review; when the video detects a person falling to the ground, and both the gas modality and the scene modality are within the normal range, the Sgb is low, the system judges it as a personnel abnormality event and triggers a local alarm or cloud review; when the scene modality detects suspected smoke, the gas modality does not detect an increase in CO and the temperature is not abnormal, the Sgs is low, the system judges it as a visual suspected false alarm event and enters the cloud review mode.
[0101] After the operation is completed, the monitoring personnel click the "End Operation" button, and the system enters the operation completion state. The evidence report generation module aggregates identity verification records, state machine transition records, gas detection data, risk assessment results, alarm records, SOS trigger records, and video clips to generate a standard format operation report, which is then distributed and stored via QR code, local storage, email, or cloud interface.
[0102] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A multimodal risk fusion management and control system for confined space operations, characterized in that, It includes an edge computing host, an all-in-one gas sampler, an explosion-proof camera, and a remote cloud platform; the edge computing host includes: The mandatory compliance control module presets identity verification status, ventilation status, pre-sampling and testing status, formal operation status and operation completion status, and uses personnel identity verification results, guardian confirmation signals, gas detection thresholds and abnormal alarm status as status transition conditions to forcibly bind the operation access process. The multi-source data fusion module is used to decode the video stream collected by the explosion-proof camera, obtain the video frame pixel matrix, draw the gas concentration data, alarm status and timestamp collected by the multi-in-one gas sampler onto the video frame pixel matrix according to preset coordinates, and encode the superimposed video frames to generate an integrated video stream in which gas data, alarm status and monitoring screen are displayed synchronously. An AI inference engine is used to output a comprehensive risk level based on gas concentration data, video behavior scene recognition results, and monitoring confirmation feedback, through a multimodal risk fusion method. The alarm linkage module is used to trigger local alarms, remote alarms, or work stoppage control according to the comprehensive risk level, and to directly trigger alarm linkage and record evidence according to the SOS button trigger signal. The evidence report generation module is used to aggregate identity verification records, state machine transition records, gas detection data, risk assessment results, alarm records, and video slices to generate an operation report after the operation is completed.
2. The multimodal risk fusion management and control system for confined space operations according to claim 1, characterized in that, The mandatory compliance control module is configured as follows: During identity verification, images of on-site personnel are captured and facial features are extracted. The facial features are then compared with a pre-defined database of authorized personnel. Entry into the ventilation state is only permitted after the comparison is successful. Receive ventilation confirmation signal from the supervisor while in ventilation mode, and receive gas concentration data while in pre-sampling and detection mode; When the gas concentration data meets the preset safety threshold and a pre-sampling confirmation signal is received from the supervisor, the formal operation state is allowed. When the gas concentration data exceeds the preset safety threshold, an alarm is triggered and the formal operation state is prohibited.
3. The multimodal risk fusion management and control system for confined space operations according to claim 1, characterized in that, The AI inference engine, when performing multimodal risk fusion, specifically includes: Calculate the gas risk sub-value, the behavioral scenario risk sub-value, and the monitoring risk sub-value respectively, and output the comprehensive risk level based on the gas risk sub-value, the behavioral scenario risk sub-value, and the monitoring risk sub-value; The gas risk sub-value is determined by the ratio of the measured values of oxygen, combustible gas, and toxic and harmful gas to the corresponding safety threshold. The risk sub-value of the behavioral scenario is determined by at least one of the following events: not wearing a safety helmet, personnel falling to the ground, personnel sleeping on duty, unauthorized use of electronic devices, smoke detection, or flame detection. The monitoring risk sub-value is determined by at least one of the following events: guardian feedback timeout, guardian leaving the monitoring area, or lack of confirmation signal.
4. The multimodal risk fusion management and control system for confined space operations according to claim 3, characterized in that, The AI inference engine also includes a credibility assessment unit and a dynamic weight adjustment unit; The credibility assessment unit calculates the credibility of the video modality, the gas modality, the network modality, and the monitoring interaction modality, respectively, where: The video modality confidence is determined based on image clarity, occlusion ratio, target detection confidence, and continuous frame stability. The reliability of the gas mode is determined based on the sensor sampling volatility, drift, response time, and calibration status. The dynamic weight adjustment unit dynamically adjusts the fusion weights of the gas risk sub-value, behavioral scenario risk sub-value, and monitoring risk sub-value in the comprehensive risk value based on the confidence level of each modality.
5. A multimodal risk fusion management and control system for confined space operations according to claim 3, characterized in that, The AI inference engine also includes a cross-modal consistency judgment unit; The cross-modal consistency judgment unit calculates the cross-modal consistency index based on the correspondence between gas modes, behavioral modes, and scene modes; When the cross-modal consistency index is lower than the preset consistency threshold, the AI inference engine control system enters the cloud review mode; Wherein, the cross-modal consistency index C = α·Sgb + β·Sgs + γ·Sbs, where Sgb represents the consistency score between the gas mode and the behavioral mode, Sgs represents the consistency score between the gas mode and the scene mode, Sbs represents the consistency score between the behavioral mode and the scene mode, and α, β, and γ are the corresponding weight coefficients; Sgb, Sgs, and Sbs are obtained by looking up the current recognition result or matching the rules through a preset modal correspondence table, and their values range from 0 to 1.
6. A multimodal risk fusion management and control system for confined space operations according to claim 5, characterized in that, The AI inference engine also includes a time-series risk trend prediction unit; The time-series risk trend prediction unit predicts the comprehensive risk level within a preset time period based on the gas concentration sequence, personnel behavior event sequence, monitoring feedback sequence, comprehensive risk value sequence, and cross-modal consistency index sequence within a preset sliding time window. When the predicted comprehensive risk level reaches the warning level, alarm level, or work stoppage level within a preset time in the future, the system will trigger a risk warning in advance or enter the cloud review mode.
7. A multimodal risk fusion management and control system for confined space operations according to claim 1, characterized in that, The multi-source data fusion module includes a data synchronization unit, a character layer generation unit, and a hardware encoding unit; The data synchronization unit associates gas concentration data, alarm status and corresponding video frames based on timestamps. The character layer generation unit converts gas concentration data, alarm status, and time information into character layers and draws them onto a preset area of the video frame pixel matrix. The hardware encoding unit performs secondary encoding on the superimposed video frames, and the resulting integrated video stream is used for local display on the host, video recording and storage, and cloud viewing.
8. A multimodal risk fusion management and control system for confined space operations according to claim 1, characterized in that, The system also includes a monitoring mechanism for guardians' absence from their posts; In the formal operation state, the edge computing host pushes an interactive confirmation request to the monitoring terminal according to a preset period T; If the guardian does not provide a confirmation signal within the preset response time M, and does not receive a valid feedback signal within N consecutive confirmation cycles, a guardian absence alarm will be generated and the alarm linkage module will be triggered. The all-in-one gas sampler is also equipped with an SOS physical alarm button. When the SOS physical alarm button is triggered, the alarm linkage module simultaneously triggers the host-side audible and visual alarm and the remote monitoring-side alarm, and writes the SOS trigger time into the evidence report.
9. A multimodal risk fusion management and control system for confined space operations according to claim 1, characterized in that, The remote cloud platform is configured as follows: Receive structured risk events, key video slices, and sensor data uploaded by the edge computing host; When the local comprehensive risk level is at the preset review level, the cross-modal consistency index is lower than the preset consistency threshold, or the predicted comprehensive risk level reaches the preset review level, the multimodal model or preset rule engine is invoked to perform risk review. The risk review results are returned to the edge computing host, including maintaining the original level, upgrading the level, downgrading the level, or marking it as a suspected false alarm.
10. A method for a multimodal risk fusion management and control system for confined space operations as described in any one of claims 1-9, characterized in that, The steps are as follows: S1. Collect images of on-site personnel and verify their identities. Once the identity verification is successful, proceed to ventilation mode. S2. Perform the ventilation process while the ventilation is in progress, and enter the pre-sampling and testing state after receiving the ventilation confirmation signal from the supervisor. S3. Collect gas concentration data in the pre-collection detection state. When the gas concentration data meets the preset safety threshold and a confirmation signal is received from the supervisor, enter the formal operation state. S4. In the formal operation state, synchronously collect video stream, gas concentration data, personnel behavior scene recognition results, monitoring confirmation feedback and SOS button trigger signal; S5. Calculate the comprehensive risk level based on gas concentration data, personnel behavior scene recognition results, and monitoring confirmation feedback, and use the SOS button trigger signal as an independent emergency trigger signal; S6. Superimpose gas concentration data and alarm status onto the video frame pixel matrix and encode to generate an integrated video stream; S7. Trigger local alarms, remote alarms, or work stoppage control based on the comprehensive risk level, and directly trigger alarm linkage and record storage based on the SOS button trigger signal; S8. After the operation is completed, an operation report is generated that includes identity verification records, state machine transition records, gas detection data, risk assessment results, alarm records, and video clips.