Dangerous environment identification method based on visual language model and dynamic scene

Through visual language model building a causal logic constraint set in dynamic scenarios, the problem of insufficient causal relationship modeling in the existing technology is solved, accurate identification and timely warning of dangerous events are achieved, and the accuracy and reliability of hazardous environment identification are improved.

CN120279499AInactive Publication Date: 2025-07-08SHENZHEN HIGHLAND BARLEY INFORMATION TECH CO LTD
View PDF 0 Cites 18 Cited by

Patent Information

Application Number
CN202510765200.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology fails to effectively model the event causal chain in the identification of hazardous environments in dynamic scenarios, resulting in the risk identification results being limited to superficial phenomena, making it difficult to accurately determine the root cause and evolution trend of risks, and misjudgment or early warning lag, which affects the timeliness and reliability of hazard handling.

Method used

Through the visual language model combined with dynamic scenes, the target motion trajectory and scene change sequence of the visual data flow are extracted, the causal logic constraint set is constructed, the spatiotemporal correlation verification and causal intensity evaluation are carried out, the causal chain of dangerous events is generated, and the identification results are generated based on the preset risk threshold.

Benefits of technology

It significantly improves the accuracy and reliability of hazardous environment identification, can deeply analyze the causal trigger mechanism between events, realize risk assessment from a global perspective, and improve the timeliness and reliability of hazard warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279499A_ABST
    Figure CN120279499A_ABST
Patent Text Reader

Abstract

The invention discloses a dangerous environment identification method based on a visual language model and a dynamic scene, particularly relates to the technical field of intelligent risk early warning, and is used for solving the problems of misjudgment and early warning lag caused by lack of event causal chain modeling in the dynamic scene in existing risk identification. According to the method, visual data flow and language description data are fused, a target motion track and a scene change sequence are extracted, a causal logic constraint set is constructed, space-time relevance is verified, a dangerous event causal chain is generated after high-confidence candidate events are screened, and finally a dynamic recognition result is generated based on causal chain overall risk quantification. Through dynamic rule adaptation and causal chain evolution rule matching, the random coexistence event and the real causal relationship are effectively distinguished, the problem of misjudgment caused by static rule stiffness in a traditional method is solved, and the reliability and the disposal timeliness of danger early warning in a complex dynamic scene are improved in combination with a global risk assessment mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent risk early warning, and more specifically, to a method for identifying dangerous environments based on a vision-language model and dynamic scenarios. Background Art

[0002] In the field of dangerous environment identification, existing technologies usually rely on independent analysis of visual or language modalities. For example, fires and obstacles are identified through video surveillance, or the risk of gas leakage is judged through sensor text. Although multi-modal methods have attempted to fuse visual and language data, in dynamic scenarios, the evolution process of dangerous signals involves complex spatio-temporal correlations (such as accident chain reactions and disaster diffusion paths). Existing methods often rely on apparent feature matching and lack in-depth modeling of event causal chains.

[0003] In multi-modal dangerous identification, existing technologies do not fully consider the causal relationships between events in dynamic scenarios (such as the relevance between accident causes and consequences), resulting in the limitation of dangerous identification results to surface phenomena and making it difficult to accurately determine the root causes and evolution trends of risks. For example, when multiple dangerous signals occur concurrently in a dynamic environment, existing methods cannot distinguish causal associations from random co-occurrences, leading to misjudgments or early warning delays, seriously affecting the timeliness and reliability of dangerous situation handling. Summary of the Invention

[0004] To overcome the above defects of the prior art, an embodiment of the present invention provides a method for identifying dangerous environments based on a vision-language model and dynamic scenarios to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions: A method for identifying dangerous environments based on a vision-language model and dynamic scenarios, comprising the following steps: S1. Obtain the visual data stream of the dynamic scenario and the corresponding language description data; S2. Extract the spatio-temporal features of dynamic objects from the visual data stream to generate the object movement trajectories and the scene change sequences; extract the causal rule texts associated with dangerous events from the language description data to construct a causal logic constraint set; S3. Perform spatio-temporal association verification on the object movement trajectories and the causal logic constraint set to screen out a candidate set of dynamic dangerous events that satisfy the causal constraints; S4. Evaluate the causal strength of the candidate set of dynamic dangerous events, calculate the spatio-temporal compactness and verify the logical completeness to generate a candidate set with high confidence; S5. Perform logical reasoning on the candidate set with high confidence based on the scene change sequences to generate a causal chain of dangerous events; S6. Generate an identification result of the dynamic dangerous environment according to the relationship between the causal chain of dangerous events and a preset risk threshold.

[0006] In a preferred embodiment, acquiring the visual data stream of the dynamic scene and the corresponding language description data includes: Performing multi-source synchronous acquisition on the dynamic scene to generate a visual data stream and language description data containing timestamps; Performing resolution adaptive adjustment and abnormal frame filtering on the visual data stream to generate a standardized visual data stream; Performing noise text cleaning and key semantic extraction on the language description data to generate structured language description data; Aligning the standardized visual data stream and the structured language description data according to the timestamps and storing them as an associated data group.

[0007] In a preferred embodiment, performing dynamic target spatio-temporal feature extraction on the visual data stream to generate a target motion trajectory and a scene change sequence, including: Performing dynamic target detection and tracking on the standardized visual data stream to generate the bounding box coordinates of the dynamic target and the displacement vectors between consecutive frames; Calculating the motion speed and direction of the dynamic target based on the displacement vectors between consecutive frames, and generating a target motion trajectory through a trajectory smoothing algorithm; Extracting the illumination intensity change gradient, background region motion vector, and dynamic target aggregation density between adjacent frames in the standardized visual data stream, and fusing them to generate a scene change sequence.

[0008] In a preferred embodiment, extracting causal rule texts associated with dangerous events from the language description data and constructing a causal logic constraint set, including: Performing semantic role annotation on the structured language description data to identify the action subject, object of action, and constraint conditions of the dangerous event; Converting the semantic role annotation results into spatio-temporal constraint rules according to a preset causal logic template to construct a causal logic constraint set.

[0009] In a preferred embodiment, performing spatio-temporal association verification on the target motion trajectory and the causal logic constraint set to screen out a candidate set of dynamic dangerous events, including: Calculating the regional overlap degree according to the spatial coverage range of the target motion trajectory and the spatial constraint conditions in the causal logic constraint set, and screening out candidate trajectories with a spatial coverage exceeding a preset overlap threshold; Verifying the temporal matching degree based on the time window of the target motion trajectory and the time constraint conditions in the causal logic constraint set, and retaining candidate trajectories with a time overlap ratio exceeding a preset temporal threshold; For the candidate trajectories passing the spatio-temporal coverage verification, performing semantic consistency matching in combination with the logical conditions in the causal logic constraint set, and eliminating the trajectories with logical conflicts to generate a candidate set of dynamic dangerous events.

[0010] In a preferred embodiment, a causal strength assessment is performed on the dynamic hazardous event candidate set, the spatio-temporal compactness is calculated, and the logical completeness is verified to generate a high-confidence candidate set, including: Calculate the spatio-temporal compactness of each trajectory in the dynamic hazardous event candidate set, including statistically analyzing the spatial distribution concentration and time series continuity of the trajectory parameters to generate a spatio-temporal compactness score; Verify the logical completeness of the dynamic hazardous event candidate set, including checking whether the causal logic conditions of the candidate trajectories are consistent with the evolution law of the scenario change sequence, and eliminating the trajectories with logical contradictions; Screen the candidate trajectories with a spatio-temporal compactness score exceeding the preset spatio-temporal compactness threshold and a logical completeness verification result of passing to generate a high-confidence candidate set.

[0011] In a preferred embodiment, logical reasoning is performed on the high-confidence candidate set based on the scenario change sequence to generate a causal chain of hazardous events, including: Extract the evolution pattern features of the scenario change sequence, including analyzing the parameter trends and mutation points in the scenario change encoding to generate a spatio-temporal evolution pattern set; Perform temporal alignment and causal association matching between the candidate events in the high-confidence candidate set and the spatio-temporal evolution pattern set to identify the causal triggering relationships between the candidate events; Chain the candidate events according to the causal triggering relationships to generate a causal chain of hazardous events.

[0012] In a preferred embodiment, the nodes of the causal chain of hazardous events are candidate events, and the edges are the spatio-temporal constraint conditions of the triggering relationships.

[0013] In a preferred embodiment, according to the relationship between the causal chain of hazardous events and the preset risk threshold, a dynamic hazardous environment identification result is generated, including: Analyze the node attributes and edge attributes of the causal chain of hazardous events, and extract the propagation path depth and node risk weights of the causal chain of hazardous events; Calculate the overall risk value of the causal chain according to the product of the propagation path depth and the node risk weights; Compare the overall risk value of the causal chain with the preset risk threshold, and generate a dynamic hazardous environment identification result when the overall risk value of the causal chain exceeds the preset risk threshold.

[0014] In a preferred embodiment, the dynamic hazardous environment identification result includes the causal chain identifier, the risk level, and the trigger location information.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. Through the deep integration of visual and language data and the iterative optimization of dynamic causal rules, the accuracy and reliability of dangerous environment recognition have been significantly improved. Aiming at the complex spatio-temporal correlation characteristics of the event causal chain in dynamic scenarios, by extracting the causal logic rules in the language description and combining the target motion trajectory and scene change sequence of the visual data stream, a multi-modal collaborative causal reasoning framework is constructed, enabling the recognition of dangerous events not only to rely on apparent feature matching, but also to deeply analyze the causal triggering mechanism between events, avoiding misjudgment problems caused by the lack of logical constraints in traditional methods, and thus effectively improving the timeliness of danger warning.

[0016] 2. Based on the overall risk assessment mechanism of the dynamic causal chain, the problem that traditional methods are overly sensitive to or miss local danger signals is solved. By constructing the causal chain of dangerous events and quantifying its overall risk, the evolution trend of dangerous environments can be identified from a global perspective. Compared with the isolated detection of danger signals in existing technologies, through multi-dimensional causal reasoning and dynamic threshold adaptation, the leap from single-event recognition to systematic risk prevention and control is achieved, providing a reliable basis for danger disposal in complex dynamic scenarios. Brief Description of the Drawings

[0017] Figure 1 It is a flowchart of a method for recognizing a dangerous environment based on a visual language model and a dynamic scene according to the present invention; Figure 2 It is a flowchart of screening a candidate set of dynamic dangerous events according to the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment: Figure 1 A method for recognizing a dangerous environment based on a visual language model and a dynamic scene according to the present invention is given, including the following steps: S1. Obtain the visual data stream of the dynamic scene and the corresponding language description data; S2. Extract the spatio-temporal features of dynamic targets from the visual data stream to generate the target motion trajectory and the scene change sequence; extract the causal rule text associated with dangerous events from the language description data to construct a causal logic constraint set; S3. Perform spatio-temporal correlation verification on the target motion trajectory and the causal logic constraint set to screen out a candidate set of dynamic dangerous events that meet the causal constraints; S4. Evaluate the causal strength of the dynamic hazard event candidate set, calculate the spatio-temporal tightness, and verify the logical completeness to generate a high-confidence candidate set; S5. Based on the scenario change sequence, perform logical reasoning on the high-confidence candidate set to generate a causal chain of hazard events; S6. Generate the identification result of the dynamic hazard environment according to the relationship between the causal chain of hazard events and the preset risk threshold.

[0020] S1. Obtain the visual data stream and the corresponding language description data of the dynamic scenario. The specific implementation is as follows: When performing multi-source synchronous acquisition of the dynamic scenario, use multiple visual acquisition devices and language acquisition devices deployed in the dynamic scenario. The visual acquisition devices include cameras and infrared sensors, and the language acquisition devices include microphones and text input interfaces. Each device controls the acquisition time error within milliseconds through a unified time synchronization protocol to generate a visual data stream and a language description data containing timestamps. The visual data stream is a continuous sequence of video frames, and the language description data is a mixed data stream containing sensor readings, manually input text, and speech-to-text. The timestamp is accurate to milliseconds and synchronized with the international standard time. The spatial layout of the visual acquisition devices and the language acquisition devices is determined according to the distribution density of the scenario hazard sources. The device spacing is less than the preset safety distance threshold, and the preset safety distance threshold is 1.5 times the effective monitoring radius of the hazard source. The effective monitoring radius of the hazard source is obtained through device calibration testing.

[0021] When adaptively adjusting the resolution of the visual data stream, dynamically adjust the resolution of the video frames according to the changes in the illumination intensity and the target distance of the dynamic scenario. Specifically, when the illumination intensity is lower than 100 lux, adjust the resolution to 1280×720 pixels and enable the contrast enhancement algorithm. When the illumination intensity is higher than 1000 lux, adjust the resolution to 1920×1080 pixels and enable the dynamic range compression algorithm. When the target distance exceeds 50 meters, adjust the resolution to 640×480 pixels and enable the super-resolution reconstruction algorithm to generate a standardized visual data stream. The illumination intensity is obtained in real time through the built-in photosensitive sensor of the camera, and the target distance is calculated through the binocular vision ranging algorithm. The contrast enhancement algorithm realizes the enhancement of dark details by adjusting the distribution of the image gray histogram. The dynamic range compression algorithm realizes the limitation of the difference between the maximum and minimum values of the image pixel brightness. The super-resolution reconstruction algorithm is realized through the bicubic interpolation algorithm. The output format of the standardized visual data stream is YUV420.

[0022] When performing abnormal frame filtering on the visual data stream, abnormal frames are detected by calculating the clarity score and the luminance distribution histogram of the video frame. Specifically, when the clarity score is lower than 0.7 or the proportion of pixels with luminance values greater than 200 in the luminance distribution histogram exceeds 5%, the video frame is determined to be an abnormal frame and discarded. The clarity score is the average value of the image gradient magnitude, which is obtained by calculating the square root of the sum of the squares of the horizontal gradient and the vertical gradient of each pixel point in the image. The horizontal gradient is calculated by convolving with the Sobel operator [-1, 0, 1], and the vertical gradient is calculated by convolving with the Sobel operator [-1, 0, 1]^T. Only the video frames that meet the clarity and luminance requirements are retained in the standardized visual data stream after abnormal frame filtering. When the number of consecutive discarded frames exceeds 10% of the total number of frames, a device calibration signal is triggered.

[0023] When performing noise text cleaning on the language description data, special symbols, meaningless characters, and redundant stop words in the text are removed through regular expression matching and stop word list filtering. For example, the exclamation mark and the hash mark in "#Alarm! Methane concentration is abnormal" are deleted, and "Methane concentration is abnormal" is retained as valid text. At the same time, repeated modal particles and filler words in the speech-to-text data are deleted. For example, "um... that... the methane concentration seems to exceed the standard" is deleted, and "Methane concentration exceeds the standard" is retained as valid text, generating the preliminarily cleaned language description data. The stop word list contains 100 predefined invalid words, and the regular expression matching rule is " / [^\u4e00-\u9fa5a-zA-Z0-9] / ".

[0024] When performing key semantic extraction on the preliminarily cleaned language description data, the core predicate-object structure in the text is identified through dependency syntax analysis, and semantic fragments associated with dangerous events are selected based on the dangerous event keyword list. For example, "Methane concentration exceeds the safety threshold" is extracted from "The sensor detects that the methane concentration exceeds the safety threshold, and it is recommended to evacuate immediately" as the causal rule text. The dangerous event keyword list includes 50 predefined dangerous action words such as "leakage", "exceeding the standard", "collision", "violation", etc. Dependency syntax analysis is achieved by identifying the core verb in the sentence and its associated subject and object. The generated causal rule text is structured language description data in the format of a "subject-action-object" triple, such as "Methane concentration - exceeds - safety threshold". When there are multiple dangerous action words in the sentence, the highest-level action word is selected according to the priority of the danger level.

[0025] When aligning the standardized visual data stream with the structured language description data by timestamp, a sliding window matching algorithm is adopted. The timestamp of each video frame in the visual data stream is subtracted from the timestamp of the structured language description data. When the timestamp difference is less than 33 milliseconds, the two are bound into an associated data group at the same time point. The step size of the sliding window is 1 millisecond. The successfully matched associated data groups are stored in JSON format, with the key being the timestamp string and the value being the corresponding standardized visual data stream frame and structured language description data entry. The storage path is compatible with the input interface of the subsequent spatio-temporal association verification step.

[0026] Extract the spatio-temporal features of dynamic objects from the visual data stream to generate the object motion trajectory and scene change sequence. The specific implementation is as follows: When performing dynamic object detection and tracking on the standardized visual data stream, the motion vectors of pixel points in consecutive video frames are calculated by the optical flow method to identify the bounding box coordinates of dynamic objects. The dynamic object is a pixel area with a moving speed exceeding the preset speed threshold, which is obtained by statistically analyzing the scene historical data. For example, it is set to 5 meters per second in a traffic scene and 2 meters per second in an industrial scene. The bounding box coordinates of the dynamic object are generated by clustering adjacent motion vectors through a clustering algorithm. The distance threshold of the clustering algorithm is 10 pixels. The displacement vector between consecutive frames is the difference in the center point coordinates of the bounding box of the same dynamic object in adjacent frames. The calculation accuracy of the displacement vector is at the pixel level. When the dynamic object is lost, it is re-initialized for tracking through a feature matching algorithm, which is implemented by comparing the SIFT feature descriptors of the target area. The matching similarity threshold is set to 0.8. When the similarity is lower than the threshold, it is determined that the object is lost and re-detection is triggered.

[0027] When calculating the motion speed and direction of a dynamic object based on the displacement vector between consecutive frames, the instantaneous speed is obtained by dividing the displacement vector by the frame interval time, which is determined according to the frame rate of the standardized visual data stream. For example, when the frame rate is 30 frames per second, the frame interval time is 33 milliseconds. The motion direction is the angular value of the displacement vector, which is obtained by calculating the ratio of the horizontal and vertical displacement components through the arctangent function. The trajectory smoothing algorithm generates the object motion trajectory by weighted averaging the motion speed and direction data of consecutive 5 frames. The weighting coefficient is dynamically adjusted according to the object motion consistency. For example, when the variance of the object speed change is less than 0.1, the weight of the historical frame is 0.7 and the weight of the current frame is 0.3. When the variance is greater than 0.5, the weight of the historical frame is 0.3 and the weight of the current frame is 0.7. The preset variance threshold is calibrated through experimental data. The object motion trajectory after trajectory smoothing is stored as a key-value pair of timestamp sequence and coordinate points, and the timestamp accuracy is at the millisecond level.

[0028] When extracting the gradient of the illumination intensity change between adjacent frames in the standardized visual data stream, the video frames are converted into grayscale images through grayscale processing. The difference matrix of adjacent grayscale images is calculated, and the change amplitude of the grayscale value of each pixel point in the difference matrix is the gradient of the illumination intensity change. The range of the gradient value is from 0 to 255. The motion vector of the background area is calculated by the block matching algorithm to obtain the overall motion offset of the non-dynamic target area in the video frame. The block size in the block matching algorithm is 16×16 pixels, and the search range is 32×32 pixels. The dynamic target aggregation density is the ratio of the total area of the bounding boxes of the dynamic targets in the scene to the area of the scene area. The dimension of the ratio is a percentage. When fusing to generate the scene change sequence, the average value of the illumination intensity change gradient, the magnitude of the motion vector of the background area, and the dynamic target aggregation density are linearly combined into the scene change code according to the preset weights. The preset weights are set according to the scene type. For example, the aggregation density weight in the traffic scene is 0.6, and the illumination gradient weight in the industrial scene is 0.5. The weight values are optimized on the historical data set through the grid search method. The output format of the scene change code is a three-dimensional vector, and the dimensions correspond to the illumination, background motion, and aggregation density parameters respectively.

[0029] Extract the causal rule text associated with the dangerous event from the language description data and construct a causal logic constraint set. The specific implementation is as follows: When performing semantic role annotation on the structured language description data, the core predicates and their associated subjects and objects in the sentence are identified through dependency syntactic analysis. For example, from the sentence "Vehicle A changed lanes illegally at the intersection, causing Vehicle B to brake suddenly", the subject "Vehicle A", the predicate "changed lanes illegally", the object "intersection", and the modifier "causing Vehicle B to brake suddenly" are extracted. The action subject of the dangerous event is the entity corresponding to the subject, the object of action is the entity corresponding to the object, and the constraint condition is the causal relationship description in the modifier. The semantic role annotation result is stored as a quadruple of "action subject - predicate - object of action - constraint condition". The dependency syntactic analysis is achieved by identifying the core verb and its dependent components in the sentence. When there are multiple predicates in the sentence, the highest-level predicate is selected according to the priority of the dangerous event. The dangerous event priority table is predefined as collision class > leakage class > violation class.

[0030] When converting the semantic role annotation results into spatio-temporal constraint rules according to a preset causal logic template, the preset causal logic template is "IF the action subject executes a predicate in the object area THEN trigger the constraint condition". For example, converting the quadruple "Vehicle A - illegal lane change - intersection - causing Vehicle B to brake suddenly" into "IF Vehicle A executes an illegal lane change in the intersection area THEN trigger the collision risk". The spatio-temporal constraint rules are further added with time window and spatial range conditions. For example, the illegal lane change operation needs to be completed within 2 seconds and cross more than 50% of the lane line. The starting point of the time window is determined according to the starting timestamp of the action subject's movement trajectory, and the spatial range condition is calculated through the bounding box coordinates of the object area. When constructing the causal logic constraint set, the spatio-temporal constraint rules are classified and stored according to the types of dangerous events, such as traffic accident rules and leakage rules. The storage format is a JSON structure, including time window, spatial range, and trigger condition fields. The storage path is compatible with the input interface of the subsequent spatio-temporal association verification step.

[0031] Figure 2 The flowchart for screening the candidate set of dynamic dangerous events of the present invention is given. S3. Perform spatio-temporal association verification on the target movement trajectory and the causal logic constraint set to screen out the candidate set of dynamic dangerous events that meet the causal constraints. The specific implementation is as follows: When calculating the regional overlap degree according to the spatial coverage range of the target movement trajectory and the spatial constraint conditions in the causal logic constraint set, the spatial coverage range of the target movement trajectory is determined by the minimum bounding rectangle of the trajectory coordinate points. The calculation method of the minimum bounding rectangle is to traverse the maximum horizontal coordinate value, minimum horizontal coordinate value, maximum vertical coordinate value, and minimum vertical coordinate value of the trajectory coordinate points to generate a rectangular area containing all the trajectory points. The spatial constraint conditions in the causal logic constraint set are predefined geofence polygons, and the vertex coordinates of the geofence polygons are pre-annotated through the scene map data. For example, in a traffic accident scene, the fence vertices of the intersection area are a quadrilateral formed by longitude 116.4074 degrees, latitude 39.9042 degrees to longitude 116.4080 degrees, latitude 39.9045 degrees. The regional overlap degree calculation is the intersection area of the bounding rectangle and the geofence polygon divided by the union area of the geofence polygon. The calculation of the intersection area and the union area is realized through polygon Boolean operations. When the proportion of the intersection area exceeds the preset spatial threshold, it is determined as a spatial match. The preset spatial threshold is obtained according to the statistical data of historical dangerous events. For example, it is set to 60% in a traffic accident scene and 50% in an industrial leakage scene. The historical dangerous event data is the minimum effective overlap ratio when a dangerous event is triggered in the same type of scene in the past year. For a trajectory partially located within the geofence, the overlapping area ratio is estimated by linear interpolation, with an interpolation step of 1 pixel, and the overlapping area ratio after interpolation is retained with two decimal places of precision.

[0032] When verifying the temporal matching degree based on the time window of the target motion trajectory and the time constraint conditions in the causal logic constraint set, the time window of the target motion trajectory is the interval between the trajectory start timestamp and the end timestamp. The precision of the timestamp is in milliseconds and is synchronized with the international standard time. The time constraint conditions in the causal logic constraint set are preset time intervals or relative time windows. When the time interval is an absolute time, such as from 09:00:00 to 09:30:00, and the relative time window is the duration after event triggering, such as 2 seconds. The temporal matching degree verification is the ratio of the intersection duration of the two time intervals to the total duration of the time constraint conditions. The calculation of the intersection duration is achieved by finding the intersection of the timestamp intervals. For example, if the trajectory time window is from 09:00:00 to 09:10:00 and the time constraint condition is from 09:05:00 to 09:15:00, then the intersection duration is 5 minutes, the total duration of the time constraint condition is 10 minutes, and the overlapping ratio is 50%. When the proportion of the intersection duration exceeds the preset time threshold, it is determined as time matching. The preset time threshold is set according to the scenario response requirements. For example, it is set to 80% in the traffic accident scenario and 70% in the industrial leakage scenario. The scenario response requirement is the minimum effective warning time coverage rate of the dangerous event handling system. The calculation of the intersection duration of the time window is accurate to the millisecond level. When the time constraint condition is an absolute time, it is aligned with the international standard timestamp, and when it is a relative time, it is calculated by offsetting based on the event triggering moment. For example, if the event triggering moment is 09:00:00 and the relative time window is 2 seconds after triggering, then the time constraint condition is from 09:00:00 to 09:00:02.

[0033] When performing semantic consistency matching on the candidate trajectories verified by spatio-temporal coverage with the logical conditions in the causal logic constraint set, the logical conditions are the predicate-argument constraint relationships defined in the causal logic constraint set. The format of the predicate-argument constraint relationship is "agent of action - action - object of action - constraint parameter", for example, "vehicle - lane change - intersection - turn signal off". Semantic consistency matching is achieved by verifying whether the associated attributes of the target motion trajectory satisfy the logical conditions. The associated attributes include the target motion direction, speed change pattern, and external device status. The target motion direction is calculated through the tangent angle of the trajectory coordinate points, the speed change pattern is calculated through the displacement vector difference between adjacent timestamps, and the external device status is obtained in real time through in-vehicle sensors or monitoring system interfaces. For example, to verify whether the turn signal status in the lane change trajectory remains off, the turn signal status is obtained through in-vehicle CAN bus signals or video analysis results. Trajectories with logical conflicts are marked and excluded by the rule engine. The decision logic of the rule engine is that if the trajectory attributes conflict with the necessary parameters of the logical conditions, it is determined to be inconsistent. For example, a lane change trajectory with the turn signal on conflicts with the condition of "turn signal off". The storage format of the dynamic dangerous event candidate set is a structured data table containing trajectory identifiers, spatio-temporal matching degrees, and logical conditions. The fields of the structured data table include trajectory ID, spatial overlap degree, time overlap ratio, logical condition ID, and matching status. The storage path is compatible with the input interface of the subsequent causal strength evaluation step. Each row of data in the structured data table corresponds to a candidate trajectory that has passed the verification.

[0034] S4. Evaluate the causal strength of the dynamic dangerous event candidate set, calculate the spatio-temporal tightness, and verify the logical completeness to generate a high-confidence candidate set. The specific implementation is as follows: When calculating the spatio-temporal compactness of each trajectory in the dynamic hazardous event candidate set, the spatial distribution concentration is calculated by the position variance of all coordinate points of the trajectory. The position variance of the trajectory coordinate points is achieved through the following steps: traverse the horizontal and vertical coordinate values of all coordinate points of the trajectory, calculate the arithmetic mean of the horizontal coordinate values and the arithmetic mean of the vertical coordinate values respectively. The horizontal coordinate variance is the average of the sum of the squares of the differences between each horizontal coordinate value and the arithmetic mean, and the vertical coordinate variance is the average of the sum of the squares of the differences between each vertical coordinate value and the arithmetic mean. The spatial distribution concentration score is the weighted sum of the horizontal coordinate variance and the vertical coordinate variance. The weights of the weighted sum are preset according to the scenario type. For example, in the traffic scenario, the weight of the horizontal coordinate variance is 0.6 and the weight of the vertical coordinate variance is 0.4. In the industrial scenario, the weights of the horizontal and vertical coordinate variances are both 0.5. The weight values are obtained based on the historical data analysis of the direction sensitivity in different scenarios. The historical data is the statistical result of the spatial distribution of hazardous events in the same type of scenario in the past three years. The time series continuity is calculated by the mean square error of the adjacent timestamp intervals. The mean square error of the adjacent timestamp intervals is achieved through the following steps: obtain the frame rate of the standardized visual data stream as the theoretical average time interval. For example, when the frame rate is 30 frames per second, the theoretical average interval is 33 milliseconds. Calculate the difference between the actual interval of adjacent timestamps in the trajectory and the theoretical average interval. The average of the sum of the squares of the differences is the mean square error value. The time series continuity score is the reciprocal of the mean square error value normalized to the range of 0 to 1. The spatio-temporal compactness score is the weighted sum of the spatial distribution concentration score and the time series continuity score. The weights of the weighted sum are set according to the type of hazardous event. For example, for traffic accident events, the spatial weight is 0.7 and the time weight is 0.3. For leakage events, the spatial weight is 0.5 and the time weight is 0.5. The weight values are optimized on a dataset containing 1,000 historical trajectories through the grid search method. The search range of the grid search method is from 0.1 to 0.9, and the step size is 0.1.

[0035] When verifying the logical completeness of the dynamic dangerous event candidate set, check whether the causal logic conditions of the candidate trajectories are consistent with the evolution law of the scenario change sequence. The causal logic conditions are the "IF-THEN" rules defined in the causal logic constraint set. For example, "IF the leakage diffusion direction is southeast wind THEN prohibit personnel from entering the southeast area". The evolution law of the scenario change sequence is determined by analyzing the parameter trends in the scenario change encoding. The parameter trends include the spatial gradient direction of the dynamic target aggregation density, the propagation direction of the illumination intensity change gradient, and the mainstream direction of the background motion vector. The spatial gradient direction is calculated by the distribution density difference of the dynamic target aggregation density. For example, the scenario is divided into a 10×10 grid, the density value of each grid is calculated and a density gradient field is generated. The evolution direction is the direction in which the gradient field drops fastest. The leakage diffusion direction is the gradient descent vector in the southeast direction. The movement direction of the candidate trajectory is calculated by the tangent angle of the trajectory coordinate points. The tangent angle is the arctangent value of the horizontal displacement and the vertical displacement of the two coordinate points at the end of the trajectory. If the angle between the movement direction of the candidate trajectory and the scenario evolution direction exceeds the preset angle threshold, it is determined as a logical contradiction. The preset angle threshold is set according to the scenario type. For example, it is set to 30 degrees in the industrial leakage scenario and 45 degrees in the traffic accident scenario. The trajectories with logical contradictions are marked and removed by the rule engine. The decision logic of the rule engine is to compare the angle difference with the preset threshold. When the threshold is exceeded, a "contradiction" label is generated, and when it is not exceeded, a "passed" label is generated. The decision result is bound and stored with the metadata of the candidate trajectory.

[0036] When screening candidate trajectories with a spatio-temporal tightness score exceeding the preset spatio-temporal tightness threshold and a logical completeness verification result of passed, the preset spatio-temporal tightness threshold is the quantile value statistically obtained from the historical dangerous event data. For example, in the traffic accident scenario, the score of the top 20% of the historical data is taken as the threshold 0.8, and in the industrial leakage scenario, the score of the top 30% of the historical data is taken as the threshold 0.7. The historical dangerous event data is the record of the spatio-temporal tightness scores when dangerous events occurred in the same type of scenario in the past five years. The screening process is achieved by traversing the candidate trajectory dataset and comparing them one by one. For example, if the spatio-temporal tightness score of a certain trajectory is 0.85 and the logical label is "passed", it is added to the high-confidence candidate set. The storage format of the high-confidence candidate set is a structured data table, and the fields include the trajectory number, spatio-temporal tightness score, logical label, scenario change encoding, and associated timestamp. The row data of the structured data table is compatible with the input interface of the subsequent logical reasoning steps. The primary key of the data table is the trajectory number, and the foreign keys are the scenario change encoding and the timestamp, ensuring the efficiency of the associated query for generating the causal chain in the subsequent steps.

[0037] S5. Perform logical reasoning on the high-confidence candidate set based on the scenario change sequence to generate the causal chain of dangerous events. The specific implementation is as follows: When extracting the evolution pattern features of the scene change sequence, the parameter trend in the scene change encoding is analyzed by calculating the moving window mean. The size of the moving window is dynamically adjusted according to the time length of the scene change sequence. For example, when the time length is 10 seconds, the window size is 1 second; when the time length is 1 minute, the window size is 5 seconds. The dynamic adjustment rule is that the window size takes one-tenth of the total time length and is rounded down to the second-level unit. The parameter trend includes the average value of the change gradient of the light intensity within the window, the average value of the modulus of the background area motion vector, and the average value of the dynamic target aggregation density. The average value of the change gradient of the light intensity is achieved by calculating the mean value of the gray difference matrix of all video frames within the window. The average value of the modulus of the background area motion vector is the arithmetic mean of the lengths of the background motion vectors of each frame within the window. The average value of the dynamic target aggregation density is the average value of the ratio of the total area of the dynamic target bounding boxes to the area of the scene area in each frame within the window. The mutation point detection is achieved by calculating the derivative change of the parameters of adjacent windows. The derivative change is the difference between the parameter value of the latter window minus the parameter value of the former window divided by the difference between the start timestamp of the latter window minus the start timestamp of the former window. For example, if the dynamic target aggregation density increases from 0.3 to 0.8 in adjacent windows and the timestamp difference is 5 seconds, then the derivative is (0.8 - 0.3) / 5 = 0.1. When the absolute value of the derivative exceeds the preset mutation threshold, it is marked as a mutation point. The preset mutation threshold is obtained based on the statistical data of historical dangerous events. For example, it is set to 0.5 in the traffic accident scene and 0.3 in the industrial leakage scene. The historical dangerous event data is the minimum valid value of the absolute value of the derivative when a dangerous event is triggered in the same type of scene in the past three years. The spatio-temporal evolution pattern set is stored as a structured data table containing trend parameters, mutation point positions, and timestamps. The primary key of the data table is the pattern number, and the foreign key is the scene change encoding. The generation rule of the pattern number is "scene type_timestamp_parameter type", for example, "traffic_20231001120000_aggregation density".

[0038] When aligning the candidate events in the high-confidence candidate set with the set of spatio-temporal evolution patterns in terms of time sequence and performing causal association matching, for time sequence alignment, the dynamic time warping algorithm is used to perform elastic matching between the time stamp sequence of the candidate event and the time stamp sequence of the spatio-temporal evolution pattern. The path cost function of the dynamic time warping algorithm is the sum of the squares of the time stamp offsets. The allowed time stamp offset tolerance is set according to the scene type. For example, the tolerance is 100 milliseconds in the traffic accident scene and 500 milliseconds in the industrial leakage scene. The basis for setting the tolerance is the requirement for scene response timeliness. Since traffic accidents require quick response, the tolerance is small, while industrial leakage tolerates a slightly higher delay, so the tolerance is large. For causal association matching, the rule engine is used to verify whether the causal logic conditions of the candidate event are consistent with the trend parameters in the spatio-temporal evolution pattern. The decision logic of the rule engine is that if the value range of the parameters in the causal logic conditions of the candidate event completely contains the trend parameter value in the spatio-temporal evolution pattern, it is determined to be a match. For example, if the causal logic condition of the candidate event is "the aggregation density is greater than 0.6" and the trend parameter in the spatio-temporal evolution pattern is 0.7, it is determined to be a match. The matching result is a boolean label "match" or "no match", and the label field is bound and stored with the metadata of the candidate event. The storage field name is "causal matching status". The metadata of the candidate event includes the event number, spatio-temporal tightness score, and associated scene change code.

[0039] When generating the causal chain of dangerous events by concatenating candidate events in a causal trigger relationship chain, the causal trigger relationship includes a time sequence relationship and a spatial coverage relationship. The time sequence relationship is verified by the time stamp sequence of candidate events. The time stamp sequence verification rule is that the end time stamp of the previous event is earlier than the start time stamp of the subsequent event. For example, if the end time stamp of event A is 09:00:05 and the start time stamp of event B is 09:00:06, then the time sequence relationship is satisfied. The spatial coverage relationship is verified by the intersection area of the position regions of candidate events. The calculation of the intersection area of the position regions is the intersection area of the minimum bounding rectangles of two events divided by the union area. When the proportion of the intersection area exceeds the preset coverage threshold, it is determined as spatially covered. The preset coverage threshold is set according to the scenario type. For example, it is 30% in the traffic accident scenario and 50% in the industrial leakage scenario. The threshold setting is based on the minimum effective spatial overlap ratio of the dangerous event chains in historical data. The nodes of the causal chain are candidate events that have passed the verification. The node attributes include event number, spatio-temporal compactness score, and causal matching status. The edges are trigger relationship labels connecting the nodes. The trigger relationship labels include time difference, spatial overlap degree, and causal logic condition. The time difference is the difference between the start time stamp of the subsequent event and the end time stamp of the previous event. The spatial overlap degree is the ratio of the intersection area. The causal logic condition is the rule ID corresponding to the causal matching status of the node event. The causal chain of dangerous events is stored as a directed graph structure. The directed graph is stored in the adjacency list format. Each row record of the adjacency list includes the start node number, end node number, time difference, spatial overlap degree, and rule ID. The storage path is compatible with the input interface of the subsequent risk threshold judgment step. The data format is a CSV file, the encoding format is UTF-8, and the field separator is a comma.

[0040] S6. Generate a dynamic dangerous environment recognition result according to the relationship between the causal chain of dangerous events and the preset risk threshold. The specific implementation is as follows: When analyzing the node attributes and edge attributes of the causal chain of dangerous events, the node attributes include the spatio-temporal tightness score and logical label of the candidate event. After normalizing the spatio-temporal tightness score to the range of 0 to 1, it is used as the weight base value. When the logical label is "passed", the weight base value is multiplied by 1.2, and when the logical label is "low confidence", it is multiplied by 0.8. For example, if the spatio-temporal tightness score of a certain node is 0.9 and the logical label is "passed", then the node risk weight is 0.9×1.2 = 1.08. The weight calculation rule is optimized according to the statistical data of historical dangerous events. The historical data is the actual risk contribution degree of dangerous event nodes in the same type of scenarios in the past five years. For example, in the traffic accident scenario, the risk contribution degree of nodes with the logical label "passed" has increased by an average of 20%. The edge attributes include the time difference of the trigger relationship, the spatial overlap degree, and the causal logic condition. The time difference unit is milliseconds, the spatial overlap degree unit is percentage, and the causal logic condition is the rule ID. The propagation path depth is calculated by traversing the directed graph structure of the causal chain to calculate the longest path length. The traversal algorithm is depth-first search, and the counting rule of the path length is the number of nodes. For example, the depth of the path A→B→C is 3. During the traversal process, if there is a loop, the loop nodes are automatically ignored to prevent infinite recursion.

[0041] When calculating the overall risk value of the causal chain based on the product of the propagation path depth and the node risk weight, the product calculation is the path depth multiplied by the arithmetic sum of the risk weights of all nodes on the path. For example, if the propagation path depth is 3 and the node weights are 1.08, 0.95, and 0.78 respectively, then the overall risk value is 3×(1.08 + 0.95 + 0.78) = 8.43. The calculation result is reserved to two decimal places. During the calculation process, if there are isolated nodes (both in-degree and out-degree are zero) in the path, they are automatically excluded. The determination of isolated nodes is realized by traversing the adjacency list of the directed graph. Each row record of the adjacency list contains the starting node, the ending node, and the edge attributes. The risk weights of isolated nodes are not included in the overall risk value, and the excluded isolated nodes are recorded as an independent log file. The log file format is CSV, and the fields include the node number, the exclusion reason, and the timestamp. The log storage path is compatible with the system monitoring module.

[0042] When comparing the overall risk value of the causal chain with the preset risk threshold, the preset risk threshold is dynamically set according to the scenario type and the danger level. The scenario types are divided into traffic accidents, industrial leaks, and others. The danger levels are divided into level 1 (high), level 2 (medium), and level 3 (low). The threshold setting rules are as follows: the level 1 threshold for the traffic accident scenario is 6.0, and the level 2 is 4.0; the level 1 threshold for the industrial leak scenario is 5.0, and the level 2 is 3.0; for other scenarios, the unified level 1 threshold is 4.0. The threshold setting is based on the hierarchical statistics of the actual losses of dangerous events in historical data. For example, in traffic accidents, the actual casualty rate of events with an overall risk value exceeding 6.0 is more than 70%; in industrial leaks, the spread range of events exceeding 5.0 is more than 100 square meters. When the overall risk value exceeds the threshold, a dynamic dangerous environment identification result is generated. The identification result includes the causal chain identifier, risk level, and trigger location information. The generation rule of the causal chain identifier is "scenario type_date_sequential number", for example, "traffic_20231001_001". The risk level is marked as "level 1" or "level 2" according to the exceeded threshold level. The trigger location information is the center point coordinates of the position area of the first trigger node in the causal chain. The center point coordinates are calculated through the center of the minimum bounding rectangle of the node position area. The center point coordinates of the minimum bounding rectangle are the upper left corner coordinates of the rectangle plus half of the width and height. The coordinate unit is pixels or longitude and latitude. The output format of the dynamic dangerous environment identification result is a JSON object, including fields: identifier, level, location, timestamp. The encoding format of the JSON object is UTF-8. The field order is fixed as identifier, level, location, timestamp. The output interface is compatible with the API of the risk warning system. The data interaction protocol is HTTP / 1.1. The transmission content encryption uses the TLS 1.2 protocol.

[0043] In the dynamic hazard environment recognition, this method solves the problem of the disconnection between static rules and dynamic scenarios in traditional methods through the deep collaboration of multi-modal data and the dynamic adaptation of causal logic. Existing technologies usually process visual or language data independently and rely on predefined fixed rules, making it difficult to capture the causal relationships in the spatio-temporal evolution of hazardous events. For example, in a traffic accident scenario, the causal chain between illegal lane changes and collision results needs to combine the spatio-temporal continuity of vehicle trajectories and the semantic constraints of traffic rules. However, due to the lack of a dynamic rule screening mechanism (such as the spatio-temporal correlation verification in S3) and causal intensity evaluation (such as the logical completeness verification in S4) in traditional methods, random parallel trajectories are often misjudged as causal relationships. This method extracts dynamic causal rules (S2) from language descriptions and combines them with the evolution law of the scenario change sequence (such as the leakage diffusion direction) to achieve real-time adaptation of rules to the environment. Further, through the overall risk quantification of the causal chain (S6), a dynamic product mechanism of path depth and node weight is introduced, breaking through the limitation of single-node risk assessment. For example, the joint optimization of dynamic rule weight assignment and causal chain path depth is a combination verified through a large number of tests in a specific scenario. In addition, the strict logical progression between steps (such as the data flow closed loop from S1→S3→S6) ensures the feasibility of the technical solution in dynamic scenarios such as industrial monitoring and autonomous driving, forming an organically unified hazard recognition framework.

[0044] In the embodiments, all the calculations involved are dimensionless numerical calculations, and the preset parameters and threshold selections in the calculations are set by those skilled in the art according to the actual situation.

[0045] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0046] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and invention constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0047] In addition, the functional modules in each embodiment of this application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0048] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.

[0049] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0050] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for identifying dangerous environments based on a vision-language model and dynamic scenes, characterized in that, The steps are as follows: S1. Obtain the visual data stream of the dynamic scene and the corresponding language description data; S2. Extract the spatio-temporal features of dynamic objects from the visual data stream to generate the target motion trajectory and the scene change sequence; extract the causal rule text associated with the dangerous event from the language description data to construct the causal logic constraint set; S3. Perform spatio-temporal association verification on the target motion trajectory and the causal logic constraint set, and filter out the candidate set of dynamic dangerous events that meet the causal constraints; S4. Evaluate the causal strength of the candidate set of dynamic dangerous events, calculate the spatio-temporal compactness and verify the logical completeness to generate a candidate set with high confidence; S5. Perform logical reasoning on the candidate set with high confidence based on the scene change sequence to generate the causal chain of dangerous events; S6. Generate the recognition result of the dynamic dangerous environment according to the relationship between the causal chain of dangerous events and the preset risk threshold.

2. The method for identifying a dangerous environment based on a vision-language model and a dynamic scene according to claim 1, wherein Obtain the visual data stream of the dynamic scene and the corresponding language description data, including: Perform multi-source synchronous acquisition on the dynamic scene to generate the visual data stream and language description data containing timestamps; Perform resolution adaptive adjustment and abnormal frame filtering on the visual data stream to generate a standardized visual data stream; Perform noise text cleaning and key semantic extraction on the language description data to generate structured language description data; Align the standardized visual data stream and the structured language description data according to the timestamp and store them as an associated data group.

3. A method for identifying dangerous environments based on a vision-language model and dynamic scenes according to claim 1, characterized in that, Extract the spatio-temporal features of dynamic objects from the visual data stream to generate the target motion trajectory and the scene change sequence, including: Perform dynamic object detection and tracking on the standardized visual data stream to generate the bounding box coordinates of the dynamic object and the displacement vector between consecutive frames; Calculate the motion speed and direction of the dynamic object based on the displacement vector between consecutive frames, and generate the target motion trajectory through the trajectory smoothing algorithm; Extract the illumination intensity change gradient, background area motion vector and dynamic object aggregation density of adjacent frames in the standardized visual data stream, and fuse them to generate the scene change sequence.

4. A method for identifying dangerous environments based on a vision-language model and dynamic scenarios according to claim 1, characterized in that, Extract the causal rule text associated with the dangerous event from the language description data to construct the causal logic constraint set, including: Perform semantic role annotation on the structured language description data to identify the action subject, object of action and constraint conditions of the dangerous event; Convert the semantic role annotation result into a spatio-temporal constraint rule according to the preset causal logic template to construct the causal logic constraint set.

5. A method for identifying dangerous environments based on a vision-language model and dynamic scenarios according to claim 1, characterized in that, Perform spatio-temporal association verification on the target motion trajectory and the causal logic constraint set, and filter out the candidate set of dynamic dangerous events that meet the causal constraints, including: Calculate the regional overlap degree according to the spatial coverage range of the target motion trajectory and the spatial constraint conditions in the causal logic constraint set, and filter out the candidate trajectories with the spatial coverage exceeding the preset overlap threshold; Verify the temporal matching degree based on the time window of the target motion trajectory and the time constraint conditions in the causal logic constraint set, and retain the candidate trajectories with the time overlap ratio exceeding the preset temporal threshold; For the candidate trajectories that pass the spatio-temporal coverage verification, perform semantic consistency matching in combination with the logical conditions in the causal logic constraint set, and eliminate the trajectories with logical conflicts to generate the candidate set of dynamic dangerous events.

6. The method for identifying a dangerous environment based on a vision language model and a dynamic scene according to claim 1, wherein Perform causal strength assessment on the candidate set of dynamic hazard events, calculate spatio-temporal compactness, and verify logical completeness to generate a high-confidence candidate set, including: Calculate the spatio-temporal compactness of each trajectory in the candidate set of dynamic hazard events, including statistically analyzing the spatial distribution concentration and time series continuity of trajectory parameters to generate a spatio-temporal compactness score; Verify the logical completeness of the candidate set of dynamic hazard events, including checking whether the causal logic conditions of candidate trajectories are consistent with the evolution law of the scenario change sequence and removing trajectories with logical contradictions; Screen candidate trajectories with spatio-temporal compactness scores exceeding the preset spatio-temporal compactness threshold and passing the logical completeness verification result to generate a high-confidence candidate set.

7. The method for identifying a dangerous environment based on a vision-language model and a dynamic scene according to claim 1, wherein Perform logical reasoning on the high-confidence candidate set based on the scenario change sequence to generate a causal chain of hazard events, including: Extract the evolution pattern features of the scenario change sequence, including analyzing the parameter trends and mutation points in the scenario change encoding to generate a set of spatio-temporal evolution patterns; Align the candidate events in the high-confidence candidate set with the set of spatio-temporal evolution patterns in terms of time series and perform causal association matching to identify the causal triggering relationships between candidate events; Chain together candidate events according to the causal triggering relationships to generate a causal chain of hazard events.

8. A method for identifying a dangerous environment based on a vision-language model and a dynamic scene according to claim 7, characterized in that, The nodes of the causal chain of hazard events are candidate events, and the edges are the spatio-temporal constraint conditions of the triggering relationships.

9. The method for identifying a dangerous environment based on a vision-language model and a dynamic scene according to claim 1, wherein Generate the identification result of the dynamic hazardous environment according to the relationship between the causal chain of hazard events and the preset risk threshold, including: Analyze the node attributes and edge attributes of the causal chain of hazard events, and extract the propagation path depth and node risk weight of the causal chain of hazard events; Calculate the overall risk value of the causal chain according to the product of the propagation path depth and the node risk weight; Compare the overall risk value of the causal chain with the preset risk threshold, and generate the identification result of the dynamic hazardous environment when the overall risk value of the causal chain exceeds the preset risk threshold.

10. The method for identifying a dangerous environment based on a vision-language model and a dynamic scene according to claim 9, wherein, The identification result of the dynamic hazardous environment includes the causal chain identifier, risk level, and trigger location information.

Citation Information

Cited By

  • Urban public security risk chain identification method and system based on graph neural network

    CN120705509A

  • Production process visualized agricultural product quality tracing generation method

    CN120746408A

  • Multi-mode-based hidden danger identifying, monitoring and early warning method and system

    CN120804618A

  • Multi-modal based hidden danger identification monitoring and early warning method and system

    CN120804618B

  • Road traffic accident restoration system and method suitable for intelligent network connection automobile

    CN120853392A