Method and device for checking safety hazards in production site

By encoding multimodal data from the production site into the same semantic space and using production site knowledge graphs and multimodal large language models for collaborative reasoning, the problems of low efficiency and poor accuracy in safety hazard investigation in existing technologies are solved, achieving efficient and accurate safety hazard detection and dynamic risk assessment.

CN121119173BActive Publication Date: 2026-02-27PEKING UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511650255.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-27
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Existing technologies have poor multimodal data analysis performance in the investigation of safety hazards on the production site, resulting in low investigation efficiency and insufficient safety. Traditional methods cannot effectively utilize multiple visual modal information, and the annotation cost is high and the accuracy is poor.

Method used

Multimodal data from the production site is encoded into the same semantic space. New graph data is constructed using production site knowledge graphs and entity information. Collaborative reasoning is performed using a multimodal large language model to obtain safety hazard investigation results. These results are then combined with a decision-making model for integrated decision-making.

Benefits of technology

It improves the efficiency and safety of identifying safety hazards on the production site, can dynamically adjust risk assessments, provide accurate hazard detection and targeted safety measures, and reduces model training costs and annotation difficulties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121119173B_ABST
    Figure CN121119173B_ABST
Patent Text Reader

Abstract

The application relates to the field of production site safety management, and provides a production site safety hidden danger checking method and device, which comprises the following steps: encoding multi-modal data of a production site into the same semantic space to obtain coded data; obtaining entity information from the coded data, constructing new graph data based on a production site knowledge graph and the entity information, and performing natural language processing on the new graph data to obtain prompt words; performing collaborative reasoning on potential safety hidden dangers of the production site based on a multi-modal large language model according to the coded data and the prompt words to obtain a safety hidden danger checking result; and the multi-modal large language model is used for predicting the correlation between different production data in the production site and various safety hidden dangers. The method significantly improves the reasoning capability of the multi-modal large language model, can more efficiently and accurately check safety hidden dangers in the production site, and thus improves the safety hidden danger checking efficiency and safety of the production site.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of production site safety management, and in particular to a production site safety hazard checking method and device. BACKGROUND

[0002] The related art is to perform multi-modal large language model inference question and answer according to multi-modal data such as text and pictures collected in a production site (such as construction, maintenance and commissioning scene operation) and natural language input of a user, and the multi-modal data is usually saved in an unstructured manner. The traditional natural language analysis model is limited by the quality of the text, and the analysis effect of the text data of the production site is not good. The prior art also trains a visual model by collecting, screening and labeling pictures of various production sites, and uses a sensing camera to monitor the production site picture in real time to detect and locate various safety hazards such as illegal operation and mechanical abnormalities. Due to the complex scene of the production site pictures and the large number of safety hazard types, the labeling cost is high, and the specific safety hazard category and position information cannot be accurately identified. In addition, the traditional large language model focuses on natural language understanding and question answering, which are based on text input. However, the actual scene will not be limited to a single text mode. The support capability for various visual modalities of the production site is insufficient, thereby reducing the accuracy and efficiency of safety hazard checking. SUMMARY

[0003] The present application provides a production site safety hazard checking method and device to solve the problem of poor inference question and answer effect of the prior art through multi-modal data of the production site, high difficulty, low efficiency of checking safety hazards, and low construction safety, thereby improving the safety hazard checking efficiency and safety of the production site.

[0004] The present application provides a production site safety hazard checking method, comprising:

[0005] Encoding the multi-modal data of the production site into the same semantic space to obtain encoded data, the multi-modal data comprising at least two of text, image, sensing data, video and context information, the context information comprising at least one of weather information, geographical information and production stage information of the production site;

[0006] Obtaining entity information from the encoded data, constructing new graph data based on a production site knowledge graph and the entity information, and performing natural language processing on the new graph data to obtain prompt words; wherein the new graph data comprises one of a knowledge subgraph and a plurality of knowledge paths;

[0007] The multi-modal large language model is used for collaborative reasoning of potential safety hazards of the production site according to the coded data and the prompt word, to obtain a safety hazard checking result; wherein the multi-modal large language model is used for predicting the association between different production data and different safety hazards in the production site.

[0008] According to the production site safety hazard checking method provided by the application, the multi-modal data includes text, image, video and context information;

[0009] The coded data obtained by coding the multi-modal data of the production site into the same semantic space includes:

[0010] The text is coded into text features, at least one of the image and the video is coded into visual features, and the context information is coded into context features;

[0011] The text features, the visual features and the context features are fused to obtain the coded data.

[0012] According to the production site safety hazard checking method provided by the application, the multi-modal data includes video, and the video includes a plurality of video frame data;

[0013] The coded data obtained by coding the multi-modal data of the production site into the same semantic space further includes:

[0014] The three-dimensional convolutional neural network is used to extract the time sequence features and the spatial features from the plurality of video frame data, and the light field technology is used to extract the light flow features from the plurality of video frame data;

[0015] The time sequence features, the spatial features and the light flow features are fused to obtain the coded data.

[0016] According to the production site safety hazard checking method provided by the application, the entity information obtained from the coded data includes:

[0017] The knowledge field corresponding to the coded data is analyzed to obtain a keyword entity;

[0018] The keyword entity and the entity set in the external knowledge graph are similarity calculated to obtain the entity information.

[0019] According to the production site safety hazard checking method provided by the application, the new graph data constructed based on the production site knowledge graph and the entity information includes:

[0020] In the case that the number of entities in the production site knowledge graph does not exceed the entity threshold, the knowledge sub-graph is constructed according to the production site knowledge graph and the entity information.

[0021] According to the method for investigating safety hazards at production sites provided by the present invention, the step of constructing new graph data based on the production site knowledge graph and the entity information further includes:

[0022] If the number of entities in the production site knowledge graph exceeds the entity threshold, the multiple knowledge paths are constructed based on the production site knowledge graph and the entity information.

[0023] According to the method for identifying safety hazards at production sites provided by the present invention, the production site knowledge graph is obtained through a dynamically updated knowledge graph database.

[0024] According to a method for identifying safety hazards at a production site provided by the present invention, after obtaining the results of the safety hazard identification, the method further includes:

[0025] The results of the safety hazard investigation are integrated based on the decision model to obtain the target decision; wherein the decision model is obtained by iterative training with sample safety hazard data as input and one or more preset sample decisions as labels.

[0026] The present invention also provides a device for detecting safety hazards at production sites, comprising:

[0027] An alignment encoding module is used to encode multimodal data from the production site into the same semantic space to obtain encoded data. The multimodal data includes at least two of the following: text, images, sensor data, video, and contextual information. The contextual information includes at least one of the following: weather information, geographic information, and production stage information from the production site.

[0028] The knowledge graph auxiliary module is used to obtain entity information from the encoded data, construct new graph data based on the production site knowledge graph and the entity information, and perform natural language processing on the new graph data to obtain prompt words; wherein, the new graph data includes a knowledge subgraph and one of multiple knowledge paths;

[0029] The reasoning module is used to perform collaborative reasoning on potential safety hazards in the production site based on the encoded data and the prompt words using a multimodal large language model, and to obtain the safety hazard investigation results; wherein, the multimodal large language model is used to predict the correlation between different production data and different safety hazards in the production site.

[0030] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the production site safety hazard investigation method described above.

[0031] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the production site safety hidden danger checking method.

[0032] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the production site safety hidden danger checking method.

[0033] The application provides a production site safety hidden danger checking method and device, which encodes multi-modal data of a production site into the same semantic space to obtain encoded data, acquires entity information from the encoded data, constructs new graph data based on a production site knowledge graph and the entity information, performs natural language processing on the new graph data to obtain prompt words, and finally performs collaborative reasoning on potential safety hidden dangers of the production site based on a multi-modal large language model according to the encoded data and the prompt words to obtain a safety hidden danger checking result, thereby enhancing the reasoning capability of the multi-modal large language model and improving the safety hidden danger checking efficiency and safety of the production site. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0035] Figure 1 is a flowchart of the production site safety hidden danger checking method provided by the application.

[0036] Figure 2 is a structural schematic diagram of the production site safety hidden danger checking device provided by the application.

[0037] Figure 3 is a structural schematic diagram of the electronic device provided by the application. DETAILED DESCRIPTION

[0038] In order to make the objects, technical solutions and advantages of the application clearer, the technical solutions in the application will be described clearly and completely below with reference to the drawings in the application. Obviously, the described embodiments are some embodiments of the application, but not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0039] The production site safety hidden danger checking method and device of the application will be described below. Figures 1-2 The production site safety hidden danger checking method and device of the application will be described below.

[0040] Figure 1 is a flowchart of a production site safety hazard checking method provided by the present application, as shown in the figure, the production site safety hazard checking method comprises the following steps: Figure 1

[0041] Step 110, encode the multi-modal data of the production site into the same semantic space to obtain encoded data, the multi-modal data includes at least two of text, image, sensor data, video and context information, and the context information includes at least one of weather information, geographic information and production stage information of the production site.

[0042] In this step, the production site includes but is not limited to construction, maintenance, debugging, management, logistics and emergency rescue scenes.

[0043] In this step, the multi-modal data includes data of different formats obtained from different scenes in the construction site, for example, it can be image or video data taken, it can also be recorded voice, text data, and it can also be sensor data obtained by related sensors.

[0044] In this step, since the information source of the production site is extensive, for example, the multi-modal data in this scene includes text data such as safety reports and work logs, visual information such as construction pictures captured by sensing cameras and mechanical operation; It also includes natural language information such as user (safety officer) demand description; By using different types of encoders to encode the above multi-modal data respectively, and then aligning the semantic spaces of the multi-modal data of different encodings, the input samples for the subsequent reasoning process can be obtained.

[0045] For example, the multi-modal data includes text data, natural language input data and image data, the text data and the natural language input data are text encoded by a text encoder, and then the picture information is encoded into the same semantic space by using a visual encoder, which is consistent with the text embedding dimension, and the image and text information is aligned, to obtain the encoded data, which can be input to the subsequent multi-modal large language model for reasoning; By fusing multi-modal information such as image and text, the multi-modal large language model can capture the details of the production site more finely and enhance the safety hazard checking capability.

[0046] In this embodiment, before data encoding, the multi-modal data of the production site can be preprocessed such as data cleaning and normalization to reduce the interference of abnormal data and improve the quality of multi-modal data.

[0047] ​Step 120, obtaining entity information from the coded data, constructing new graph data based on the production site knowledge graph and the entity information, and performing natural language processing on the new graph data to obtain prompt words; wherein the new graph data includes one of a knowledge subgraph and a plurality of knowledge paths.

[0048] It should be noted that the pre-training of the multi-modal large language model is often to meet the reasoning in general fields, but in the application in specific fields such as construction fields, the pre-training process will face problems such as knowledge shortage and insufficient learning of knowledge in the field, and the cost of fine-tuning training using the knowledge graph in the construction field is relatively high.

[0049] In this step, the production site knowledge graph can be constructed by the association between a plurality of entities in the construction scene, or can be directly obtained from the knowledge graph database.

[0050] In this embodiment, the key entities are extracted from the coded data, and the semantic association information and implicit contact information associated with the key entities are mined from the external knowledge graph, i.e. new graph data. These graph data can be complete knowledge subgraphs or multiple knowledge paths. Finally, the obtained new graph data is converted into natural language and used as a prompt word for subsequent auxiliary multi-modal large language model reasoning.

[0051] In this embodiment, by inputting external knowledge as a prompt word, the model fine-tuning cost is reduced, and the robustness and transferability of the model are enhanced. The prompt word is used to guide the multi-modal large language model to perform accurate reasoning, so that the model can unify the semantics in the knowledge graph, eliminate ambiguity, and thus improve the accuracy of the production site question and answer.

[0052] Step 130, based on the multi-modal large language model, the potential safety hazards in the production site are cooperatively reasoned according to the coded data and the prompt words, and a safety hazard investigation result is obtained; wherein the multi-modal large language model is used to predict the association between different production data and different safety hazards in the production site.

[0053] In this step, the multi-modal large language model fuses the natural language information in the coded data and the picture visual information, and understands the semantic relationship between the image and the text in the common vector space. Through the "picture-to-text" question and answer mode (i.e. explaining according to the input image and question), it answers whether there is a safety hazard in the construction field, and if there is, what solution should be taken.

[0054] In this embodiment, the multi-modal large language model takes the encoded data corresponding content as the reasoning object, takes the prompt word as the reasoning guide, and then inputs the question and the converted knowledge subgraph according to the pre-defined prompt word template. The model understands and enhances the input to build its own mind map for reasoning, and obtains the safety hazard investigation result, which can cope with complex and variable construction scenes.

[0055] The production site safety hazard investigation method provided by the embodiment of the application encodes the multi-modal data of the production site into the same semantic space to obtain encoded data, obtains entity information from the encoded data, constructs new graph data based on the production site knowledge graph and the entity information, performs natural language processing on the new graph data to obtain a prompt word, and finally performs collaborative reasoning on potential safety hazards of the production site based on the multi-modal large language model according to the encoded data and the prompt word to obtain a safety hazard investigation result. The reasoning ability of the multi-modal large language model is enhanced, and the safety hazard investigation efficiency and safety of the production site are improved.

[0056] In some embodiments, the multi-modal data includes text, images, videos, and context information; encoding the multi-modal data of the production site into the same semantic space to obtain encoded data includes: encoding the text into text features, encoding at least one of the images and the videos into visual features, and encoding the context information into context features; and performing feature fusion on the text features, the visual features, and the context features to obtain the encoded data.

[0057] It should be noted that the existing production site safety hazard investigation method is mainly based on general scenarios and lacks in-depth analysis of context information. For example, in rainy weather, the risk of a slippery area should be increased, but the traditional method does not consider the context information such as weather conditions, geographical location, and production stage, which leads to insufficient accuracy and effectiveness of hazard detection, and cannot timely adjust the risk rating to provide effective safety warning.

[0058] Specifically, the existing production site safety hazard investigation cannot dynamically adjust risk assessment, and adverse weather such as rain and snow increases construction risk; the existing production site safety hazard investigation does not consider geographical location, and the environmental differences (such as terrain and surrounding facilities) of different construction sites are not considered, resulting in poor generalization ability of hazard detection; the existing method cannot adjust the detection focus according to the construction progress, and different stages of the construction project may involve different risk points.

[0059] In this embodiment, the multi-modal fusion encoded data is obtained through the following steps:

[0060] (1) Collecting text reports, images, videos, and contextual information (geographical location, weather, time, production stage, etc.) from the production site, and frame sequence labeling the videos to ensure synchronization of multi-modal data in time and space.

[0061] (2) Encoding the contextual information into numerical features, encoding the image and video into visual features, and encoding the text data into text features.

[0062] For example, weather conditions can be represented by temperature ( ), humidity ( ), rainfall probability ( ), etc.

[0063] ;

[0064] wherein, Concat is a feature concatenation operation, is the contextual feature, is the text feature, is the visual feature.

[0065] (3) In the input layer of the multi-modal large language model, the contextual feature is fused with the visual and text features to form a comprehensive feature vector .

[0066] In this embodiment, using the multi-modal large language model to perform safety hazard reasoning based on the encoded data and prompt words can improve the risk rating of a specific area according to real-time weather conditions, such as increasing the risk of slippery areas in rainy weather, can use geographical location information (such as local terrain and environmental factors) to provide more targeted hazard detection, and can adjust the risk points of concern according to the construction progress, for example, in the high-altitude operation stage, the risk of high-altitude falling is monitored.

[0067] The production site safety hazard investigation method provided by the embodiment of the application can dynamically adjust risk assessment and provide more accurate hazard detection by fusing text features, visual features, and contextual features for multi-modal large language model safety hazard reasoning with prompt words, and can also provide targeted safety measures and suggestions according to specific construction environment and conditions, improving the effectiveness of safety management.

[0068] In some embodiments, the multi-modal data includes a video, and the video includes a plurality of video frame data; encoding the multi-modal data of the production site into the same semantic space to obtain the encoded data further includes: extracting time sequence features and spatial features from the plurality of video frame data using a three-dimensional convolutional neural network, and extracting optical flow features from the plurality of video frame data using a light field technology; fusing the time sequence features, the spatial features, and the optical flow features to obtain the encoded data.

[0069] It should be noted that the existing method usually down-samples the video into individual images for frame-by-frame analysis when processing the video data; this way ignores the time continuity of the video and cannot capture the security risks brought by dynamic changes; for example: the moving process of large mechanical equipment may have potential risks, but single-frame image cannot fully reflect it; the building structure may change during construction, and frame-by-frame analysis cannot identify the accumulated deformation or offset.

[0070] In this embodiment, a three-dimensional convolutional neural network (C3D) is used to process video data to capture time and spatial features; in the C3D model, the three-dimensional convolution operation is defined as:

[0071] ;

[0072] wherein, is the output feature value of the i-th layer at position (x, y, t); is the input feature value of the i-th layer at position (x, y, t); i,j,k is the weight of the convolution kernel of the j-th channel, with a size of (d x, d y, d t); is the bias of the i-th layer; M is the number of input channels; are the sizes of the convolution kernel in the time and spatial dimensions, respectively. m i+p,j+q,k+r In this embodiment, the three-dimensional convolution model can capture both spatial features (such as object shape, texture) and temporal features (such as motion trajectory, speed change) of the video. m In this embodiment, the optical flow between video frames is calculated by optical flow technology to obtain the direction and speed information of object motion, and then the optical flow features are fused with the features extracted by C3D to form a more comprehensive motion representation, i.e., the above-mentioned encoding data; the specific splicing process is represented by the following formula: P,Q,R ; P,Q,R wherein,

[0073] is a feature fusion operation (such as weighted sum, splicing, etc.), is the optical flow feature,

[0074] is the spatial and temporal features captured by the three-dimensional convolution model; is the encoding data.

[0075]

[0076]

[0077] ​​​​​​​​In this embodiment, the model simultaneously performs target detection of static images and time series analysis of videos, improving detection accuracy.

[0078] The corresponding model loss function is represented by the following formula:

[0079] ;

[0080] wherein, is a target detection loss, is a classification loss, is a regression loss; 、 and are loss weight coefficients corresponding to the three different losses respectively; L is a model loss function.

[0081] In this embodiment, the detected hidden dangers are risk rated in combination with context information. The risk score R can be represented as:

[0082] ;

[0083] wherein, is a hidden danger feature, is a risk assessment function (such as a neural network); R is a risk score; is a context feature.

[0084] The production site safety hidden danger investigation method provided by the embodiment of the application extracts time series features and spatial features from multiple video frame data by using a three-dimensional convolutional neural network, extracts optical flow features from multiple video frame data by using light field technology, and fuses the time series features, spatial features and optical flow features. The model can identify dynamic hidden dangers that cannot be captured by static images, such as risks during equipment movement, progressive deformation of structures, unsafe behaviors of personnel, etc., improving the comprehensiveness and depth of hidden danger detection.

[0085] In some embodiments, obtaining the entity information from the coded data includes: analyzing a knowledge field corresponding to the coded data to obtain a keyword entity; and performing similarity calculation on the keyword entity and an entity set in an external knowledge graph to obtain the entity information.

[0086] In this embodiment, a large language model (LLM) can be used to analyze the field to which the question belongs and extract a keyword set in the input question Q and a triple set T; and perform semantic similarity calculation on M and an entity set in an external knowledge graph : by comparing each keyword and each entity Perform embedding encoding (embedding) to obtain the vector representation and ; this embodiment uses cosine similarity to calculate the similarity between keywords and knowledge graph entities, and the specific calculation formula is as follows:

[0087] ;

[0088] wherein: is the embedding representation vector of the keyword ; is the embedding representation vector of the entity ; and are the Euclidean norms of vectors and , respectively; is the inner product of two vectors.

[0089] In this embodiment, by selecting the entity with the highest cosine similarity, the entity set is formed, and the calculation method is as follows:

[0090] ;

[0091] wherein, is the set of the highest similarity entities matched for each input keyword , and the knowledge subgraph in the subsequent step is generated based on the set.

[0092] In this embodiment, when the new graph data is a knowledge subgraph; constructing the new graph data based on the production site knowledge graph and the entity information includes: in a case where the number of entities in the production site knowledge graph does not exceed an entity threshold, constructing a knowledge subgraph according to the production site knowledge graph and the entity information.

[0093] In this embodiment, the entity threshold can be set according to user demand, for example, the entity threshold can be 50.

[0094] In this embodiment, after the required entity set (the above entity information) is screened out, the knowledge subgraph can be constructed by the path connection method and the neighbor expansion method, respectively.

[0095] Specifically, the following two methods are used to construct the knowledge subgraph from the source knowledge graph :

[0096] (1) Path connection method: this method connects any two nodes and in the entity set to form a path, and then connects all entity nodes and generates a subgraph , which contains all the paths between entities. The specific calculation method is as follows:

[0097] ;

[0098] , wherein: is the filtered entity set; represents the path between entity and entity ; and U represents merging all paths into a complete subgraph .

[0099] (2) Neighbor expansion method: for each entity in the entity set , find its neighbor node , and add the corresponding relationship to the subgraph. Through this method, the subgraph not only contains the nodes in the initial entity set, but also introduces new neighbor nodes through 1-hop expansion, and updates the entity set . The specific calculation method is as follows:

[0100] ;

[0101] , wherein: represents the neighbor node set of entity ; is the relationship between entity and neighbor node (i.e. the edge in the graph); represents the subgraph generated by the neighbor expansion method, which contains the triples between node and its neighbor nodes.

[0102] In this embodiment, the connection between each entity can be obtained from the above knowledge subgraph, and then each subgraph is formatted into an entity chain, and each entity chain is converted into a natural language description through a multi-modal large language model. According to a pre-defined prompt word template, input the question and the converted knowledge subgraph respectively, the model understands and enhances the input to build its own mind map for reasoning.

[0103] The production site safety hidden danger investigation method provided by the embodiment of the application efficiently extracts the knowledge subgraph related to the entity through the graph mining algorithm, enriches the semantic association and extracts the implicit knowledge, improves the reasoning ability and knowledge coverage range of the model, and the system can realize accurate reasoning by using the existing external knowledge base without relying on a large amount of data retraining; reduces the cost of model training and adjustment, and improves the knowledge reuse rate.

[0104] In some embodiments, when the new graph data is multiple knowledge paths, the constructing of the new graph data based on the production site knowledge graph and the entity information further comprises: in a case where the number of entities in the production site knowledge graph exceeds an entity threshold, constructing multiple knowledge paths according to the production site knowledge graph and the entity information.

[0105] In this embodiment, when the number of entities in the knowledge graph is large, the computational overhead of subgraph mining is large; in this embodiment, a plurality of key paths are generated by sampling instead of complete subgraphs, so as to obtain sufficient knowledge graph information at a lower computational cost.

[0106] For example, when the number of entities of the production site knowledge graph exceeds 50, the above entity information extracts 10 key knowledge paths from the knowledge graph, converts each knowledge path into a natural language description through a multi-modal large language model, and inputs the question and the converted knowledge subgraph according to a pre-defined prompt word template, so that the model understands and enhances the input to build its own mind map for reasoning.

[0107] The production site safety hidden danger investigation method provided by the embodiment of the application can effectively improve the reasoning efficiency under the premise of ensuring the reasoning quality by constructing multiple knowledge paths according to the production site knowledge graph and the entity information in a case where the number of entities in the production site knowledge graph exceeds an entity threshold, and is suitable for scenes with high time efficiency requirements.

[0108] In some embodiments, the production site knowledge graph is obtained through a dynamically updated knowledge graph database.

[0109] In this embodiment, in order to cope with rapidly changing fields (such as laws, regulations, etc.), a dynamic knowledge updating mechanism can be introduced, so that the production site knowledge graph and other knowledge graphs can be updated periodically or in real time.

[0110] The production site safety hidden danger investigation method provided by the embodiment of the application can ensure that the large model uses the latest domain knowledge when reasoning, maintain the accuracy and timeliness of reasoning, and reduce the workload of manual graph updating.

[0111] In some embodiments, in the case of limited resources or high real-time requirements, a lightweight large language model can be used, and the computational cost can be reduced by reducing the parameter size or using quantization technology.

[0112] In this embodiment, for specific tasks in different fields, a small field-specific model can also be introduced in combination with a knowledge graph to further optimize the reasoning efficiency.

[0113] In some embodiments, after obtaining the safety hazard investigation result, the method further comprises: integrating the safety hazard investigation result based on a decision model to obtain a target decision; wherein the decision model is obtained by iterative training based on sample safety hazard data as input and one or more sample decisions as labels.

[0114] In this embodiment, the decision model can be obtained through multiple pre-experiments, for example, by fitting the linear or nonlinear correlation between different safety hazard factors (density, equipment power consumption parameters, service life, climate, etc.) and historical emergency measures to determine the decision model.

[0115] In this embodiment, the safety hazard investigation result obtained by the above reasoning is input into the trained decision model, and the corresponding decision scheme can be directly output without the help of artificial analysis, further improving the safety hazard investigation efficiency.

[0116] In some embodiments, the detection result and risk assessment can also be converted into a human-readable report, i.e., a target decision, by using a multi-modal large language model; for example, generating the following safety suggestion: "detecting that the high-altitude worker does not wear a safety belt, the risk level is high, and it is recommended to correct immediately and strengthen safety training".

[0117] The present embodiment illustrates the following two scenarios:

[0118] Scenario one: in a production site on a rainy day, the model obtains weather information (such as rainfall probability ), and improves the risk rating of the wet ground area If it is detected that the worker does not wear anti-slip shoes, a high-risk alarm is generated, and anti-slip measures are recommended.

[0119] Scenario two: by analyzing the operation video of the tower crane equipment, the model captures the abnormal shaking (motion feature anomaly) of the equipment during movement by using the C3D network. Combined with the equipment operation state, production stage information, and the knowledge of the large model and the knowledge graph, the model evaluates the risk level as high and recommends immediate shutdown for maintenance.

[0120] The production site safety hazard investigation method provided by the embodiment of the present application integrates the safety hazard investigation result by a decision model to obtain a target decision, which can provide an intelligent solution based on multi-dimensional information, help decision makers make more intelligent and safer decisions, and improve the efficiency and safety of overall construction management.

[0121] The production site safety hazard investigation device provided by the present application is described below, and the production site safety hazard investigation device described below can be correspondingly referred to the production site safety hazard investigation method described above.

[0122] Figure 2 This is a schematic diagram of the production site safety hazard investigation device provided by the present invention, as shown below. Figure 2 As shown, the safety hazard investigation device for the production site includes: an alignment coding module 210, a knowledge graph auxiliary module 220, and a reasoning module 230.

[0123] Alignment encoding module 210 is used to encode multimodal data from the production site into the same semantic space to obtain encoded data. The multimodal data includes at least two of the following: text, images, sensor data, video, and contextual information. The contextual information includes at least one of the following: weather information, geographic information, and production stage information from the production site.

[0124] The knowledge graph auxiliary module 220 is used to obtain entity information from the encoded data, construct new graph data based on the production site knowledge graph and entity information, and perform natural language processing on the new graph data to obtain prompt words; wherein, the new graph data includes a knowledge subgraph and one of multiple knowledge paths;

[0125] The reasoning module 230 is used to perform collaborative reasoning on potential safety hazards in the production site based on the encoded data and prompt words using a multimodal large language model, and to obtain the safety hazard investigation results; wherein, the multimodal large language model is used to predict the correlation between different production data and different safety hazards in the production site.

[0126] The production site safety hazard investigation device provided in this invention encodes multimodal data from the production site into the same semantic space to obtain encoded data. Entity information is then extracted from the encoded data. New graph data is constructed based on the production site knowledge graph and entity information. Natural language processing is then performed on the new graph data to obtain prompt words. Finally, a multimodal large language model is used to perform collaborative reasoning on potential safety hazards in the production site based on the encoded data and prompt words to obtain safety hazard investigation results. This enhances the model's reasoning ability and thus improves the efficiency and safety of safety hazard investigation in the production site.

[0127] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 3As shown, the electronic device can include a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can invoke the logical instructions in the memory 330 to execute the production site safety hazard investigation method, which includes: encoding multi-modal data of a production site into the same semantic space to obtain encoded data, the multi-modal data including at least two of text, image, sensor data, video, and context information, and the context information including at least one of weather information, geographic information, and production stage information of the production site; obtaining entity information from the encoded data, constructing new graph data based on a production site knowledge graph and the entity information, and performing natural language processing on the new graph data to obtain a prompt word; wherein the new graph data includes one of a knowledge subgraph and multiple knowledge paths; based on a multi-modal large language model, performing collaborative reasoning on potential safety hazards of the production site according to the encoded data and the prompt word to obtain a safety hazard investigation result; wherein the multi-modal large language model is used to predict the association between different production data and different safety hazards in the production site.

[0128] In addition, the logical instructions in the memory 330 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part of the prior art that essentially contributes or the part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0129] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program being stored on a non-transitory computer readable storage medium, and the computer program being executable by a processor to enable a computer to perform the production site safety hazard checking method provided by the above method, the method comprising: encoding multi-modal data of a production site into the same semantic space to obtain encoded data, the multi-modal data comprising at least two of text, image, sensor data, video and context information, the context information comprising at least one of weather information, geographical information and production stage information of the production site; obtaining entity information from the encoded data, constructing new graph data based on a production site knowledge graph and the entity information, and performing natural language processing on the new graph data to obtain a prompt word; wherein the new graph data comprises one of a knowledge subgraph and a plurality of knowledge paths; and performing collaborative reasoning on potential safety hazards of the production site based on a multi-modal large language model according to the encoded data and the prompt word to obtain a safety hazard checking result; wherein the multi-modal large language model is used to predict the association between different production data and different safety hazards in the production site.

[0130] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the production site safety hazard checking method provided by the above method, the method comprising: encoding multi-modal data of a production site into the same semantic space to obtain encoded data, the multi-modal data comprising at least two of text, image, sensor data, video and context information, the context information comprising at least one of weather information, geographical information and production stage information of the production site; obtaining entity information from the encoded data, constructing new graph data based on a production site knowledge graph and the entity information, and performing natural language processing on the new graph data to obtain a prompt word; wherein the new graph data comprises one of a knowledge subgraph and a plurality of knowledge paths; and performing collaborative reasoning on potential safety hazards of the production site based on a multi-modal large language model according to the encoded data and the prompt word to obtain a safety hazard checking result; wherein the multi-modal large language model is used to predict the association between different production data and different safety hazards in the production site.

[0131] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0132] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0133] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for identifying safety hazards in a production site, characterized in that, The method comprises the following steps: encoding multi-modal data of a production site into the same semantic space to obtain encoded data, wherein the multi-modal data comprises at least two of text, image, sensor data, video and context information, and the context information comprises at least one of weather information, geographical information and production stage information of the production site; obtaining entity information from the encoded data, constructing new graph data based on a production site knowledge graph and the entity information, and performing natural language processing on the new graph data to obtain prompt words; wherein the new graph data comprises one of a knowledge subgraph and multiple knowledge paths; based on a multi-modal large language model, performing collaborative reasoning on potential safety hazards of the production site according to the encoded data and the prompt words to obtain a safety hazard investigation result; wherein the multi-modal large language model is used to predict the association between different production data and different safety hazards in the production site; constructing new graph data based on the production site knowledge graph and the entity information comprises: in the case that the number of entities in the production site knowledge graph does not exceed an entity threshold, constructing the knowledge subgraph according to the production site knowledge graph and the entity information; in the case that the number of entities in the production site knowledge graph exceeds the entity threshold, constructing the multiple knowledge paths according to the production site knowledge graph and the entity information.

2. The production site safety hazard investigation method according to claim 1, characterized by, The multi-modal data comprises text, image, video and context information. The encoding of the multi-modal data of the production site into the same semantic space to obtain the encoded data comprises: encoding the text into text features, encoding at least one of the image and the video into visual features, and encoding the context information into context features; performing feature fusion on the text features, the visual features and the context features to obtain the encoded data.

3. The method according to claim 1, wherein The multi-modal data comprises video, and the video comprises multiple video frame data. The encoding of the multi-modal data of the production site into the same semantic space to obtain the encoded data further comprises: extracting time sequence features and spatial features from the multiple video frame data by using a three-dimensional convolutional neural network, and extracting optical flow features from the multiple video frame data by using a light field technology; performing fusion on the time sequence features, the spatial features and the optical flow features to obtain the encoded data.

4. The method according to claim 1, wherein The obtaining of the entity information from the encoded data comprises: analyzing a knowledge field corresponding to the encoded data to obtain a keyword entity; performing similarity calculation on the keyword entity and an entity set in an external knowledge graph to obtain the entity information.

5. The method of claim 1, wherein, The production site knowledge graph is obtained through a dynamically updated knowledge graph database.

6. The method of claim 1, wherein, After obtaining the safety hazard investigation result, the method further comprises: integrating the safety hazard investigation result based on a decision model to obtain a target decision; wherein the decision model is obtained by iterative training based on sample safety hazard data as input and one or more preset sample decisions as labels.

7. A production site safety hazard investigation device characterized by comprising: The method comprises the following steps: An alignment encoding module is configured to encode multi-modal data of a production site into a same semantic space to obtain encoded data, the multi-modal data including at least two of text, image, sensor data, video and context information, and the context information including at least one of weather information, geographical information and production stage information of the production site; A knowledge graph auxiliary module is configured to acquire entity information from the encoded data, construct new graph data based on a production site knowledge graph and the entity information, and perform natural language processing on the new graph data to obtain a prompt word; wherein the new graph data includes one of a knowledge subgraph and multiple knowledge paths; A reasoning module is configured to perform collaborative reasoning on potential safety hazards of the production site based on a multi-modal large language model according to the encoded data and the prompt word to obtain a safety hazard elimination result; wherein the multi-modal large language model is used to predict the association between different production data and different safety hazards in the production site. The knowledge graph auxiliary module is specifically configured to: construct the knowledge subgraph according to the production site knowledge graph and the entity information when the number of entities in the production site knowledge graph does not exceed an entity threshold value; and construct the multiple knowledge paths according to the production site knowledge graph and the entity information when the number of entities in the production site knowledge graph exceeds the entity threshold value.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the production site safety hazard elimination method according to any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the production site safety hazard elimination method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal reasoning method and device based on large language model and knowledge graph

    CN118193684A