Training method and detection method of potential safety hazard detection model and electronic equipment
Through multimodal data acquisition and association processing, the safety hazard detection model is trained, which solves the accuracy problem of a single visual detection method in complex environments, and achieves higher safety hazard detection accuracy.
Patent Information
- Application Number
- CN202411911199.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-13
Smart Images

Figure CN119989252A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of engineering safety management, and in particular to a training method, a detection method and an electronic device for a safety hazard detection model. Background Art
[0002] Engineering safety analysis is the process of identifying potential safety hazards or risk factors by monitoring and analyzing data from the construction environment. Currently, engineering safety analysis is usually performed using a detection method based on visual algorithms, that is, by collecting images of the construction environment and using deep learning algorithms to analyze the images to identify safety hazards.
[0003] In the process of implementing this application, the inventors found that there are at least the following problems in the prior art: the environment of the construction site is complex and changeable, and a single visual detection method is easily affected by factors such as shooting conditions, resulting in errors in the identification of safety hazards, thereby reducing the accuracy of safety hazard detection. Summary of the invention
[0004] The embodiments of the present application provide a training method, a detection method and an electronic device for a safety hazard detection model to solve the technical problem of low safety hazard detection accuracy caused by a single modality detection method.
[0005] The embodiments of the present application provide the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a method for training a potential safety hazard detection model, comprising:
[0007] Acquire first multimodal data of engineering environments;
[0008] Performing data association processing on the first multimodal data to construct an instruction data set;
[0009] Train a safety hazard detection model based on the instruction dataset.
[0010] In some embodiments, the first multimodal data includes first visual data, first audio data, first positional relationship data, and first device state data;
[0011] Get first-of-its-kind multimodal data of your engineering environment, including:
[0012] Acquire first visual data collected by a camera device in an engineering environment, wherein the first visual data includes at least one of individual protection information, working environment information, and worker behavior information;
[0013] Acquire first audio data collected by a sound collection device in an engineering environment, wherein the first audio data includes at least one of equipment operation sound, gas release sound, and liquid flow sound;
[0014] Acquire first position relationship data collected by a position collection device in an engineering environment, wherein the first position relationship data includes at least one of worker position information and construction equipment position information;
[0015] First equipment status data collected by an equipment status monitoring device in an engineering environment is acquired, wherein the first equipment status data includes operation status information of the construction equipment.
[0016] In some embodiments, performing data association processing on the first multimodal data to construct an instruction data set includes:
[0017] Performing a preprocessing operation on the first multimodal data to obtain second multimodal data;
[0018] Performing time synchronization and spatial registration operations on each modality data in the second multimodal data to obtain multimodal spatiotemporal correlation data, wherein the spatiotemporal correlation data includes second visual data, second audio data, second position relationship data, and second device status data;
[0019] Annotating the spatiotemporal correlation data to obtain a first data set;
[0020] An instruction dataset is constructed based on the first dataset and the first language model.
[0021] In some embodiments, labeling the spatiotemporal correlation data includes:
[0022] Annotating visual information and time sequence information in the second visual data, wherein the visual information includes the potential safety hazard object and the location and status information of the potential safety hazard object;
[0023] extracting and annotating audio information in the second audio data, wherein the audio information includes abnormal sounds;
[0024] Calculate, based on the second position relationship data, a first distance between the worker's position and the first preset area, a second distance between the construction equipment position and the second preset area, and a third distance between different construction equipment positions, and mark the first distance, the second distance, and the third distance;
[0025] Annotating semantic information in the second device status data, wherein the semantic information includes a device operating speed and / or a device load status;
[0026] Annotate the safety hazard types corresponding to the second visual data, the second audio data, the second position relationship data, and the second device status data;
[0027] The different modal data corresponding to the same safety hazard type, time and space in the spatiotemporal correlation data are associated and labeled.
[0028] In some embodiments, the first data set includes annotated spatiotemporal correlation data, the instruction data set includes instruction data, each instruction data includes a detection instruction and a set of annotated spatiotemporal correlation data and a detection result, and the detection result includes a safety hazard type;
[0029] Based on the first data set and the first language model, an instruction data set is constructed, including:
[0030] Parsing the annotated spatiotemporal correlation data based on the first language model, and generating detection instructions in combination with a preset instruction template;
[0031] Associating the detection instruction with the annotated spatiotemporal correlation data and the detection result to generate instruction data;
[0032] Among them, the detection instructions include at least one of checking personnel safety hazards, checking material safety hazards, checking construction process hazards, checking equipment safety hazards, checking spatial location hazards and checking environmental safety hazards, and the types of safety hazard include at least one of personnel safety hazards, material safety hazards, equipment safety hazards, environmental safety hazards, construction process hazards and spatial location hazards.
[0033] In some embodiments, the security risk detection model includes a multimodal encoder, an interface layer, and a second language model;
[0034] Before training the safety hazard detection model based on the instruction dataset, the method further includes:
[0035] Obtain a pre-trained multimodal encoder and a pre-trained second language model;
[0036] Keeping parameters of the multimodal encoder and the second language model unchanged, training the potential safety hazard detection model based on the second data set to update parameters of the interface layer, wherein the second data set is different from the instruction data set;
[0037] The safety hazard detection model is trained based on the instruction dataset, including:
[0038] Keeping the parameters of the multimodal encoder unchanged, the safety hazard detection model is trained based on the instruction dataset to update the parameters of the interface layer and the second language model.
[0039] In some embodiments, the multimodal encoder includes a text encoder, a visual encoder, and an audio encoder;
[0040] The text encoder is used to extract features from the second position relationship data and the second device status data to obtain a first text feature vector;
[0041] The visual encoder is used to extract features from the second visual data to obtain a visual feature vector;
[0042] The audio encoder is used to extract features from the second audio data to obtain an audio feature vector;
[0043] The interface layer is used to align the feature space of the visual feature vector and the audio feature vector with the text feature space of the second language model to obtain a second text feature vector;
[0044] The second language model is used to extract text features from the first text feature vector and the second text feature vector based on the detection instruction to detect whether there are corresponding safety hazards, and output the corresponding safety hazard type when safety hazards are detected.
[0045] In a second aspect, an embodiment of the present application provides a method for detecting potential safety hazards, including:
[0046] Acquire target multimodal data in target engineering environment;
[0047] Detect target multimodal data based on a potential safety hazard detection model to determine whether a target engineering environment has potential safety hazards, wherein the potential safety hazard detection model is trained using the potential safety hazard detection model training method of the first aspect.
[0048] In some embodiments, before detecting the target multimodal data based on the potential safety hazard detection model, the potential safety hazard detection method further includes:
[0049] Performing preprocessing operations on the target multimodal data to obtain multimodal data to be detected;
[0050] Obtaining a detection instruction, and inputting the detection instruction and the multimodal data to be detected into a potential safety hazard detection model, so that the potential safety hazard detection model detects the multimodal data to be detected;
[0051] When it is determined that there are safety hazards in the target engineering environment, the safety hazard detection method also includes:
[0052] Output the safety hazard type based on the safety hazard detection model;
[0053] Perform early warning processing based on the type of safety hazard and preset early warning strategies.
[0054] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the processor implements the steps of any method for training a safety hazard detection model as proposed in the first aspect, or the steps of any method for detecting safety hazard as proposed in the second aspect.
[0055] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions executable by a processor are stored. When the computer program instructions are called by the processor, the processor executes any one of the training methods for the safety hazard detection model proposed in the first aspect, or any one of the safety hazard detection methods proposed in the second aspect.
[0056] The beneficial effect of the implementation mode of the present application is: different from the prior art, the implementation mode of the present application provides a training method for a safety hazard detection model, and the training method for the safety hazard detection model includes: obtaining first multimodal data of an engineering environment; performing data association processing on the first multimodal data to construct an instruction data set; and training the safety hazard detection model based on the instruction data set.
[0057] By acquiring the first multimodal data of the engineering environment, performing data association processing on the first multimodal data to construct an instruction data set, and training a safety hazard detection model based on the instruction data set, the present application can comprehensively consider various modal data in the engineering environment, so that the safety hazard detection model can realize hazard detection based on the correlation between different modal data, avoid the limitations of single modal detection, and improve the accuracy of safety hazard detection in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] One or more embodiments are exemplarily described by corresponding drawings, which do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and the figures in the drawings do not constitute proportional limitations unless otherwise stated.
[0059] Figure 1 It is a schematic diagram of an application environment of a training method for a safety hazard detection model provided in an embodiment of the present application;
[0060] Figure 2 It is a flowchart of a training method for a safety hazard detection model provided in an embodiment of the present application;
[0061] Figure 3 It is a structural schematic diagram of a safety hazard detection model provided in an embodiment of the present application;
[0062] Figure 4 It is a detailed structural diagram of a safety hazard detection model provided in an embodiment of the present application;
[0063] Figure 5 It is a structural schematic diagram of a training device for a safety hazard detection model provided in an embodiment of the present application;
[0064] Figure 6 It is a flowchart of a safety hazard detection method provided in an embodiment of the present application;
[0065] Figure 7 It is a structural schematic diagram of a safety hazard detection device provided in an embodiment of the present application;
[0066] Figure 8 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application.
[0067] Description of Figure Numbers:
[0068]
[0069] DETAILED DESCRIPTION
[0070] In order to facilitate the understanding of the present application, the present application is described in more detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that when an element is described as "fixed to" another element, it can be directly on the other element, or there can be one or more centered elements therebetween. When an element is described as "connected to" another element, it can be directly connected to the other element, or there can be one or more centered elements therebetween. The terms "vertical", "horizontal", "left", "right" and similar expressions used in this specification are for illustrative purposes only.
[0071] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.
[0072] The technical solution of the present application is described in detail below in conjunction with the accompanying drawings:
[0073] See also Figure 1 , Figure 1 It is a schematic diagram of the application environment of a training method for a safety hazard detection model provided in an embodiment of the present application.
[0074] like Figure 1 As shown, the application environment 100 includes: a data acquisition device 10 and an electronic device 20. The data acquisition device 10 is connected to the electronic device 20 via a network communication, and the network includes a wired network and / or a wireless network. It is understood that the network includes wireless networks such as 2G, 3G, 4G, 5G, wireless LAN, Bluetooth, etc., and may also include wired networks such as serial cables and network cables.
[0075] The data acquisition device 10 is communicatively connected to the electronic device 20, and is used to collect multimodal data in an engineering environment and send the collected multimodal data to the electronic device 20. The engineering environment refers to an actual workplace established to complete a specific engineering task, including a collection of related equipment, systems, processes and personnel. Multimodal data refers to a collection of information covering multiple forms and dimensions obtained through multiple different types of sensors or data acquisition methods, for example: multimodal data includes visual data, audio data, position relationship data and equipment status data.
[0076] In the embodiment of the present application, the engineering environment includes but is not limited to the environment of the construction site, and the data acquisition device 10 includes a camera 11, a sound acquisition device 12, a position acquisition device 13 and an equipment status monitoring device 14.
[0077] The electronic device 20 is communicatively connected to the data acquisition device 10 and is used to execute the training method of the safety hazard detection model or the safety hazard detection method in any of the following embodiments. The electronic device 20 includes a terminal or a server, wherein the terminal includes but is not limited to various terminals with computing processing capabilities such as a laptop computer, a desktop computer or a mobile device, and the server includes but is not limited to a tower server, a rack server, a blade server, and a cloud server.
[0078] In some embodiments, the multimodal data includes first multimodal data, and the first multimodal data is multimodal data for training a safety hazard detection model. The data acquisition device 10 is used to acquire the first multimodal data and send the first multimodal data to the electronic device 20. The electronic device 20 is used to receive the first multimodal data and execute the training method of the safety hazard detection model to train the safety hazard detection model based on the first multimodal data to obtain a trained safety hazard detection model.
[0079] In some embodiments, the multimodal data includes target multimodal data, which is multimodal data that needs to be detected by a trained safety hazard detection model. The data acquisition device 10 is used to acquire the target multimodal data and send the target multimodal data to the electronic device 20. The electronic device 20 is used to receive the target multimodal data and execute the safety hazard detection method to detect the target multimodal data based on the safety hazard detection model to determine whether there are safety hazards in the target engineering environment.
[0080] See also Figure 2 , Figure 2 It is a flowchart of a training method for a safety hazard detection model provided in an embodiment of the present application.
[0081] Among them, the training method of the potential safety hazard detection model is applied to electronic devices, such as: Figure 1In the electronic device 20 shown, specifically, the execution subject of the training method of the potential safety hazard detection model is one or at least two processors of the electronic device.
[0082] like Figure 2 As shown, the training method of the potential safety hazard detection model includes steps S201 to S203:
[0083] Step S201: Acquire first multimodal data of an engineering environment.
[0084] The engineering environment includes but is not limited to the environment of the construction site. The following takes the environment of the construction site as an example to illustrate the specific implementation of the training method of the safety hazard detection model. For other types of engineering environments, the specific implementation of the training method of the safety hazard detection model is similar to this and will not be repeated here.
[0085] Among them, the first multimodal data is multimodal data collected by a data acquisition device and used to train a safety hazard detection model. The first multimodal data includes first visual data, first audio data, first position relationship data and first device status data.
[0086] Specifically, first multimodal data collected by a data collection device in an engineering environment is obtained. The data collection device is deployed in the engineering environment, and the data collection device includes a camera, a sound collection device, a position collection device, and an equipment status monitoring device. The camera is used to collect first visual data, the sound collection device is used to collect first audio data, the position collection device is used to collect first position relationship data, and the equipment status monitoring device is used to collect first equipment status data.
[0087] In the embodiment of the present application, obtaining the first multimodal data of the engineering environment specifically includes steps S211 to S214:
[0088] Step S211: Acquire first visual data collected by the camera device in the engineering environment.
[0089] Among them, the first visual data is a picture or video collected by a camera device and used to train a safety hazard detection model. The camera device includes but is not limited to a camera or other device used to collect images or videos.
[0090] Specifically, after the camera device deployed in the engineering environment collects the first visual data, the first visual data is received.
[0091] The first visual data includes at least one of individual protection information, work environment information, and worker behavior information. The individual protection information includes information such as the appearance, color, shape, and wearing position of protective equipment worn by workers to ensure personal safety. The protective equipment includes but is not limited to safety helmets, safety vests, goggles, and safety belts. Exemplarily, the individual protection information includes at least one of the appearance, color, and wearing position of a safety helmet and / or safety vest.
[0092] The work environment information includes information such as the physical environment and spatial layout related to the work activities in the engineering environment. The physical environment includes but is not limited to environmental factors such as on-site lighting intensity and on-site visibility. The spatial layout includes but is not limited to the placement of construction equipment, the placement of safety signs, and the placement of building materials. Exemplarily, the work environment information includes at least one of the physical environment and spatial layout of the construction site.
[0093] Worker behavior information includes information such as the actions and postures of workers in the engineering environment, including but not limited to walking, climbing, squatting, carrying, and operating construction equipment. Exemplarily, worker behavior information includes the actions and postures of workers at the construction site.
[0094] Step S212: Acquire first audio data collected by the sound collection device in the engineering environment.
[0095] Among them, the first audio data is the sound collected by the sound collection device and used to train the safety hazard detection model. The sound collection device includes but is not limited to a sound sensor and other devices for collecting sound.
[0096] Specifically, after the camera device deployed in the engineering environment collects the first audio data, the first audio data is received.
[0097] The first audio data includes at least one of equipment operation sound, gas release sound, and liquid flow sound. Equipment operation sound includes the sound generated by construction equipment during operation, gas release sound refers to the sound generated when gas is released through leakage, discharge, or injection, and liquid flow sound refers to the sound generated when liquid flows in a pipe, container, or equipment.
[0098] In some embodiments, the first audio data includes at least one of the equipment operation sound, gas release sound, and liquid flow sound collected when there is a safety hazard in the engineering environment corresponding to the first visual data. For example, the first audio data includes at least one of the equipment operation sound, gas release sound, and liquid flow sound when the safety helmet is not worn properly.
[0099] Step S213: Acquire first position relationship data collected by the position collection device in the engineering environment.
[0100] Among them, the first position relationship data is the position information collected by the position acquisition device and used to train the safety hazard detection model. The position acquisition device includes but is not limited to a position sensor, a positioning device using the Global Positioning System (GPS) or Radio Frequency Identification (RFID), an inertial measurement unit (IMU), and other devices for collecting position information.
[0101] Specifically, after the position acquisition device deployed in the engineering environment acquires the first position relationship data, the first position relationship data is received.
[0102] The first position relationship data includes at least one of worker position information and construction equipment position information, the worker position information includes real-time or historical position information of the worker in the engineering environment, and the construction equipment position information includes real-time or historical position information of the construction equipment in the engineering environment. Exemplarily, the worker position information includes at least one of the coordinate position and / or movement trajectory of the worker in the construction environment, and the construction equipment position information includes at least one of the coordinate position and / or movement trajectory of the construction equipment in the construction environment.
[0103] Step S214: Acquire first equipment status data collected by the equipment status monitoring device in the engineering environment.
[0104] Among them, the first equipment status data is the status information of the construction equipment collected by the equipment status monitoring device and used to train the safety hazard detection model. The equipment status monitoring device includes but is not limited to equipment status monitoring sensors and other devices used to monitor equipment status.
[0105] Specifically, after the equipment status monitoring device deployed in the engineering environment collects the first equipment status data, the first equipment status data is received.
[0106] The first equipment status data includes the operating status information of the construction equipment, and the construction equipment includes but is not limited to earthwork equipment, lifting and hoisting equipment, concrete engineering equipment, steel bar and formwork engineering equipment, road construction equipment, drilling and pile foundation equipment, power and welding equipment. The construction equipment is determined by the specific engineering environment and is not limited here. Exemplarily, the construction equipment includes lifting and hoisting equipment, and the operating status information of the construction equipment includes at least one of the operating speed and / or the weight carried by the lifting and hoisting equipment.
[0107] In the embodiment of the present application, the above steps S211 to S214 can be performed simultaneously without any restriction on the order.
[0108] In some embodiments, each piece of data in the first multimodal data corresponds to a timestamp, and the timestamp is time information added to the data by the data acquisition device when acquiring the data.
[0109] In some embodiments, obtaining the first multimodal data of the engineering environment further includes: obtaining the first multimodal data of the same engineering task in different construction stages and different construction scenes. Among them, the construction stage includes the foundation construction stage, the structure construction stage and the high-altitude operation stage. In the foundation construction stage, the construction scene includes foundation pit excavation, pile foundation construction or earthwork transportation. In the structure construction stage, the construction scene includes steel bar binding, concrete pouring or formwork support. In the high-altitude operation stage, the construction scene includes tower crane operation or facade construction.
[0110] In view of the diversity of construction site scenarios, different types of projects, different construction stages and task characteristics are different. By obtaining the first multimodal data of the same engineering task in different construction stages and different construction scenes, this application can train the safety hazard detection model to learn multimodal data in multiple construction scenarios, improve the generalization ability and adaptability of the model, and be able to quickly adapt to new construction environments and task requirements without the need for large-scale retraining for each specific scenario, thereby improving the efficiency of safety hazard detection.
[0111] Step S202: performing data association processing on the first multimodal data to construct an instruction data set.
[0112] The data association processing is used to perform temporal association and spatial association on each modal data in the first multimodal data to integrate the temporal and spatial consistency of the multimodal data. The types of modal data include first visual data, first audio data, first positional relationship data, or first device status data. The instruction data set is a data set used in the training of the safety hazard detection model.
[0113] Specifically, data association processing is performed on the first multimodal data to obtain multimodal spatiotemporal association data, the multimodal spatiotemporal association data is annotated to obtain a first data set, and an instruction data set is constructed based on the first data set. The multimodal spatiotemporal association data is multimodal data that is consistent in time and space, and the first data set is different from the instruction data set.
[0114] In an embodiment of the present application, before performing data association processing on the first multimodal data, the training method of the safety hazard detection model also includes: classifying the first multimodal data according to the type of device used when collecting the data.
[0115] Specifically, the device type includes a camera device, a sound collection device, a position collection device or a device status monitoring device. According to the corresponding device type, the acquired first multimodal data is divided into first visual data, first audio data, first position relationship data and first device status data.
[0116] In the embodiment of the present application, data association processing is performed on the first multimodal data to construct an instruction data set, which specifically includes steps S221 to S224:
[0117] Step S221: performing a preprocessing operation on the first multimodal data to obtain second multimodal data.
[0118] The preprocessing operation includes at least one of data cleaning, data transformation, signal processing and data standardization, and the second multimodal data is the first multimodal data after the preprocessing operation.
[0119] Specifically, the first multimodal data is cleaned to obtain cleaned first multimodal data. The data cleaning is used to remove noise, erroneous data, and data irrelevant to safety hazard detection. The cleaned first multimodal data is processed according to the type of modal data to obtain second multimodal data. The second multimodal data includes the pre-processed first visual data, the first audio data, the first position relationship data, and the first device status data.
[0120] Among them, according to the type of modal data, each modal data is processed, including: performing data transformation on the first visual data to obtain preprocessed first visual data; performing signal processing on the first audio data to obtain preprocessed first audio data; performing data standardization on the first position relationship data to obtain preprocessed first position relationship data; performing data standardization on the first device status data to obtain preprocessed first device status data.
[0121] Among them, data transformation includes but is not limited to operations such as cropping and scaling, signal processing includes but is not limited to filtering, and data standardization includes but is not limited to operations such as format conversion and calibration.
[0122] Step S222: performing time synchronization and spatial registration operations on each modality data in the second multimodal data to obtain multimodal spatiotemporal correlation data.
[0123] Among them, the time synchronization is used to uniformly map each modality data to the same time base, and the spatial registration operation is used to uniformly map each modality data to the same spatial coordinate system to ensure the correspondence in spatial position. The multimodal spatiotemporal correlation data is the second multimodal data after the time synchronization and spatial registration operations.
[0124] Specifically, each modality data in the second multimodal data is aligned through the timestamp corresponding to each data, so that they correspond at the same time point, thereby achieving time sequence synchronization. Then, through calibration, coordinate transformation and data fusion, these data are mapped to the same spatial coordinate system to obtain multimodal spatiotemporal correlation data.
[0125] For example, the coordinate transformation matrix between the camera device and other devices is obtained through camera calibration technology, and all modal data are mapped to the same spatial coordinate system through the coordinate transformation matrix. This method can make different modal data consistent in space, for example, the camera perspective corresponds to the position relationship data, so that the position marked by the person or equipment in the image is consistent with the coordinates in the position relationship data.
[0126] Among them, the spatiotemporal correlation data includes second visual data, second audio data, second positional relationship data, and second device status data. The second visual data is the visual data obtained by performing time synchronization and spatial registration operations on the preprocessed first visual data. The second audio data is the audio data obtained by performing time synchronization and spatial registration operations on the preprocessed first audio data. The second positional relationship data is the positional relationship data obtained by performing time synchronization and spatial registration operations on the preprocessed first positional relationship data. The second device status data is the device status data obtained by performing time synchronization and spatial registration operations on the preprocessed first device status data.
[0127] Step S223: labeling the spatiotemporal correlation data to obtain a first data set.
[0128] The first data set includes annotated spatiotemporal correlation data, where annotation refers to the process of adding semantic information or description labels to the data.
[0129] Specifically, the spatiotemporal correlation data are labeled according to the types of modal data to obtain labeled spatiotemporal correlation data. Several groups of labeled spatiotemporal correlation data constitute a first data set.
[0130] In the embodiment of the present application, the spatiotemporal correlation data is labeled, specifically including steps S2231 to S2236:
[0131] Step S2231: labeling the visual information and timing information in the second visual data.
[0132] The visual information includes the potential safety hazard object, the location and status information of the potential safety hazard object, the potential safety hazard object includes but is not limited to workers, safety helmets, safety ropes, construction equipment, and building materials, and the status information of the potential safety hazard object includes but is not limited to worker status information, protective equipment status information, construction equipment status information, and building material status information. The timing information includes the front-to-back sequence between image frames in the second visual data. The protective equipment status information includes but is not limited to the status information of the safety helmet and / or the status information of the safety rope.
[0133] Worker status information includes but is not limited to the worker's behavior, posture, and whether they are wearing protective equipment. Protective equipment status information includes but is not limited to whether they are wearing, the wearing position, whether the wearing position is correct, and whether they are damaged. For example, the status information of a safety helmet includes but is not limited to whether it is wearing, the wearing position, and whether it is damaged. The status information of a safety rope includes but is not limited to whether it is in use, whether it is broken, and whether it is firmly connected. The status information of construction equipment includes but is not limited to whether the construction equipment is in operation and whether it has a fault. The status information of building materials includes but is not limited to whether the building materials are scattered and stacked, and whether the stacking height is greater than the preset stacking height. Among them, the preset stacking height can be set by a person skilled in the art according to the actual engineering environment, and is not limited here.
[0134] Specifically, based on the image deep learning algorithm, the safety hazard object in the second visual data is identified, the relevant visual information is annotated, and the timing information of the continuous image frames is annotated according to the timestamp. Among them, the image deep learning algorithm includes but is not limited to a pre-trained convolutional neural network. Exemplarily, the image deep learning algorithm is a pre-trained region-based convolutional neural network (Region-Based Convolutional Neural Network, R-CNN). The pre-trained convolutional neural network can be trained by a technician in this field based on a relevant image data set. The method of training the model through the data set belongs to the prior art and will not be repeated here.
[0135] Step S2232: extract and annotate the audio information in the second audio data.
[0136] The audio information includes abnormal sounds, and the abnormal sounds include at least one of abnormal equipment operation sounds, abnormal gas release sounds, and abnormal liquid flow sounds. The abnormal equipment operation sounds include abnormal sounds or warning signals generated by the construction equipment during operation, for example, the abnormal equipment operation sounds are the impact or shaking sounds caused by the loosening of internal parts of the mechanical equipment, or the sharp noises caused by motor failure.
[0137] Abnormal sound of gas release refers to the abnormal sound made when gas is released through leakage, discharge or injection, for example: the sharp sound or injection sound when high-pressure gas leaks due to pipeline rupture. Abnormal sound of liquid flow refers to the abnormal sound made when liquid flows in pipelines, containers or equipment, for example: the vibration sound produced when the valve is blocked and the liquid oscillates.
[0138] Specifically, the second audio data is analyzed by a sound processing algorithm to extract and annotate the audio information, wherein the sound processing algorithm includes but is not limited to frequency domain and time domain analysis methods, signal feature extraction algorithms, deep learning and machine learning algorithms.
[0139] Step S2233: Based on the second position relationship data, calculate the first distance between the worker position and the first preset area, the second distance between the construction equipment position and the second preset area, and the third distance between different construction equipment positions, and mark the first distance, second distance, and third distance.
[0140] The first preset area is a prohibited area or a dangerous area for workers in the engineering environment, and the second preset area is a specific functional area or a restricted area set for construction equipment. The worker position is the position of the worker, the construction equipment position is the position of the construction equipment, the first distance is the shortest distance between the position of the worker and the first preset area, the second distance is the shortest distance between the position of the construction equipment and the second preset area, and the third distance is the distance between different construction equipment positions. The coordinate ranges of the first preset area and the second preset area can be set by those skilled in the art according to the actual engineering environment, and are not limited here.
[0141] Specifically, the coordinate position of each worker and / or construction equipment is obtained from the second position relationship data. According to the coordinate position of each worker and the coordinate range of the first preset area, the first distance between each worker position and the first preset area is calculated. According to the coordinate position of each construction equipment and the coordinate range of the second preset area, the second distance between each construction equipment position and the second preset area is calculated. According to the coordinate position of each construction equipment, the third distance between each two different construction equipment positions is calculated. The first distance, the second distance and the third distance are marked correspondingly in the second position relationship data.
[0142] In some embodiments, calculating the first distance between each worker's position and the first preset area includes: when the first preset area is a circular area, calculating the fourth distance between the coordinate position of each worker and the center coordinates of the first preset area; when the fourth distance is less than the radius of the first preset area, the worker is located in the first preset area, and the first distance is set to zero; when the fourth distance is equal to the radius of the first preset area, the worker is located at the boundary of the first preset area, and the first distance is set to the radius of the first preset area; when the fourth distance is greater than the radius of the first preset area, the worker is outside the first preset area, and the first distance is set to the difference between the fourth distance and the radius of the first preset area. Among them, the fourth distance is the distance between the coordinate position of the worker and the center coordinates of the first preset area. The center coordinates and radius of the first preset area can be set by those skilled in the art according to the actual engineering environment, and are not limited here.
[0143] In some embodiments, calculating the first distance between each worker position and the first preset area includes: when the first preset area is a non-circular area, dividing the boundary of the first preset area into a plurality of line segments; for the coordinate position of any worker, traversing each line segment, calculating the shortest distance from the coordinate position to the current line segment, taking the shortest distance as the fifth distance, taking the minimum value of the plurality of fifth distances as the first distance between the worker position and the first preset area, and repeating the traversal steps until the first distance between each worker position and the first preset area is calculated. The fifth distance is the shortest distance between the coordinate position of any worker and any line segment of the first preset area, and the calculation method of the fifth distance includes but is not limited to the shortest distance formula between a point and a line segment, such as the vector projection method.
[0144] In some embodiments, calculating the second distance between each construction equipment position and the second preset area includes: when the first preset area is a circular area, calculating the sixth distance between the coordinate position of each construction equipment and the center coordinates of the second preset area; when the sixth distance is less than the radius of the second preset area, the construction equipment is located in the second preset area, and the second distance is set to zero; when the sixth distance is equal to the radius of the second preset area, the construction equipment is located at the boundary of the second preset area, and the second distance is set to the radius of the second preset area; when the sixth distance is greater than the radius of the second preset area, the construction equipment is located outside the second preset area, and the second distance is set to the difference between the sixth distance and the radius of the second preset area. Among them, the sixth distance is the distance between the coordinate position of the construction equipment and the center coordinates of the second preset area. The center coordinates and radius of the second preset area can be set by those skilled in the art according to the actual engineering environment, and are not limited here.
[0145] In some embodiments, the second distance between each construction equipment position and the second preset area is calculated, including: when the second preset area is a non-circular area, dividing the boundary of the second preset area into a plurality of line segments; for the coordinate position of any construction equipment, traversing each line segment, calculating the shortest distance from the coordinate position to the current line segment, taking the shortest distance as the seventh distance, taking the minimum value of the plurality of seventh distances as the second distance between the construction equipment position and the second preset area, and repeating the traversal step until the second distance between each construction equipment position and the second preset area is calculated. The seventh distance is the shortest distance between the coordinate position of any construction equipment and any line segment of the second preset area, and the calculation method of the seventh distance includes but is not limited to the shortest distance formula between a point and a line segment such as the vector projection method.
[0146] In some embodiments, calculating the third distance between each two different construction equipment positions includes: calculating the distance between each two different construction equipment coordinate positions based on the Euclidean geometric distance formula, and using the distance as the third distance between the two different construction equipment. The Euclidean geometric distance formula is an existing formula and will not be described in detail here.
[0147] Step S2234: label the semantic information in the second device status data.
[0148] The semantic information includes the equipment operating speed and / or the equipment load status. The equipment operating speed is the operating speed of the construction equipment. The equipment load status includes a normal state or an overweight state. In the normal state, the weight carried by the construction equipment is less than or equal to the preset load; in the overweight state, the weight carried by the construction equipment is greater than the preset load. The preset load is the maximum load of the construction equipment under safe construction conditions. The preset load can be set by technicians in this field according to the safe operation standards and engineering requirements of the construction equipment, and is not limited here.
[0149] Specifically, the second equipment status data is analyzed to determine whether the operating speed of the construction equipment is greater than the preset operating speed. When the operating speed of the construction equipment is less than or equal to the preset operating speed, the operating speed of the construction equipment is marked as normal speed. When the operating speed of the construction equipment is greater than the preset operating speed, the operating speed of the construction equipment is marked as overspeeding. The preset speed is the maximum speed of the construction equipment under safe construction conditions, the normal speed is the operating speed less than or equal to the preset speed, and the overspeeding state is the state where the operating speed of the construction equipment is greater than the preset speed. The preset speed can be set by those skilled in the art according to the safe operation standards and engineering requirements of the construction equipment, and is not limited here.
[0150] Determine whether the weight carried by the construction equipment is greater than the preset load. When the weight carried by the construction equipment is less than or equal to the preset load, mark the weight carried by the construction equipment as normal; when the weight carried by the construction equipment is greater than the preset load, mark the weight carried by the construction equipment as overweight.
[0151] Among them, the methods for labeling semantic information include but are not limited to rule-based threshold judgment methods, unsupervised learning-based anomaly detection algorithms, and machine learning-based classification models.
[0152] In an embodiment of the present application, the above steps S2231 to S2234 can be performed simultaneously without any restriction on the order of precedence.
[0153] Step S2235: label the safety hazard types corresponding to the second visual data, the second audio data, the second position relationship data and the second device status data.
[0154] Among them, the safety hazard type is the type of safety hazard that is easy to exist in the engineering environment. The safety hazard refers to the potential risk factor that may cause harm and loss to personnel, equipment, environment, etc. The safety hazard type can be set by those skilled in the art according to safety standards and is not limited here. Exemplarily, the safety hazard type includes at least one of personnel safety hazard, material safety hazard, equipment safety hazard, environmental safety hazard, construction process hazard and spatial location hazard.
[0155] Personnel safety hazards refer to potential safety risks directly related to the safety of staff, material safety hazards refer to potential safety risks related to the materials used, equipment safety hazards refer to safety hazards related to the operation, operation and maintenance of equipment and facilities, environmental safety hazards refer to potential safety risks related to the environment, construction process hazards refer to potential safety risks caused by factors such as non-standard construction processes, unreasonable steps, and inadequate management, and spatial location hazards refer to potential safety risks related to spatial layout, location arrangement or site design.
[0156] Specifically, the type of safety hazard corresponding to each piece of data in the second visual data, the second audio data, the second position relationship data, and the second device status data is respectively labeled by a labeling tool, wherein the labeling tool includes but is not limited to a detection model or a classification model based on machine learning.
[0157] For example, when the visual information shows that the worker is not wearing protective equipment or the protective equipment is worn in the wrong position, the corresponding second visual data is marked as a safety hazard for personnel. When the marking tool detects that the first distance between the worker's position and the first preset area is less than the first preset safety distance, the corresponding second position relationship data is marked as a safety hazard for personnel. The first preset safety distance is the maximum safety distance between the worker's position and the first preset area.
[0158] When the visual information shows that the building materials are randomly stacked and / or the height of the stacking is greater than the preset stacking height, the safety hazard type of the corresponding second visual data is marked as a material safety hazard. When the audio information contains abnormal sounds of gas release and / or abnormal sounds of liquid flow, the safety hazard type of the corresponding second audio data is marked as an environmental safety hazard.
[0159] When the semantic information includes that the construction equipment is in an overspeed state and / or the construction equipment is in an overweight state, the safety hazard type of the corresponding second equipment status data is marked as equipment safety hazard. When the audio information includes abnormal equipment operation sound, the safety hazard type of the corresponding second audio data is marked as equipment safety hazard.
[0160] When damage to protective equipment and / or construction equipment failure occurs in the visual information, the safety hazard type of the corresponding second visual data is marked as a construction process hazard. When the behavior or posture of any worker in the visual information is identified as a preset dangerous action by the annotation tool, the safety hazard type of the corresponding second visual data is marked as a construction process hazard. Among them, the preset dangerous action can be set by those skilled in the art according to the safety specifications and standards of the engineering environment, and is not limited here.
[0161] When the marking tool detects that the second distance between the construction equipment position and the second preset area is less than the second preset safety distance, and / or the third distance between different construction equipment positions is less than the third preset safety distance, the safety hazard type of the corresponding second position relationship data is marked as a spatial position hazard. The second preset safety distance is the maximum safety distance between the construction equipment position and the second preset area, and the third preset safety distance is the maximum safety distance between different construction equipment positions.
[0162] Step S2236: associate and annotate different modal data corresponding to the same safety hazard type, time and space in the spatiotemporal correlation data.
[0163] Specifically, different types of modal data with the same safety hazard type, the same timestamp, and the same coordinate position in the same spatial coordinate system are associated and annotated to make the multimodal data consistent in time, space, and safety hazard type.
[0164] In some embodiments, the above steps S2231 to S2236 can be performed by the same pre-trained model to complete the data annotation. The pre-trained model uses a deep learning algorithm, and the pre-trained model can be trained by those skilled in the art based on relevant image data sets, audio data sets, and semantic information data sets. The method of training the model through the data set belongs to the prior art and will not be repeated here.
[0165] In some embodiments, after the pre-trained model completes labeling of the spatiotemporal correlation data, relevant technical personnel check the labeled spatiotemporal correlation data to correct erroneous labeling information and obtain a first data set.
[0166] Step S224: constructing an instruction data set based on the first data set and the first language model.
[0167] The first language model is used to parse the annotated spatiotemporal correlation data to generate a detection instruction. The first language model includes but is not limited to a large language model (LLM), and the detection instruction is used to instruct the safety hazard detection model to perform a corresponding type of safety hazard detection on the input multimodal data. The detection instruction can be set by a person skilled in the art according to the type of safety hazard, and is not limited here.
[0168] The instruction data set is used to train the safety hazard detection model. The instruction data set includes instruction data. The instruction data consists of an optional detection instruction and a set of input-output pairs of the safety hazard detection model. That is, each instruction data includes a detection instruction and a set of annotated spatiotemporal correlation data and detection results. The detection result includes whether there are safety hazards of the corresponding type in the engineering environment. When there are safety hazards in the engineering environment, the detection result includes the type of safety hazards.
[0169] Specifically, a detection instruction is generated based on the annotated spatiotemporal correlation data in the first data set and the first language model, and instruction data is generated according to the detection instruction, the annotated spatiotemporal correlation data and the detection result. Several instruction data constitute an instruction data set.
[0170] In the embodiment of the present application, based on the first data set and the first language model, an instruction data set is constructed, which specifically includes steps S2241 and S2242:
[0171] Step S2241: parsing the annotated spatiotemporal correlation data based on the first language model, and generating a detection instruction in combination with a preset instruction template.
[0172] Among them, the detection instructions include at least one of checking personnel safety hazards, checking material safety hazards, checking construction process hazards, checking equipment safety hazards, checking spatial location hazards, and checking environmental safety hazards. The preset instruction template includes preset input content, preset output instructions, and output format of the first language model, the preset input content includes a description text of the multimodal data, the preset output instructions include the detection instructions corresponding to the preset input content, and the output format is the format of the detection instructions output by the first language model. The preset output instructions can be set by a technician in this field according to the preset input content, and there is no limitation here.
[0173] Specifically, the annotated spatiotemporal correlation data is parsed through the first language model, the annotation information of the data is converted into a detailed description text of the multimodal data, and the preset instruction template is matched, the key information in the detailed description text is mapped to the corresponding field in the preset instruction template, and a standardized detection instruction is generated.
[0174] Step S2242: Associating the detection instruction with the annotated spatiotemporal correlation data and the detection result to generate instruction data.
[0175] Specifically, each detection instruction is integrated with the annotated spatiotemporal correlation data and the detection results to establish an association relationship, thereby generating each complete instruction data.
[0176] See also Figure 3 , Figure 3 It is a structural diagram of a safety hazard detection model provided in an embodiment of the present application.
[0177] like Figure 3 As shown, the safety hazard detection model 300 includes a multimodal encoder 301 , an interface layer 302 and a second language model 303 .
[0178] The multimodal encoder 301 is respectively connected to the interface layer 302 and the second language model 303, and is used to obtain multimodal data and perform feature extraction on different types of modal data to obtain corresponding feature vectors. The interface layer 302 is respectively connected to the multimodal encoder 301 and the second language model 303, and is used to perform feature space alignment on the feature vectors input to the interface layer 302, so that the feature spaces of different feature vectors input to the second language model 303 are compatible. The second language model 303 is respectively connected to the multimodal encoder 301 and the interface layer 302, and is used to obtain detection instructions and perform text feature extraction on the feature vectors input to the second language model 303 to detect whether there are corresponding safety hazards.
[0179] See also Figure 4 , Figure 4 It is a detailed structural diagram of a safety hazard detection model provided in an embodiment of the present application.
[0180] like Figure 4 As shown, the safety hazard detection model 300 includes a multimodal encoder 301 , an interface layer 302 , and a second language model 303 . The multimodal encoder 301 includes a text encoder 311 , a visual encoder 312 , and an audio encoder 313 .
[0181] The text encoder 311 is connected to the second language model 303 and is used to extract features from the second position relationship data and the second device status data to obtain a first text feature vector. The first text feature vector is a feature vector obtained by extracting features from the second position relationship data and the second device status data, and the feature vector is used to encode key information in the text data (i.e., the second position relationship data and the second device status data), such as semantic information, word frequency features, text distribution features, etc.
[0182] The visual encoder 312 is connected to the interface layer 302 and is used to extract features from the second visual data to obtain a visual feature vector, wherein the visual feature vector is a feature vector obtained by extracting features from the second visual data, and the feature vector is used to encode key information in the visual data, such as shape features, spatial features, spatiotemporal features, etc.
[0183] The audio encoder 313 is connected to the interface layer 302 and is used to extract features from the second audio data to obtain an audio feature vector. The audio feature vector is a feature vector obtained by extracting features from the second audio data, and the feature vector is used to encode key information in the audio data, such as time domain energy, spectrum features, pitch features, etc.
[0184] The interface layer 302 is connected to the visual encoder 312, the audio encoder 313 and the second language model 303 respectively, and is used to align the feature space of the visual feature vector and the audio feature vector with the text feature space of the second language model to obtain a second text feature vector. The second text feature vector is a unified text feature representation obtained by mapping the visual feature vector and the audio feature vector to the text feature space. The interface layer 302 includes but is not limited to a neural network layer such as a fully connected layer or a linear mapping layer for aligning the feature spaces of different feature vectors.
[0185] The second language model 303 is connected to the text encoder 311 and the interface layer 302 respectively, and is used to extract text features from the first text feature vector and the second text feature vector based on the detection instruction to detect whether there is a corresponding security risk, and output the corresponding security risk type when the security risk is detected. The second language model 303 includes but is not limited to a large language model.
[0186] By using an interface layer to align the feature space of the visual feature vector and the audio feature vector with the text feature space of the second language model, the present application can transform the features of the non-text modality into a feature space that is compatible with the text features of the second language model, thereby achieving alignment of the text feature space and other feature spaces.
[0187] In the embodiment of the present application, before training the potential safety hazard detection model based on the instruction data set, the potential safety hazard detection model training method further includes steps S1 and S2:
[0188] Step S1: Obtain a pre-trained multimodal encoder and a pre-trained second language model.
[0189] Specifically, the multimodal encoder and the second language model are pre-trained respectively using a public dataset to obtain a pre-trained multimodal encoder and a pre-trained second language model.
[0190] Alternatively, a pre-trained text encoder, pre-trained visual encoder, pre-trained audio encoder, and pre-trained second language model are obtained through the Internet. For example, the pre-trained visual encoder uses the Contrastive Language-Image Pre-training (CLIP) model available on the Internet, and the pre-trained second language model uses the Large Language Model Meta AI (LLaMA) available on the Internet.
[0191] In some embodiments, obtaining a pre-trained multimodal encoder includes: jointly training a text encoder and a visual encoder through an image-text dataset, and jointly training a text encoder and an audio encoder through an audio-text dataset, thereby obtaining a pre-trained text encoder, a pre-trained visual encoder, and a pre-trained audio encoder. Among them, the image-text dataset includes several image-text pairs, and the audio-text dataset includes several audio-text pairs. The image-text dataset includes but is not limited to public datasets such as LAION-400M, and the audio-text dataset includes but is not limited to public datasets such as WavCaps.
[0192] By jointly training the text encoder and the visual encoder through the image-text dataset, and jointly training the text encoder and the audio encoder through the audio-text dataset, the present application can enable the multimodal encoder to obtain low-level modal information and high-level modal information through pre-training. Among them, low-level modal information is mainly information at the perception level, including: basic features of images, audio, and text. High-level modal information is more abstract, semantic-level information, including: object categories in images, relationships between objects, semantic content of audio, contextual semantics of text, sentence meaning, described scenes, etc.
[0193] In some embodiments, obtaining a pre-trained second language model includes: training the second language model through a text dataset to obtain the pre-trained second language model. The text dataset may be a text dataset publicly available on the Internet.
[0194] Step S2: Keep the parameters of the multimodal encoder and the second language model unchanged, and train the potential safety hazard detection model based on the second data set to update the parameters of the interface layer.
[0195] The second data set is different from the first data set, the second data set is different from the instruction data set, and the second data set is a public data set on the Internet, such as an image-text data set or an audio-text data set. The second data set can be selected by a person skilled in the art according to the data set used when the multimodal encoder is pre-trained, and is not limited here.
[0196] Specifically, the parameters of the pre-trained multimodal encoder and the pre-trained second language model are kept unchanged, and the safety hazard detection model is trained based on the second data set to update the parameters of the interface layer. Among them, the method of training the model through the data set belongs to the prior art and will not be repeated here.
[0197] By keeping the parameters of the pre-trained multimodal encoder and the pre-trained second language model unchanged and training only updates the parameters of the interface layer, the present application can achieve the alignment of feature vectors of different modal data in the feature space, thereby effectively integrating multimodal information without affecting the original performance of the second language model, improving the performance of cross-modal tasks, reducing computational overhead and resource consumption, and improving the efficiency and stability of model training.
[0198] Step S203: training a safety hazard detection model based on the instruction data set.
[0199] Specifically, after executing step S2, the safety hazard detection model is trained based on the instruction data set to adjust the parameters of the safety hazard detection model to improve the generalization of the model.
[0200] In an embodiment of the present application, training a safety hazard detection model based on an instruction data set specifically includes: keeping the parameters of the multimodal encoder unchanged, training the safety hazard detection model based on the instruction data set to update the parameters of the interface layer and the second language model.
[0201] Specifically, after executing step S2, the parameters of the pre-trained multimodal encoder are kept unchanged, and the potential safety hazard detection model is trained based on the instruction data set to update the parameters of the interface layer and the second language model.
[0202] The first multimodal data includes first visual data, first audio data, first position relationship data and first device status data, and data association processing is performed on the first multimodal data to construct an instruction data set, and a safety hazard detection model is trained based on the instruction data set. On the one hand, the present application can comprehensively consider various modal data in the engineering environment, so that the safety hazard detection model can realize hazard detection based on the correlation between different modal data, avoid the limitations of single modal detection, improve the accuracy of safety hazard detection in complex environments, and reduce the probability of false alarms and missed alarms.
[0203] On the other hand, compared with the existing text-based status analysis system, which can only provide quantitative construction area status information and is difficult to judge whether there are safety hazards based on the multi-faceted information of the construction site, this application can integrate multimodal data to train the model in view of the complex and changeable characteristics of the construction site environment, so that the safety hazard detection model can adapt to safety hazard monitoring under different conditions such as lighting, noise, and equipment operation status, and improve the stability and reliability of the model in various extreme scenarios. And through comprehensive analysis of multimodal data, it can detect common safety hazards and potential safety hazards, thereby improving the comprehensiveness of safety hazard detection.
[0204] In some embodiments, the training method of the safety hazard detection model further includes: regularly acquiring third multimodal data of the engineering environment; performing data association processing on the third multimodal data to construct a new instruction data set; and training the safety hazard detection model based on the new instruction data set to adjust the parameters of the safety hazard detection model. The third multimodal data is different from the first multimodal data, and the third multimodal data is multimodal data collected by the data collection device and used to update the parameters of the safety hazard detection model.
[0205] Specifically, the third multimodal data of the engineering environment is collected regularly through the data collection equipment, and the third multimodal data is subjected to data association processing to construct a new instruction data set, and then the safety hazard detection model is trained based on the new instruction data set to adjust the parameters of the safety hazard detection model. Among them, the specific implementation method of performing data association processing on the third multimodal data is similar to the specific implementation method of performing data association processing on the first multimodal data, and will not be repeated here. The specific implementation method of training the safety hazard detection model based on the new instruction data set is similar to the specific implementation method of step S203, and will not be repeated here.
[0206] In some embodiments, the third multimodal data of the engineering environment is regularly collected by the data collection device, including: when new construction equipment or construction technology is introduced into the engineering environment, the third multimodal data of the engineering environment is collected by the data collection device.
[0207] By regularly acquiring the third multimodal data of the engineering environment, performing data association processing on the third multimodal data to construct a new instruction data set, and training the safety hazard detection model based on the new instruction data set, the present application can perform data association processing on the newly collected multimodal data as the equipment or construction process of the engineering environment changes, update the instruction data set and retrain the model, thereby improving the adaptability of the safety hazard detection model to new engineering environments or detection requirements.
[0208] In an embodiment of the present application, a training method for a safety hazard detection model is provided, and the training method for the safety hazard detection model includes: acquiring first multimodal data of an engineering environment; performing data association processing on the first multimodal data to construct an instruction data set; and training the safety hazard detection model based on the instruction data set.
[0209] By acquiring the first multimodal data of the engineering environment, performing data association processing on the first multimodal data to construct an instruction data set, and training a safety hazard detection model based on the instruction data set, the present application can comprehensively consider various modal data in the engineering environment, so that the safety hazard detection model can realize hazard detection based on the correlation between different modal data, avoid the limitations of single modal detection, and improve the accuracy of safety hazard detection in complex environments.
[0210] See also Figure 5 , Figure 5 It is a structural schematic diagram of a training device for a safety hazard detection model provided in an embodiment of the present application.
[0211] Among them, the training device of the safety hazard detection model is applied to an electronic device, such as a terminal or a server. Specifically, the training device of the safety hazard detection model is configured on the electronic device.
[0212] like Figure 5 As shown, the training device 500 of the potential safety hazard detection model includes:
[0213] The first data acquisition unit 501 is used to acquire first multimodal data of the engineering environment.
[0214] The data processing unit 502 is used to perform data association processing on the first multimodal data to construct an instruction data set.
[0215] The model training unit 503 is used to train the safety hazard detection model based on the instruction data set.
[0216] In some embodiments of the present application, the first multimodal data includes first visual data, first audio data, first position relationship data, and first equipment status data. The data acquisition unit 501 is specifically used to: acquire the first visual data collected by the camera device in the engineering environment, the first visual data includes at least one of individual protection information, working environment information, and worker behavior information; acquire the first audio data collected by the sound acquisition device in the engineering environment, the first audio data includes at least one of equipment operation sound, gas release sound, and liquid flow sound; acquire the first position relationship data collected by the position acquisition device in the engineering environment, the first position relationship data includes at least one of worker position information and construction equipment position information; acquire the first equipment status data collected by the equipment status monitoring device in the engineering environment, the first equipment status data includes the operation status information of the construction equipment.
[0217] In some embodiments of the present application, the data processing unit 502 is specifically used to: perform preprocessing operations on the first multimodal data to obtain second multimodal data; perform time synchronization and spatial registration operations on each modal data in the second multimodal data to obtain multimodal spatiotemporal correlation data, wherein the spatiotemporal correlation data includes second visual data, second audio data, second position relationship data, and second device status data; annotate the spatiotemporal correlation data to obtain a first data set; and construct an instruction data set based on the first data set and the first language model.
[0218] In some embodiments of the present application, the data processing unit 502 is also used to: annotate the visual information and timing information in the second visual data, the visual information including the safety hazard object, the position and status information of the safety hazard object; extract and annotate the audio information in the second audio data, the audio information including abnormal sounds; calculate the first distance between the worker position and the first preset area, the second distance between the construction equipment position and the second preset area, and the third distance between different construction equipment positions according to the second position relationship data, and annotate the first distance, the second distance, and the third distance; annotate the semantic information in the second equipment status data, the semantic information including the equipment operating speed and / or the equipment load status; annotate the safety hazard types corresponding to the second visual data, the second audio data, the second position relationship data, and the second equipment status data; and associate and annotate the different modal data corresponding to the same safety hazard type, time, and space in the spatiotemporal correlation data.
[0219] In some embodiments of the present application, the first data set includes annotated spatiotemporal correlation data, the instruction data set includes instruction data, each instruction data includes a detection instruction and a set of annotated spatiotemporal correlation data and detection results, and the detection results include safety hazard types. The data processing unit 502 is also used to: parse the annotated spatiotemporal correlation data based on the first language model, and generate detection instructions in combination with a preset instruction template; associate the detection instructions with the annotated spatiotemporal correlation data and the detection results to generate instruction data; wherein the detection instructions include at least one of checking personnel safety hazards, checking material safety hazards, checking construction process hazards, checking equipment safety hazards, checking spatial location hazards, and checking environmental safety hazards, and the safety hazard types include at least one of personnel safety hazards, material safety hazards, equipment safety hazards, environmental safety hazards, construction process hazards, and spatial location hazards.
[0220] In some embodiments of the present application, the safety hazard detection model includes a multimodal encoder, an interface layer, and a second language model. Before training the safety hazard detection model based on the instruction data set, the model training unit 503 is also used to: obtain a pre-trained multimodal encoder and a pre-trained second language model; keep the parameters of the multimodal encoder and the second language model unchanged, and train the safety hazard detection model based on the second data set to update the parameters of the interface layer, wherein the second data set is different from the instruction data set.
[0221] In some implementations of the present application, the model training unit 503 is specifically used to: keep the parameters of the multimodal encoder unchanged, train the safety hazard detection model based on the instruction data set, and update the parameters of the interface layer and the second language model.
[0222] In some embodiments of the present application, the multimodal encoder includes a text encoder, a visual encoder, and an audio encoder. The text encoder is used to perform feature extraction on the second position relationship data and the second device status data to obtain a first text feature vector; the visual encoder is used to perform feature extraction on the second visual data to obtain a visual feature vector; the audio encoder is used to perform feature extraction on the second audio data to obtain an audio feature vector; the interface layer is used to align the feature space of the visual feature vector and the audio feature vector with the text feature space of the second language model to obtain a second text feature vector; the second language model is used to perform text feature extraction on the first text feature vector and the second text feature vector based on the detection instruction to detect whether there is a corresponding safety hazard, and output the corresponding safety hazard type when a safety hazard is detected.
[0223] It can be understood that the implementation principle and technical effects of the training device 500 for the safety hazard detection model proposed in the present application can all refer to the implementation principle and technical effects of the steps of the training method for the safety hazard detection model proposed in the present application, and will not be repeated here.
[0224] See also Figure 6 , Figure 6 It is a flowchart of a safety hazard detection method provided in an embodiment of the present application.
[0225] The potential safety hazard detection method is applied to electronic devices, for example: Figure 1 In the electronic device 20 shown, specifically, the execution subject of the potential safety hazard detection method is one or at least two processors of the electronic device.
[0226] like Figure 6 As shown, the potential safety hazard detection method includes steps S601 and S602:
[0227] Step S601: Acquire target multimodal data in a target engineering environment.
[0228] Among them, the target engineering environment is an engineering environment that needs to be detected by the trained safety hazard detection model to see if there are safety hazards, and the target multimodal data is multimodal data that needs to be detected by the trained safety hazard detection model.
[0229] Specifically, target multimodal data is collected by a data collection device deployed in a target engineering environment.
[0230] Among them, the data acquisition equipment includes a camera device, a sound acquisition device, a position acquisition device and an equipment status monitoring device, and the target multimodal data includes target visual data, target audio data, target position relationship data and target equipment status data. The target visual data is the picture or video collected by the camera device and needs to be detected by the trained safety hazard detection model, the target audio data is the sound collected by the sound acquisition device and needs to be detected by the trained safety hazard detection model, the target position relationship data is the position information collected by the position acquisition device and needs to be detected by the trained safety hazard detection model, and the target equipment status data is the status information of the construction equipment collected by the equipment status monitoring device and needs to be detected by the trained safety hazard detection model.
[0231] In the embodiment of the present application, obtaining target multimodal data in the target engineering environment includes: obtaining target visual data collected by a camera device in the target engineering environment; obtaining target audio data collected by a sound collection device in the target engineering environment; obtaining target position relationship data collected by a position collection device in the target engineering environment; obtaining target device status data collected by a device status monitoring device in the target engineering environment. The specific implementation method of obtaining target multimodal data in the target engineering environment is similar to step S201 and will not be repeated here.
[0232] In the embodiment of the present application, before detecting the target multimodal data based on the potential safety hazard detection model, the potential safety hazard detection method further includes steps S3 and S4:
[0233] Step S3: Preprocessing the target multimodal data to obtain the multimodal data to be detected.
[0234] The multimodal data to be detected is the multimodal data obtained after preprocessing the target multimodal data.
[0235] Specifically, the target multimodal data is cleaned to obtain cleaned target multimodal data. The cleaned target multimodal data is processed according to the type of modal data to obtain the multimodal data to be detected. This step is similar to the specific implementation of step S221 and will not be repeated here.
[0236] Step S4: Acquire a detection instruction, and input the detection instruction and the multimodal data to be detected into a potential safety hazard detection model, so that the potential safety hazard detection model detects the multimodal data to be detected.
[0237] The detection instructions include at least one of checking personnel safety hazards, checking material safety hazards, checking construction process hazards, checking equipment safety hazards, checking space location hazards, and checking environmental safety hazards. The safety hazard detection model is trained by the training method of the safety hazard detection model in any of the above embodiments.
[0238] Specifically, a detection instruction is obtained, and the detection instruction and the multimodal data to be detected are input into a trained potential safety hazard detection model, so that the potential safety hazard detection model detects the multimodal data to be detected according to the detection instruction.
[0239] Step S602: Detect target multimodal data based on the potential safety hazard detection model to determine whether there are potential safety hazards in the target engineering environment.
[0240] Specifically, the safety hazard detection model takes the target multimodal data as input and outputs the corresponding detection results, which include whether there are safety hazards in the target engineering environment.
[0241] In some embodiments, the safety hazard detection model takes the target multimodal data and the detection instruction as input, detects the multimodal data to be detected according to the detection instruction, and outputs the corresponding detection result.
[0242] In some embodiments, the trained safety hazard detection model is deployed on the monitoring equipment of the target engineering environment, so as to detect safety hazard on the multimodal data collected in real time by the data acquisition equipment. The monitoring equipment includes but is not limited to edge computing equipment and field servers.
[0243] When it is determined that the target engineering environment has potential safety hazards, the potential safety hazard detection method further includes steps S5 and S6:
[0244] Step S5: Outputting the safety hazard type based on the safety hazard detection model.
[0245] Among them, the types of safety hazards include at least one of personnel safety hazards, material safety hazards, equipment safety hazards, environmental safety hazards, construction process hazards and spatial location hazards.
[0246] Specifically, when the safety hazard detection model detects that there are safety hazards in the target engineering environment, the corresponding safety hazard type is output.
[0247] Step S6: Perform early warning processing based on the type of potential safety hazard and the preset early warning strategy.
[0248] The preset warning strategy refers to the response and disposal rules formulated according to different types of safety hazards. The preset warning strategy can be set by those skilled in the art according to the types of safety hazards, and is not limited here.
[0249] Specifically, according to the safety hazard type output by the safety hazard detection model, a corresponding preset warning strategy is matched, and a corresponding warning operation is performed according to the preset warning strategy.
[0250] In some embodiments, the electronic device is connected to at least one communication device, and the communication device is an electronic device used by staff in the target engineering environment. Early warning processing is performed based on the type of safety hazard and the preset early warning strategy, including: when the type of safety hazard is a personnel safety hazard or an equipment safety hazard, an alarm notification is sent to the communication device. The alarm notification is used to notify the staff to take timely measures to deal with the safety hazard, and the alarm notification includes the type of safety hazard, the location of the safety hazard, and related target multimodal data.
[0251] In some embodiments, the electronic device is in communication with a display screen, and the display screen is located in a target engineering environment. Early warning processing is performed based on the type of safety hazard and a preset early warning strategy, including: when the type of safety hazard is a material safety hazard or a construction process hazard, a warning message is sent to the display screen so that the display screen displays the warning message. The warning information includes the type of safety hazard, the location of the safety hazard, and related target multimodal data.
[0252] In some embodiments, the electronic device is communicatively connected to a speaker, and the speaker is located in the target engineering environment. Early warning processing is performed based on the type of potential safety hazard and the preset early warning strategy, including: when the type of potential safety hazard is an environmental safety hazard or a spatial location hazard, controlling the speaker to emit a sound alarm. The sound alarm is used to remind the staff of the target engineering environment to pay attention to the current dangerous situation and evacuate in time or take safety protection measures.
[0253] In some embodiments, the safety hazard detection method further includes: when determining that the target engineering environment has a safety hazard, recording the safety hazard information in a database. The database is used to store the safety hazard information, and the safety hazard information includes but is not limited to the type of safety hazard, the location of the safety hazard, the time when the safety hazard occurs, and related target multimodal data.
[0254] In an embodiment of the present application, a safety hazard detection method is provided, which includes: acquiring target multimodal data in a target engineering environment; detecting the target multimodal data based on a safety hazard detection model to determine whether there are safety hazards in the target engineering environment, wherein the safety hazard detection model is trained by the training method of the safety hazard detection model in any of the above embodiments.
[0255] The target multimodal data in the target engineering environment is detected by a safety hazard detection model. The model is trained by the safety hazard detection model training method in any of the above embodiments. The present application can comprehensively consider various modal data in the engineering environment and improve the accuracy of safety hazard detection in the target engineering environment.
[0256] See also Figure 7 , Figure 7 It is a structural schematic diagram of a safety hazard detection device provided in an embodiment of the present application.
[0257] The potential safety hazard detection device is applied to an electronic device, such as a terminal or a server. Specifically, the potential safety hazard detection device is configured in the electronic device.
[0258] like Figure 7 As shown, the potential safety hazard detection device 700 includes:
[0259] The second data acquisition unit 701 is used to acquire target multimodal data in a target engineering environment.
[0260] The data detection unit 702 is used to detect the target multimodal data based on the potential safety hazard detection model to determine whether there are potential safety hazards in the target engineering environment.
[0261] In some embodiments of the present application, the second data acquisition unit 701 is also used to: perform preprocessing operations on the target multimodal data to obtain multimodal data to be detected; obtain detection instructions, and input the detection instructions and the multimodal data to be detected into the safety hazard detection model, so that the safety hazard detection model detects the multimodal data to be detected.
[0262] In some embodiments of the present application, when it is determined that there are safety hazards in the target engineering environment, the data detection unit 702 is also used to: output the safety hazard type based on the safety hazard detection model; and perform warning processing based on the safety hazard type and a preset warning strategy.
[0263] It can be understood that the implementation principle and technical effects of the safety hazard detection device 700 proposed in the present application can refer to the implementation principle and technical effects of the steps of the safety hazard detection method proposed in the present application, and will not be repeated here.
[0264] See also Figure 8 , Figure 8 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application.
[0265] like Figure 8 As shown, the electronic device 20 includes one or more processors 21 and a memory 22, and the electronic device 20 may be a terminal or a server. Figure 8 A processor 21 is taken as an example.
[0266] The processor 21 and the memory 22 may be connected via a bus or other means. Figure 8 The example of connecting through bus is taken in the following.
[0267] The processor 21 is used to provide computing and control capabilities to control the electronic device 20 to perform corresponding tasks, for example, to control the electronic device 20 to execute the training method of the safety hazard detection model in any of the above method embodiments, or to control the electronic device 20 to execute the safety hazard detection method in any of the above method embodiments.
[0268] The processor 21 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip or any combination thereof; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0269] The memory 22, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the training method of the safety hazard detection model or the program instructions / modules corresponding to the safety hazard detection method in the embodiment of the present application. The processor 21 can implement the training method of the safety hazard detection model or the safety hazard detection method in any of the above method embodiments by running the non-transitory software programs, instructions and modules stored in the memory 22. Specifically, the memory 22 may include a volatile memory (VM), such as a random access memory (RAM); the memory 22 may also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD) or other non-transitory solid-state storage device; the memory 22 may also include a combination of the above types of memories.
[0270] The memory 22 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 may optionally include a memory remotely arranged relative to the processor 21, and these remote memories may be connected to the processor 21 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0271] One or more modules are stored in the memory 22, and when executed by one or more processors 21, the training method of the safety hazard detection model or the safety hazard detection method in any of the above method embodiments is executed, for example, the above described Figure 2 or Figure 6 The steps shown can also be implemented Figure 5 or Figure 7 The functions of each module or unit.
[0272] In the embodiment of the present application, the electronic device 20 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The electronic device 20 may also include other components for realizing device functions, which will not be described in detail here.
[0273] The present application also provides a computer-readable storage medium, such as a memory including a control program code of an electronic device, and the control program code of the electronic device can be executed by a processor to complete the training method of the safety hazard detection model or the safety hazard detection method in the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, and an optical data storage device.
[0274] The present application also provides a computer program product, which includes one or more program codes stored in a computer-readable storage medium. The processor of the electronic device reads the program code from the computer-readable storage medium, and the processor executes the program code to complete the training method of the safety hazard detection model or the method steps of the safety hazard detection method provided in the above embodiment.
[0275] A person skilled in the art will appreciate that all or part of the steps for implementing the above embodiments may be accomplished by hardware or by hardware associated with a program code, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0276] Through the description of the above implementation methods, ordinary technicians in this field can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Ordinary technicians in this field can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0277] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes in different aspects of the present application as mentioned above, which are not provided in detail for the sake of simplicity. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A training method for a safety hazard detection model, characterized in that: include: Acquire first multimodal data of engineering environments; Performing data association processing on the first multimodal data to construct an instruction data set; The safety hazard detection model is trained based on the instruction data set.
2. The method according to claim 1, characterized in that The first multimodal data includes first visual data, first audio data, first position relationship data, and first device status data; Get first-of-its-kind multimodal data of your engineering environment, including: Acquire first visual data collected by a camera device in an engineering environment, wherein the first visual data includes at least one of individual protection information, working environment information, and worker behavior information; Acquire first audio data collected by a sound collection device in an engineering environment, wherein the first audio data includes at least one of equipment operation sound, gas release sound, and liquid flow sound; Acquire first position relationship data collected by a position collection device in an engineering environment, wherein the first position relationship data includes at least one of worker position information and construction equipment position information; First equipment status data collected by an equipment status monitoring device in an engineering environment is acquired, wherein the first equipment status data includes operation status information of the construction equipment.
3. The method according to claim 1, characterized in that Performing data association processing on the first multimodal data to construct an instruction data set includes: Performing a preprocessing operation on the first multimodal data to obtain second multimodal data; Performing time synchronization and spatial registration operations on each modality data in the second multimodal data to obtain multimodal spatiotemporal correlation data, wherein the spatiotemporal correlation data includes second visual data, second audio data, second position relationship data, and second device status data; Annotating the spatiotemporal correlation data to obtain a first data set; An instruction dataset is constructed based on the first dataset and the first language model.
4. The method according to claim 3, characterized in that The spatiotemporal correlation data is labeled, including: Annotating visual information and time sequence information in the second visual data, wherein the visual information includes a potential safety hazard object and location and status information of the potential safety hazard object; extracting and annotating audio information in the second audio data, wherein the audio information includes abnormal sounds; Calculate, based on the second position relationship data, a first distance between the worker's position and the first preset area, a second distance between the construction equipment position and the second preset area, and a third distance between different construction equipment positions, and mark the first distance, the second distance, and the third distance; Annotating semantic information in the second device status data, wherein the semantic information includes device operating speed and / or device load status; Marking the potential safety hazard types corresponding to the second visual data, the second audio data, the second position relationship data, and the second device status data; The different modal data corresponding to the same safety hazard type, time and space in the spatiotemporal correlation data are associated and labeled.
5. The method according to claim 3, characterized in that: The first data set includes annotated spatiotemporal correlation data, the instruction data set includes instruction data, each of the instruction data includes a detection instruction and a set of annotated spatiotemporal correlation data and a detection result, and the detection result includes a safety hazard type; Constructing an instruction dataset based on the first dataset and the first language model, including: Parsing the annotated spatiotemporal correlation data based on the first language model, and generating a detection instruction in combination with a preset instruction template; Associating the detection instruction with the annotated spatiotemporal correlation data and the detection result to generate instruction data; Among them, the detection instructions include at least one of checking personnel safety hazards, checking material safety hazards, checking construction process hazards, checking equipment safety hazards, checking spatial location hazards and checking environmental safety hazards, and the safety hazard types include at least one of personnel safety hazards, material safety hazards, equipment safety hazards, environmental safety hazards, construction process hazards and spatial location hazards.
6. The method according to claim 1, characterized in that The potential safety hazard detection model includes a multimodal encoder, an interface layer, and a second language model; Before training the potential safety hazard detection model based on the instruction data set, the method further includes: Obtain a pre-trained multimodal encoder and a pre-trained second language model; Keeping parameters of the multimodal encoder and the second language model unchanged, training the potential safety hazard detection model based on a second data set to update parameters of the interface layer, wherein the second data set is different from the instruction data set; Training the potential safety hazard detection model based on the instruction data set includes: The parameters of the multimodal encoder are kept unchanged, and the potential safety hazard detection model is trained based on the instruction data set to update the parameters of the interface layer and the second language model.
7. The method according to claim 6, characterized in that The multimodal encoder includes a text encoder, a visual encoder, and an audio encoder; The text encoder is used to extract features from the second position relationship data and the second device status data to obtain a first text feature vector; The visual encoder is used to extract features from the second visual data to obtain a visual feature vector; The audio encoder is used to extract features from the second audio data to obtain an audio feature vector; The interface layer is used to align the feature space of the visual feature vector and the audio feature vector with the text feature space of the second language model to obtain a second text feature vector; The second language model is used to perform text feature extraction on the first text feature vector and the second text feature vector based on the detection instruction to detect whether there is a corresponding security risk, and output a corresponding security risk type when a security risk is detected.
8. A method for detecting potential safety hazards, characterized in that: include: Acquire target multimodal data in target engineering environment; The target multimodal data is detected based on a safety hazard detection model to determine whether there are safety hazards in the target engineering environment, wherein the safety hazard detection model is trained by the training method of the safety hazard detection model as described in any one of claims 1 to 7.
9. The method according to claim 8, characterized in that Before detecting the target multimodal data based on the potential safety hazard detection model, the method further includes: Performing a preprocessing operation on the target multimodal data to obtain multimodal data to be detected; Obtaining a detection instruction, and inputting the detection instruction and the multimodal data to be detected into the potential safety hazard detection model, so that the potential safety hazard detection model detects the multimodal data to be detected; When it is determined that the target engineering environment has a potential safety hazard, the method further includes: Outputting the type of safety hazard based on the safety hazard detection model; Perform early warning processing based on the safety hazard type and preset early warning strategy.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the training method of the safety hazard detection model as described in any one of claims 1 to 7, or the steps of the safety hazard detection method as described in claim 8 or 9 are implemented.
Citation Information
Cited By
Industrial safety grading early warning system and method based on multi-type DNN joint reasoning
CN121527980A