A method, device, electronic device, and storage medium for identifying risks and hazards based on a large model.
By using a large-scale model-based risk and hazard identification method, feature extraction and fusion are performed using a large-scale hazard identification model. Combined with sensor signals and virtual image training, this method solves the problem of insufficient accuracy in hazard identification in complex scenarios in existing technologies. It achieves efficient hazard identification and early handling in complex scenarios, thereby improving production safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA NAT BUILDING MATERIALS TECH CO LTD
- Filing Date
- 2025-03-26
- Publication Date
- 2026-05-26
Smart Images

Figure CN120298836B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device and storage medium for identifying risks and hazards based on a large model. Background Technology
[0002] Production safety is directly related to the life safety of employees and the stable development of enterprises. Only under the premise of safety can production activities proceed smoothly and avoid personal injury and property loss caused by accidents.
[0003] In the production process, timely detection and elimination of potential risks are crucial for preventing accidents. Currently, AI-powered identification methods can be used to recognize common risks and improve the production safety level of enterprises.
[0004] In the process of realizing this disclosure, it was found that the prior art has at least the following technical problems: the current AI model has limited recognition accuracy and cannot effectively recognize complex and new scenarios, resulting in poor recognition performance. Summary of the Invention
[0005] This disclosure provides a risk and hazard identification method, device, electronic device, and storage medium based on a large model to improve the identification effect of hazards.
[0006] According to one aspect of this disclosure, a risk hazard identification method based on a large model is provided, including:
[0007] Acquire a detection image, the detection image including an image at at least one acquisition angle;
[0008] The detected image is processed based on the trained hazard identification model to obtain the hazard identification result corresponding to the detected image. The hazard identification includes one or more of the hazard type and hazard degree.
[0009] The large-scale hazard identification model includes a first encoder and a decoder;
[0010] The image at the at least one acquisition angle is input into the first encoder for feature extraction to obtain global features and multi-scale local features at each acquisition angle; the global features and multi-scale local features at the at least one angle are input into the decoder for feature fusion to obtain the hazard identification result corresponding to the detected image;
[0011] The hazard identification model is trained based on sample images from multiple scenarios, including historical images collected in real-world scenarios and virtual images in virtual scenarios; the virtual scenarios are newly added scenarios relative to the real-world scenarios.
[0012] Optionally, the large-scale hazard identification model may also include a second encoder;
[0013] The method further includes: acquiring sensor acquisition signals, the sensor acquisition signals including one or more of audio signals, odor signals, temperature signals, vibration signals and pressure signals; transmitting the sensor acquisition signals to the second encoder for feature extraction to obtain auxiliary features; and inputting the auxiliary features, the global features at at least one angle and the multi-scale local features to the decoder for feature fusion to obtain the hazard identification result corresponding to the detection image.
[0014] Optionally, the method for generating virtual images in the virtual scene includes one or more of the following:
[0015] Obtain scene description text and hazard description text, wherein the hazard description text includes one or more of the hazard type and hazard degree; generate a virtual image in the virtual scene based on the scene description text and the hazard description text using a first image generation model, wherein the virtual image includes an image corresponding to at least one angle;
[0016] Obtain scene description text and historically acquired images; generate virtual images in a virtual scene based on the scene description text and historically acquired images using a second image generation model. The virtual images correspond to different scenes than the historically acquired images, and the virtual images and historically acquired images have the same hazard labels.
[0017] Optionally, the virtual scene includes a composite scene; correspondingly, the virtual image includes a virtual image within the composite scene;
[0018] The virtual image in the composite scene is generated by the first image generation model based on the composite scene description text and the hazard description text; or, it is generated by the second image generation model based on the composite scene description text and the historically acquired image.
[0019] Optionally, the training method for the large-scale hazard identification model includes:
[0020] Obtain an initial training set, and pre-train the large-scale hazard identification model to be trained based on the initial training set to obtain the initial large-scale hazard identification model.
[0021] Based on sample images from the various scenarios, the initial hazard identification model is subjected to transfer learning to obtain a trained hazard identification model.
[0022] Optionally, the application process of the large-scale hazard identification model also includes multiple incremental update stages;
[0023] The method further includes: during the inference process of the hazard identification big model, determining the similarity data between the detected image and the sample image corresponding to the hazard identification big model; when the similarity data meets the sample screening conditions, determining the detected image as an incremental sample image; and / or, obtaining the scene description text of the newly added scene and generating the incremental sample image corresponding to the newly added scene.
[0024] In any of the incremental update stages, the trained hazard identification model is optimized and trained based on the incremental sample images to obtain an updated hazard identification model; wherein, the model parameters of the updated hazard identification model are expressed as follows: Where θ represents the model parameters of the large-scale hazard identification model before the incremental update phase; θ * The model parameters for the updated hazard identification model; Let D be the loss function. new Let x be the incremental sample image, y be the hazard label in the incremental sample image.
[0025] Optionally, the method further includes: creating a hazard display item based on the hazard identification result; displaying the hazard display item and its warning information on a display page, wherein the warning information is set according to the hazard degree and / or hazard type in the hazard display item.
[0026] According to another aspect of this disclosure, a risk hazard identification device based on a large model is provided, comprising:
[0027] The image acquisition module is used to acquire a detection image, which includes an image at at least one acquisition angle;
[0028] The hazard identification module is used to process the detected image based on a trained hazard identification model to obtain the hazard identification result corresponding to the detected image. The hazard identification includes one or more of the hazard type and hazard degree.
[0029] The large-scale hazard identification model includes a first encoder and a decoder;
[0030] The hazard identification module is used to input the image at the at least one acquisition angle into the first encoder for feature extraction, to obtain global features and multi-scale local features at each acquisition angle; and to input the global features and multi-scale local features at the at least one angle into the decoder for feature fusion, to obtain the hazard identification result corresponding to the detected image.
[0031] The hazard identification model is trained based on sample images from multiple scenarios, including historical images collected in real-world scenarios and virtual images in virtual scenarios; the virtual scenarios are newly added scenarios relative to the real-world scenarios.
[0032] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising:
[0033] At least one processor; and
[0034] A memory communicatively connected to the at least one processor; wherein,
[0035] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the risk identification method based on a large model as described in any embodiment of this disclosure.
[0036] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute and implement the risk identification method based on a large model as described in any embodiment of this disclosure.
[0037] The technical solution of this disclosure utilizes an end-to-end hazard identification model to extract and fuse features from detection images at multiple angles and scales, improving the comprehensiveness of features and enabling hazard identification in complex scenarios, thereby enhancing the accuracy of hazard identification results. The hazard identification results obtained through the large-scale hazard identification model include the hazard type and severity, displaying the identification results and indicating the hazard type and severity. This facilitates early detection and handling of hazards in the production process, improving production safety.
[0038] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a flowchart of a risk and hazard identification method based on a large model provided in an embodiment of this disclosure;
[0041] Figure 2 This is a flowchart of a training method for a large-scale hazard identification model provided in an embodiment of this disclosure;
[0042] Figure 3 This is a flowchart of a risk and hazard identification method based on a large model provided in an embodiment of this disclosure;
[0043] Figure 4 This is a structural diagram of a large-scale hazard identification model according to an embodiment of the present invention;
[0044] Figure 5 This is a schematic diagram of the structure of a risk and hazard identification device based on a large model provided in an embodiment of this disclosure;
[0045] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0046] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0047] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0048] Figure 1This is a flowchart of a risk hazard identification method based on a large model, provided in an embodiment of this disclosure. This embodiment is applicable to situations where risk hazard detection is performed on inspection images using a large model. The method can be executed by a risk hazard identification device based on a large model, which can be implemented in hardware and / or software. This device can be configured in inspection robots, mobile terminals, servers, and computers, wherein the mobile terminal may include, but is not limited to, mobile phones and tablets. Figure 1 As shown, the method includes:
[0049] S110. Acquire a detection image, wherein the detection image includes an image at at least one acquisition angle.
[0050] In this embodiment, risk and hazard detection is performed on the environment to be tested. This environment may include, but is not limited to, a factory environment. The factory environment can include multiple detection scenarios, such as, but not limited to, equipment operation scenarios, material placement scenarios, personnel work scenarios, and vehicle transportation scenarios. Detection images are acquired for multiple detection scenarios within the environment to be tested. These images are then used to determine whether any safety hazards exist in the environment, enabling early detection and handling of hazards and improving production safety.
[0051] Optionally, inspection personnel can conduct inspections using handheld devices, capturing images of various inspection scenarios during the process. These handheld devices can be electronic devices with image capture capabilities, such as mobile phones, cameras, or tablets. Images or videos captured by the handheld devices are used as inspection images, or a set number of images are extracted from the video as inspection images. For example, the handheld device can capture images of the inspection scenario from multiple angles, obtaining images from at least one capture angle; or, the handheld device can capture videos of the inspection scenario from multiple angles, extracting images from at least one capture angle. For example, each inspection scenario may include one or more inspection points, and for each inspection point, images from at least one capture angle are captured; the acquisition method is not described in detail here. It is understood that the handheld device can execute the risk and hazard identification method based on a large model; alternatively, the handheld device can transmit the captured images or videos to the electronic device in this embodiment to execute the risk and hazard identification method based on a large model.
[0052] Optionally, an inspection robot inspects the environment to be inspected and acquires inspection images at each inspection scene. The inspection robot is pre-configured with an inspection path and inspects the environment according to this path. The inspection path includes multiple inspection points, and each inspection scene may include one or more inspection points. The inspection robot acquires images or videos at each inspection point to determine the inspection image for each point. Similarly, the inspection robot acquires images from multiple angles at each inspection point to obtain images from at least one angle. It is understood that the inspection robot may be equipped with a processor that can execute the large-model-based risk identification method provided in this embodiment; or, the inspection robot may be equipped with a communication component that communicates with the electronic device in this embodiment to transmit the acquired images or videos to the electronic device for executing the large-model-based risk identification method.
[0053] Optionally, cameras are installed in each detection scenario of the environment to be inspected. These cameras can be connected to the electronic devices in this embodiment via communication. For example, the camera receives acquisition signals transmitted by the electronic device, acquires images or videos in the detection scenario, and transmits them to the electronic device to obtain a detection image. Alternatively, the camera periodically acquires images or videos and transmits them to the electronic device to obtain a detection image. Or, the camera acquires images or videos in the detection scenario in real time and transmits them to the electronic device, which then determines the detection image from the images or videos of each detection environment. One or more cameras can be installed in each detection environment to improve the comprehensiveness of the inspection. Each camera can acquire data from multiple angles, improving the comprehensiveness of the image acquisition angles and further enhancing the comprehensiveness of hazard detection.
[0054] In this embodiment, by acquiring images from at least one angle, multi-angle risk and hazard detection can be achieved, thereby improving the comprehensiveness and accuracy of the detection.
[0055] S120. The detected image is processed based on the trained hazard identification model to obtain the hazard identification result corresponding to the detected image. The hazard identification includes one or more of the hazard type and hazard degree.
[0056] In this embodiment, an end-to-end large-scale hazard identification model is used to detect hazards in the detection images, thereby improving detection efficiency. Optionally, the detection image of any detection scene (i.e., an image of the detection scene from at least one angle) is input into the aforementioned trained large-scale hazard identification model to obtain the hazard identification result for that detection scene. Optionally, the detection image of any detection point in the detection scene (i.e., an image of the detection point from at least one angle) is input into the aforementioned trained large-scale hazard identification model to obtain the hazard identification result for that detection point.
[0057] The type and specific structure of the large-scale hazard identification model are not limited here, and it can be pre-built during the model training phase. The large-scale hazard identification model can be a neural network model, such as, but not limited to, the transformer model.
[0058] Optionally, the large-scale hazard identification model includes a first encoder and a decoder; the first encoder and decoder may each include multiple network layers, such as one or more of convolutional layers, pooling layers, normalization layers, and activation function layers. Optionally, the first encoder and / or decoder may include residual network blocks.
[0059] The first encoder is used to extract image features from the input image. Images from at least one acquisition angle are input to the first encoder for feature extraction, obtaining global features and multi-scale local features for each acquisition angle. Global features represent the overall features of the input image, while local features represent the features of local regions within the input image. By extracting global and local features from images at each acquisition angle, the comprehensiveness of the features is improved, facilitating the extraction of complex semantic information from the image. Furthermore, extracting multi-scale local features effectively captures features and semantic information in the input image, providing rich feature information for hazard identification and improving the accuracy of hazard identification.
[0060] The decoder is used to fuse the extracted features and obtain the hazard identification result. The global features at at least one angle and the local features at multiple scales are input into the decoder for feature fusion to obtain the hazard identification result corresponding to the detected image. Specifically, the global features at at least one angle and the local features at multiple scales are fused to obtain the target fused feature, and the hazard identification result is identified based on the fused feature. The global features corresponding to at least one acquisition angle are fused to obtain the global fused feature, and the local features at any scale corresponding to at least one acquisition angle are fused to obtain the local fused feature at that scale. The local fused features at each scale are upsampled to achieve the fusion of the global fused feature and the local fused features at each scale, obtaining the target fused feature. For example, the local fused feature at the first scale is upsampled, and the upsampled local fused feature is fused with the local fused feature at the second scale to obtain the second-scale fused feature. The second-scale fused feature is upsampled, and the upsampled fused feature is fused with the local fused feature at the third scale, and so on, until the local fused features at each scale and the global fused feature are fused to obtain the target fused feature. The first, second, and third scales increase sequentially.
[0061] By identifying the target fusion features, the hazard identification result is obtained. By fusing multi-scale image features corresponding to at least one acquisition angle, the extraction and fusion of multi-angle and multi-scale features are realized, enriching the feature information of the image. This enables hazard identification in complex scenes and provides a basis for the accuracy of hazard identification results.
[0062] The hazard identification results include hazard type and hazard severity. In this embodiment, the large-scale hazard identification model can identify multiple hazard types with high integration. The detected image can correspond to one or more hazard types. The large-scale hazard identification model can obtain the confidence level corresponding to each hazard type. If the confidence level corresponding to any hazard type is greater than a confidence threshold, it outputs the hazard type and its severity. This severity level can be understood as the urgency of the hazard type, or the probability that the hazard of that type will transform into an accident. The higher the severity level, the greater the probability of an accident occurring.
[0063] The technical solution in this embodiment uses an end-to-end hazard identification model to extract and fuse features from the detected images from multiple angles and scales, improving the comprehensiveness of the features and enabling hazard identification in complex scenarios, thus improving the accuracy of the hazard identification results. The hazard identification results obtained through the hazard identification model include the hazard type and hazard severity, displaying the hazard identification results and indicating the hazard type and severity, facilitating early detection and handling of hazards in the production process, and improving production safety.
[0064] Based on the above embodiments, the large-scale hazard identification model is pre-trained and optional. The large-scale hazard identification model is trained based on sample images from multiple scenarios, including historical images collected in actual scenarios and virtual images in virtual scenarios; the virtual scenarios are newly added scenarios relative to the actual scenarios.
[0065] In this context, the actual scene can be a real scene within the environment to be detected, and there can be one or more actual scenes. Historically acquired images can be detection images of actual scenes within the environment to be detected over a historical period. The virtual scene is a different scene from the actual scene; it is a newly added scene relative to the actual scene. The virtual scene supplements the actual scene by obtaining virtual images within the virtual scene, thus enriching the sample image set and increasing the scene diversity of the sample images. The large-scale hazard identification model trained with diverse sample images can identify hazards in multiple scenes, improving its adaptability to new scenes and avoiding the limitation of scene constraints that could affect the robustness and generalization of the large-scale hazard identification model in different scenarios.
[0066] Setting up virtual images within a virtual scene simplifies the sample image acquisition process, avoids the problem of not being able to comprehensively acquire sample images from different scenes, and reduces the amount of sample data to be acquired, thus lowering the difficulty of acquiring sample images.
[0067] Optionally, the method for generating virtual images in a virtual scene may include: obtaining scene description text and hazard description text, wherein the hazard description text includes one or more of the hazard type and hazard degree; generating virtual images in a virtual scene based on the scene description text and the hazard description text using a first image generation model, wherein the virtual images include images corresponding to at least one angle.
[0068] The scene description text can be understood as text describing the virtual scene; the hazard description text can be understood as text describing the hazards existing in the virtual scene. Both the scene description text and the hazard description text can be collected through an interactive page and input by the operator. The scene description text and the hazard description text are then input into a pre-trained first image generation model to obtain a virtual image of the virtual scene. This first image generation model can be a neural network model with image generation capabilities, trained through generative adversarial training using a generator in a generative adversarial network.
[0069] The first image generation model can generate an image corresponding to at least one angle in a virtual scene based on scene description text and hazard description text, facilitating the training of the large-scale hazard identification model with multi-angle images. Based on this, hazard labels corresponding to the virtual images are set to form sample images.
[0070] Optionally, the method for generating virtual images in a virtual scene may include: acquiring scene description text and historically acquired images; generating virtual images in a virtual scene based on the scene description text and the historically acquired images using a second image generation model, wherein the virtual images and the historically acquired images correspond to different scenes, and the virtual images and the historically acquired images correspond to the same hazard labels.
[0071] The second image generation model can be a neural network model with image generation capabilities obtained through generative adversarial training, using a generator in a generative adversarial network. By inputting scene description text and historically acquired images into the second image generation model, a virtual image in a virtual scene is obtained. This virtual image can be understood as a scene transformation of the historically acquired images, transforming the actual scene corresponding to the historically acquired images into the virtual scene corresponding to the scene description text. The hazard labels corresponding to the virtual image and the historically acquired images are consistent, eliminating the need to set hazard labels for the virtual image. Instead, the hazard labels corresponding to the historically acquired images are used as the hazard labels for the virtual image, simplifying the hazard label setting process.
[0072] Based on the above embodiments, the virtual scene includes a composite scene, and correspondingly, the virtual image includes a virtual image within the composite scene. A composite scene can be understood as a mixed scene comprising two or more single scenes. For example, a single scene may include, but is not limited to, equipment operation scenes, material placement scenes, personnel work scenes, and vehicle transportation scenes. A composite scene may be, for example, a mixed scene of personnel work scenes and vehicle transportation scenes, a mixed scene of material placement scenes and personnel work scenes, a mixed scene of equipment operation scenes and personnel work scenes, a mixed scene of equipment operation scenes and material placement scenes, etc. It is understood that composite scenes are complex scenes. By setting virtual images within composite scenes, it is convenient to train the large-scale hazard identification model for hazard identification in complex scenes, enabling accurate hazard identification of detected images in complex scenes through the large-scale hazard identification model, thereby improving the hazard identification accuracy and adaptability of the large-scale hazard identification model to complex scenes.
[0073] The virtual image in the composite scene is generated using the first image generation model based on the composite scene description text and the hazard description text; or, it is generated using the second image generation model based on the composite scene description text and the historically acquired images. During the virtual image generation process, by transforming the scene description text into composite scene description text, which can immediately serve as a description of the composite scene, a virtual image corresponding to the composite scene description text is generated based on either the first or second image generation model, thus realizing the generation of the virtual image in the composite scene.
[0074] It is understandable that virtual scenes may include one or more detection scenes, and real scenes may include one or more detection scenes. Correspondingly, the sample image set includes sample images of multiple scenes. A large-scale hazard identification model is obtained by training sample images of multiple scenes, which can realize hazard identification in multiple scenes and hazard identification in changing scenes, so that the large-scale hazard identification model can effectively identify hazards in new or complex scenes.
[0075] Based on the above embodiments, the training process of the large-scale hazard identification model includes an initial training phase and multiple incremental update phases. The initial training phase can be understood as the phase of training the large-scale hazard identification model to obtain a trained large-scale hazard identification model. The incremental update phase can be understood as the phase of optimizing the trained large-scale hazard identification model during the inference phase. It can be understood that the large-scale hazard identification model can be optimized multiple times, with each optimization training serving as an incremental update phase.
[0076] For example, Figure 2This is a flowchart illustrating a training method for a large-scale hazard identification model provided in this disclosure. In the initial training phase, the training method for the large-scale hazard identification model includes:
[0077] S210. Obtain an initial training set, and pre-train the large-scale hazard identification model to be trained based on the initial training set to obtain the initial large-scale hazard identification model.
[0078] S220. Based on the sample images of the various scenarios, perform transfer learning on the initial hazard identification model to obtain a trained hazard identification model.
[0079] In this embodiment, a large-scale hazard identification model is trained using a transfer learning mechanism. The initial training set can be understood as a general sample image training set. This initial training set is used to pre-train the large-scale hazard identification model to be trained, resulting in an initial large-scale hazard identification model with general feature extraction and processing capabilities.
[0080] Based on the initial large-scale hazard identification model, sample images from various scenarios are used as a specific set of sample images corresponding to the environment to be detected. The initial large-scale hazard identification model is then fine-tuned to obtain a large-scale hazard identification model that is adapted to various scenarios in the environment to be detected. Through pre-training, the number of sample images required in the specific sample image set can be reduced, the difficulty of collecting sample images can be simplified, and the dependence on new labeled data can be reduced, thereby improving the adaptability to various scenarios.
[0081] In the training process described above, a loss function is generated from the hazard identification results and corresponding hazard labels of the sample images using the initial hazard identification model. The parameters of the initial hazard identification model are then adjusted using this loss function. This training process is repeated until a well-trained hazard identification model is obtained. The hazard labels include hazard type and hazard severity labels. A first loss term is generated using the hazard type and its label from the hazard identification results, and a second loss term is generated using the hazard severity and its label. A loss function is then generated based on both the first and second loss terms. These first and second loss terms can be obtained using the cross-entropy function.
[0082] Based on the above embodiments, the application process of the large-scale hazard identification model also includes multiple incremental update stages; in any of the incremental update stages, incremental sample images are acquired, and the trained large-scale hazard identification model is optimized and trained based on the incremental sample images to obtain the updated large-scale hazard identification model.
[0083] In this context, the incremental sample image can be understood as a newly added sample image based on the sample image before the incremental update phase. This incremental sample image is different from the sample image before the incremental update phase.
[0084] Optionally, the incremental sample image can be determined as follows: during the inference process of the hazard identification model, the similarity data between the detected image and the sample image corresponding to the hazard identification model is determined, and when the similarity data meets the sample screening conditions, the detected image is determined as the incremental sample image.
[0085] The sample images corresponding to the large-scale hazard identification model can be understood as the sample images used in the initial training phase or the completed incremental update phase. The sample selection criteria can be understood as the selection criteria based on similarity data, used to select incremental sample images that are different from the sample images that have already been used.
[0086] A vector database is pre-created, storing semantic vectors corresponding to the used sample images. These semantic vectors can be extracted using a semantic extraction model. Sample images used in the initial training phase or the completed incremental update phase are converted into semantic vectors and stored in the vector database.
[0087] In the reasoning process of the large-scale hazard identification model, detection images are acquired, and the semantic vectors of the detection images are matched with semantic vectors in the vector database to obtain similarity data. The methods for determining the similarity data include, but are not limited to, cosine similarity calculation.
[0088] The sample selection criterion can be that the maximum similarity data is less than a preset threshold. The maximum similarity data between the semantic vector of the detected image and the semantic vectors in the vector database is determined. If the maximum similarity data is less than the preset threshold, it indicates that the detected image differs significantly from the used sample images, and the detected image can be used as an incremental sample image. If the maximum similarity data is greater than or equal to the preset threshold, it indicates that the detected image is similar to one or more of the used sample images, and it is unnecessary to use the detected image as an incremental sample image.
[0089] By detecting the similarity data between the image and the sample image corresponding to the hazard identification model, and by selecting incremental sample images that differ significantly from the already used sample images through sample selection conditions, the ineffective training of the hazard identification model with repeated sample images is avoided, which leads to a waste of training resources and improves the effectiveness of the optimization training of the hazard identification model in the incremental update stage.
[0090] Optionally, the incremental sample image can also be determined by: obtaining the scene description text of the newly added scene and generating the incremental sample image corresponding to the newly added scene. Specifically, the incremental sample image is obtained by processing the scene description text and hazard description text of the newly added scene using a first image generation model; or the incremental sample image is obtained by processing the scene description text of the newly added scene and the used sample image using a second image generation model.
[0091] In the incremental update phase, the model parameters of the large-scale hazard identification model are updated based on the incremental sample images. The updated model parameters of the large-scale hazard identification model are expressed as follows: Where θ represents the model parameters of the large-scale hazard identification model before the incremental update phase; θ * The model parameters for the updated hazard identification model; Let D be the loss function. new Let x be the incremental sample image, and y be the hazard label in the incremental sample image. The specific type of loss function is not limited here.
[0092] The technical solution provided in this embodiment updates the large-scale hazard identification model based on the incremental sample images corresponding to the new scenarios when new scenarios appear or are about to appear in the environment to be detected, thereby achieving the adaptability of the large-scale hazard identification model to the new scenarios and effective hazard identification.
[0093] Figure 3 This is a flowchart of a risk and hazard identification method based on a large model provided in this disclosure embodiment. It is an optimization based on the above embodiment, and the method specifically includes:
[0094] S310. Acquire a detection image, wherein the detection image includes an image at at least one acquisition angle.
[0095] S320. Input the image at the at least one acquisition angle into the first encoder in the hazard identification big data model for feature extraction to obtain global features and multi-scale local features at each acquisition angle.
[0096] S330. Acquire sensor acquisition signals, which include one or more of audio signals, odor signals, temperature signals, vibration signals, and pressure signals; transmit the sensor acquisition signals to the second encoder in the hazard identification big data model for feature extraction to obtain auxiliary features.
[0097] S340. The auxiliary features, the global features at at least one angle, and the multi-scale local features are input into the decoder in the hazard identification model for feature fusion to obtain the hazard identification result corresponding to the detected image.
[0098] It is understandable that the environment to be tested includes different types of hazards, and each hazard may be accompanied by other auxiliary judgment information. For example, fire hazards may be accompanied by auxiliary judgment information such as smell and temperature, and electrical hazards may be accompanied by auxiliary judgment information such as smell and sound.
[0099] Sensors are set up in various detection scenarios within the environment to be inspected to collect sensor signals from these scenarios. Optionally, different sensors can be set up for different detection scenarios; for example, sensors can be set up based on the frequency of occurrence of each type of hazard in different detection scenarios. Specifically, different sensors correspond to different types of hazard.
[0100] In each detection scenario, the sensors are connected to the electronic devices in this embodiment to transmit the signals collected by the sensors to the electronic devices in order to execute the risk and hazard identification method based on the large model in this embodiment.
[0101] In this embodiment, the large-scale hazard identification model includes a first encoder, a second encoder, and a decoder, wherein the first encoder and the second encoder are respectively connected to the decoder. For example, see [link to example]. Figure 4 , Figure 4 This is a structural diagram of a large-scale hazard identification model according to an embodiment of the present invention.
[0102] By using the detection images and sensor signals corresponding to the detection scene, the potential hazards in the detection scene can be identified. Alternatively, by using the detection images and sensor signals of any detection point in the detection scene, the potential hazards at that detection point can be identified.
[0103] The second encoder is used to extract auxiliary features corresponding to the sensor-acquired signals. The decoder then fuses the auxiliary features with the global features and multi-scale local features corresponding to the detected image to obtain the hazard identification results, thereby improving the comprehensiveness of the features and further enhancing the accuracy of the hazard identification results.
[0104] Based on the above embodiments, a hazard display item is created according to the hazard identification result; the hazard display item and the warning information of the hazard display item are displayed on the display page, and the warning information is set according to the hazard degree and / or hazard type in the hazard display item.
[0105] If the hazard type in the hazard identification result is not empty, a hazard display item is created. This hazard display item is used to display each hazard in the detection image. If the hazard identification result includes two or more hazard types, a hazard display item is created for each hazard type.
[0106] The display page shows potential hazards, which may include one or more of the following: hazard type, hazard severity, detection time, hazard location, and detection image. Displaying these hazards allows relevant personnel to address identified hazards promptly, preventing accidents from occurring.
[0107] The display page also shows warning information for the hazard items. This warning information comes in different formats; for example, different levels of hazard severity correspond to different warning messages, or different hazard types correspond to different warning messages. For instance, the warning information can be in graphical form, such as an exclamation mark image, with different colors corresponding to different hazard severity levels.
[0108] Optionally, a hazard handling work order can be generated for each hazard displayed item, and the work order can be transmitted to the associated device of the handling personnel so that the personnel can handle the hazard handling work order and achieve the effect of eliminating the hazard. Different hazard handling work orders can be sorted according to the priority of hazard type or hazard severity to facilitate the timely handling of urgent hazards.
[0109] The technical solution of this embodiment provides auxiliary features for the hazard identification process of the detected image by acquiring signals from the sensor, thereby improving the accuracy of hazard detection.
[0110] Figure 5 This is a schematic diagram of the structure of a risk and hazard identification device based on a large model provided in an embodiment of this disclosure. Figure 5 As shown, the device includes:
[0111] Image acquisition module 410 is used to acquire a detection image, the detection image including an image at at least one acquisition angle;
[0112] The hazard identification module 420 is used to process the detected image based on a trained hazard identification model to obtain the hazard identification result corresponding to the detected image. The hazard identification includes one or more of the hazard type and hazard degree.
[0113] Optionally, the large-scale hazard identification model includes a first encoder and a decoder;
[0114] The hazard identification module 420 is used to input the image at the at least one acquisition angle into the first encoder for feature extraction, to obtain global features and multi-scale local features at each acquisition angle; and to input the global features and multi-scale local features at the at least one angle into the decoder for feature fusion, to obtain the hazard identification result corresponding to the detected image.
[0115] Optionally, the large-scale hazard identification model is trained based on sample images from multiple scenarios, including historically collected images from actual scenarios and virtual images from virtual scenarios; the virtual scenarios are newly added scenarios relative to the actual scenarios.
[0116] The technical solution in this embodiment uses an end-to-end hazard identification model to extract and fuse features from the detected images from multiple angles and scales, improving the comprehensiveness of the features and the accuracy of the hazard identification results. The hazard identification results identified by the hazard identification model include the hazard type and hazard severity, displaying the hazard identification results and indicating the hazard type and severity, facilitating early detection and handling of hazards in the production process, and improving production safety.
[0117] Optionally, based on the above embodiments, the large-scale hazard identification model may further include a second encoder;
[0118] The device also includes: a sensor signal acquisition module, used to acquire sensor-collected signals, wherein the sensor-collected signals include one or more of audio signals, odor signals, temperature signals, vibration signals, and pressure signals;
[0119] The hazard identification module 420 is also used to: transmit the sensor-acquired signal to the second encoder for feature extraction to obtain auxiliary features; input the auxiliary features, the global features at at least one angle and the multi-scale local features to the decoder for feature fusion to obtain the hazard identification result corresponding to the detected image.
[0120] Optionally, the device further includes a virtual image generation module for acquiring scene description text and hazard description text, wherein the hazard description text includes one or more of the hazard type and hazard degree; and generating a virtual image in a virtual scene based on the scene description text and the hazard description text using a first image generation model, wherein the virtual image includes an image corresponding to at least one angle.
[0121] Optionally, the virtual image generation module is also used to: acquire scene description text and historically acquired images; and generate a virtual image in a virtual scene based on the scene description text and the historically acquired images using a second image generation model, wherein the virtual image and the historically acquired images correspond to different scenes, and the virtual image and the historically acquired images correspond to the same hazard label.
[0122] Optionally, the virtual scene includes a composite scene; correspondingly, the virtual image includes a virtual image within the composite scene;
[0123] The virtual image in the composite scene is generated by the first image generation model based on the composite scene description text and the hazard description text; or, it is generated by the second image generation model based on the composite scene description text and the historically acquired image.
[0124] Optionally, the device further includes: a model training module, used to acquire an initial training set, pre-train the large-scale hazard identification model to be trained based on the initial training set to obtain an initial large-scale hazard identification model; and perform transfer learning on the initial large-scale hazard identification model based on sample images of the various scenarios to obtain a trained large-scale hazard identification model.
[0125] Optionally, the application of the large-scale hazard identification model also includes multiple incremental update stages.
[0126] Optionally, the model training module is further configured to: determine similarity data between the detected image and the sample image corresponding to the hazard identification model during the inference process of the hazard identification model; when the similarity data meets the sample screening conditions, determine the detected image as an incremental sample image; and / or, obtain scene description text of the newly added scene and generate the incremental sample image corresponding to the newly added scene.
[0127] In any of the incremental update stages, the trained hazard identification model is optimized and trained based on the incremental sample images to obtain an updated hazard identification model; wherein, the model parameters of the updated hazard identification model are expressed as follows: Where θ represents the model parameters of the large-scale hazard identification model before the incremental update phase; θ * The model parameters for the updated hazard identification model; Let D be the loss function. new Let x be the incremental sample image, y be the hazard label in the incremental sample image.
[0128] Optionally, the device further includes: a display module, used to create hazard display items based on the hazard identification results; and to display the hazard display items and their warning information on a display page, wherein the warning information is set according to the degree and / or type of hazard in the hazard display items.
[0129] The risk and hazard identification device based on a large model provided in this disclosure can execute the risk and hazard identification method based on a large model provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method.
[0130] Figure 6This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0131] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0132] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0133] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as risk hazard identification methods based on large models.
[0134] In some embodiments, the large-model-based risk identification method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the large-model-based risk identification method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the large-model-based risk identification method by any other suitable means (e.g., by means of firmware).
[0135] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0136] Computer programs used to implement the large-model-based risk identification method of this disclosure can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are performed. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0137] This disclosure also provides a computer-readable storage medium storing computer instructions for causing a processor to execute a risk identification method based on a large model, the method comprising:
[0138] Acquire a detection image, which includes an image from at least one acquisition angle; process the detection image based on a trained hazard identification model to obtain a hazard identification result corresponding to the detection image, wherein the hazard identification includes one or more of the hazard type and hazard degree;
[0139] The hazard identification model includes a first encoder and a decoder. The image at the at least one acquisition angle is input into the first encoder for feature extraction to obtain global features and multi-scale local features at each acquisition angle. The global features and multi-scale local features at the at least one angle are input into the decoder for feature fusion to obtain the hazard identification result corresponding to the detected image.
[0140] The hazard identification model is trained based on sample images from multiple scenarios, including historical images collected in real-world scenarios and virtual images in virtual scenarios; the virtual scenarios are newly added scenarios relative to the real-world scenarios.
[0141] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0143] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0144] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0145] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0146] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A risk and hazard identification method based on a large model, characterized in that, include: Acquire a detection image, the detection image including an image at at least one acquisition angle; The detected image is processed based on the trained hazard identification model to obtain the hazard identification result corresponding to the detected image. The hazard identification includes one or more of the hazard type and hazard degree. The large-scale hazard identification model includes a first encoder, a second encoder, and a decoder. The image at at least one acquisition angle is input to the first encoder for feature extraction to obtain global features and multi-scale local features at each acquisition angle; sensor acquisition signals are acquired, including one or more of audio signals, odor signals, temperature signals, vibration signals, and pressure signals; the sensor acquisition signals are transmitted to the second encoder for feature extraction to obtain auxiliary features; the auxiliary features, the global features at the at least one angle, and the multi-scale local features are input to the decoder for feature fusion to obtain the hazard identification result corresponding to the detected image; The hazard identification model is trained based on sample images from multiple scenarios, including historical images collected in real-world scenarios and virtual images in virtual scenarios; the virtual scenarios are newly added scenarios relative to the real-world scenarios.
2. The method according to claim 1, characterized in that, The method for generating virtual images in the virtual scene includes one or more of the following: Obtain scene description text and hazard description text, wherein the hazard description text includes one or more of the hazard type and hazard degree; generate a virtual image in the virtual scene based on the scene description text and the hazard description text using a first image generation model, wherein the virtual image includes an image corresponding to at least one angle; Obtain scene description text and historically acquired images; generate virtual images in a virtual scene based on the scene description text and historically acquired images using a second image generation model. The virtual images correspond to different scenes than the historically acquired images, and the virtual images and historically acquired images have the same hazard labels.
3. The method according to claim 2, characterized in that, The virtual scene includes a composite scene; correspondingly, the virtual image includes a virtual image within the composite scene. The virtual image in the composite scene is generated by the first image generation model based on the composite scene description text and the hazard description text; or, it is generated by the second image generation model based on the composite scene description text and the historically acquired image.
4. The method according to claim 1, characterized in that, The training methods for the large-scale hazard identification model include: Obtain an initial training set, and pre-train the large-scale hazard identification model to be trained based on the initial training set to obtain the initial large-scale hazard identification model. Based on sample images from the various scenarios, the initial hazard identification model is subjected to transfer learning to obtain a trained hazard identification model.
5. The method according to claim 1 or 4, characterized in that, The application of the aforementioned hazard identification model also includes multiple incremental update stages; The method further includes: during the inference process of the hazard identification big model, determining the similarity data between the detected image and the sample image corresponding to the hazard identification big model; when the similarity data meets the sample screening conditions, determining the detected image as an incremental sample image; and / or, obtaining the scene description text of the newly added scene and generating the incremental sample image corresponding to the newly added scene. In any of the incremental update stages, the trained hazard identification model is optimized and trained based on the incremental sample images to obtain an updated hazard identification model; wherein, the model parameters of the updated hazard identification model are expressed as follows: ,in, The model parameters for the large model of hidden danger identification before the incremental update phase; The model parameters for the updated hazard identification model; For loss function, For the incremental sample image, For the incremental sample image, The hazard labels in the incremental sample images.
6. The method according to claim 1, characterized in that, The method further includes: Create hazard display items based on the hazard identification results; The display page shows the hazard items and their warning information, which are set according to the degree and / or type of hazard in the hazard items.
7. A risk and hazard identification device based on a large model, characterized in that, include: The image acquisition module is used to acquire a detection image, which includes an image at at least one acquisition angle; The hazard identification module is used to process the detected image based on a trained hazard identification model to obtain the hazard identification result corresponding to the detected image. The hazard identification includes one or more of the hazard type and hazard degree. The large-scale hazard identification model includes a first encoder and a decoder; The hazard identification module is used to input the image at the at least one acquisition angle into the first encoder for feature extraction, to obtain global features and multi-scale local features at each acquisition angle; and to input the global features and multi-scale local features at the at least one angle into the decoder for feature fusion, to obtain the hazard identification result corresponding to the detected image. The hazard identification model is trained based on sample images from multiple scenarios, including historical images collected in real-world scenarios and virtual images in virtual scenarios; the virtual scenarios are newly added scenarios relative to the real-world scenarios. The large-scale hazard identification model also includes a second encoder; The device further includes: a sensor signal acquisition module, used to acquire sensor-collected signals, wherein the sensor-collected signals include one or more of audio signals, odor signals, temperature signals, vibration signals, and pressure signals; The hazard identification module is further configured to: transmit the sensor-acquired signal to the second encoder for feature extraction to obtain auxiliary features; input the auxiliary features, the global features at at least one angle, and the multi-scale local features to the decoder for feature fusion to obtain the hazard identification result corresponding to the detected image.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the risk identification method based on a large model as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the risk identification method based on a large model as described in any one of claims 1-6.