Method for constructing an ontology-based knowledge base for dangerous chemicals
By constructing an ontology-based hazardous chemicals knowledge base and combining text and image data, the characteristics of chemicals are comprehensively extracted, solving the problem of the lack of integration of image data in existing technologies. This achieves the comprehensiveness and accuracy of the hazardous chemicals knowledge base and supports efficient querying and management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHONGSHUN INT ENG DESIGN CO LTD
- Filing Date
- 2025-09-07
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies are insufficient to fully capture all the characteristics of hazardous chemicals and fail to effectively integrate image and text data, resulting in biased hazard assessments and incomplete knowledge bases.
An ontology-based method for constructing a hazardous chemicals knowledge base involves acquiring textual information and label image data of hazardous chemicals, using a pre-trained semantic extraction model to analyze semantic triple sets, extract attribute sets and hazard feature sets, and then constructing a hazardous chemicals knowledge base.
It enables comprehensive identification and analysis of the characteristics of hazardous chemicals, improves the accuracy and completeness of the knowledge base, supports efficient retrieval and management, and enhances the intelligence level of query and management.
Smart Images

Figure CN121168587B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of chemical knowledge base construction technology, specifically to an ontology-based method for constructing a hazardous chemical knowledge base. Background Technology
[0002] The production, use, and storage of hazardous chemicals are becoming increasingly common, which has led to a gradual increase in their potential hazards and risks to humans and the environment in daily life. Hazardous chemicals may not only cause dangerous incidents such as fires, explosions, and leaks during use, but may also affect ecosystems and human health through long-term environmental pollution and toxicity accumulation. In order to improve the management of hazardous chemicals, ontology-based knowledge base construction methods have gradually attracted attention in recent years. As a method of expressing and managing knowledge, ontology can systematically and structurally represent various types of information about hazardous chemicals and provide a foundation for intelligent reasoning and knowledge updating, enabling it to accurately describe the basic information, physicochemical properties, and hazard characteristics of chemicals.
[0003] The limitations of existing technologies include at least the following problems: Existing technologies cannot comprehensively acquire all the characteristics of hazardous chemicals. For example, they fail to effectively integrate image data. Many chemical information is richly expressed in label images, but existing technologies have not fully explored and integrated this image data. For instance, chemical label images contain important information about their hazards and usage conditions, but this image information has not been effectively combined with text data, leading to a one-sided chemical hazard assessment. Since label images and illustrations can provide multi-dimensional characteristics of chemicals, existing technologies have failed to integrate them into the knowledge base, which not only limits the depth of hazard assessment but also hinders the accuracy of the knowledge base, thus affecting the completeness of the hazardous chemical knowledge base. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides an ontology-based method for constructing a hazardous chemicals knowledge base, which solves the problem that existing technologies comprehensively acquire all characteristics of hazardous chemicals, thereby affecting the completeness of the hazardous chemicals knowledge base construction.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for constructing a hazardous chemicals knowledge base based on ontology, comprising the following steps: acquiring textual information and label image data of several hazardous chemicals, and inputting them into a pre-trained semantic extraction model to analyze the semantic triplet set of each hazardous chemical; extracting the attribute set of each hazardous chemical based on the semantic triplet set of each hazardous chemical, and analyzing the hazard feature set of each hazardous chemical, including fire and explosion hazard feature values, diffusion and pollution risk feature values, and reaction stability feature values; analyzing the comprehensive hazard feature value of each hazardous chemical based on the hazard feature set of each hazardous chemical; and constructing a hazardous chemicals knowledge base based on the comprehensive hazard feature value and the semantic triplet set of each hazardous chemical.
[0006] Furthermore, the label image data specifically refers to the pixel value and two-dimensional coordinates of each pixel in the label image, and the semantic extraction model includes a text recognition subnetwork, an image recognition subnetwork, and a fusion output subnetwork.
[0007] Furthermore, the specific steps for analyzing the semantic triplet set of each hazardous chemical are as follows: In the text recognition sub-network of the semantic extraction model, the text information of each hazardous chemical is received, and the structural semantic representation vector of each hazardous chemical is analyzed; in the image recognition sub-network of the semantic extraction model, the label image data of each hazardous chemical is received, and the image semantic representation vector of each hazardous chemical is analyzed; in the fusion output sub-network of the semantic extraction model, the structural semantic representation vector and image semantic representation vector of each hazardous chemical are received, and the semantic triplet set of each hazardous chemical is analyzed.
[0008] Furthermore, the text recognition subnetwork includes a semantic encoding layer, an entity recognition layer, and a text output layer. The specific steps for analyzing the high-dimensional semantic vector of each hazardous chemical are as follows: In the semantic encoding layer of the text recognition subnetwork, the text information of each hazardous chemical is encoded; in the entity recognition layer of the text recognition subnetwork, the semantic entity vector of each hazardous chemical is analyzed based on the encoded text information; in the text output layer of the text recognition subnetwork, the structural semantic representation vector of each hazardous chemical is analyzed based on the semantic entity vector of each hazardous chemical.
[0009] Furthermore, the image recognition subnetwork includes a preprocessing layer, a warning recognition layer, and an image output layer. The specific steps for analyzing the image semantic representation vector of each hazardous chemical are as follows: In the preprocessing layer of the image recognition subnetwork, the label image data of each hazardous chemical is preprocessed; in the warning recognition layer of the image recognition subnetwork, the warning semantic entity vector of each hazardous chemical is analyzed based on the preprocessed label image data of each hazardous chemical; in the image output layer of the image recognition subnetwork, the image semantic representation vector of each hazardous chemical is analyzed based on the warning semantic entity vector of each hazardous chemical.
[0010] Furthermore, the fusion output subnetwork includes a filling fusion layer and an output layer. The specific content of the fusion output subnetwork is as follows: In the filling fusion layer of the fusion output subnetwork, the fusion semantic representation vector of each hazardous chemical is analyzed based on the structural semantic representation vector and image semantic representation vector of each hazardous chemical; In the output layer of the fusion output subnetwork, the fusion semantic representation vector of each hazardous chemical is output to obtain the semantic triple set of each hazardous chemical.
[0011] Furthermore, the attribute set includes flash point value, auto-ignition point value, toxicity threshold, vapor diffusion gradient value, conductivity change rate value, gas-liquid interface energy consumption value, thermal radiation absorption value, reactivity value, and gas-liquid interface tension value.
[0012] Furthermore, the specific steps for analyzing the hazard characteristic set of each hazardous chemical are as follows: Based on the flash point and auto-ignition point of each hazardous chemical, analyze the fire and explosion hazard characteristic value of each hazardous chemical; based on the conductivity change rate, vapor diffusion gradient, toxicity threshold, and gas-liquid interface energy consumption value of each hazardous chemical, analyze the diffusion pollution risk characteristic value of each hazardous chemical; based on the thermal radiation absorption value, reactivity value, and gas-liquid interfacial tension value of each hazardous chemical, analyze the reaction stability characteristic value of each hazardous chemical.
[0013] Furthermore, the specific formula for calculating the comprehensive hazard characteristic value of a certain hazardous chemical is as follows:
[0014] ;in, This represents the comprehensive hazard characteristic value of a certain hazardous chemical. This refers to the fire and explosion hazard characteristic value of a certain hazardous chemical. The explosion coefficient is stored in the database. This represents the characteristic value of the diffusion and pollution risk of a certain hazardous chemical. The pollution coefficient is stored in the database. This is a characteristic value for the reactivity of a certain hazardous chemical. The stability coefficients are stored in the database. These are adjustment coefficients stored in the database. The interaction coefficients are stored in the database. These are the collaboration coefficients stored in the database.
[0015] Furthermore, the specific steps for constructing a hazardous chemicals knowledge base are as follows: First, sort the chemicals in descending order based on their comprehensive hazard characteristic values to generate a chemical hazard ranking table. Then, construct the hazardous chemicals knowledge base based on the chemical hazard ranking table and the semantic triplet set for each hazardous chemical.
[0016] The present invention has the following beneficial effects:
[0017] (1) The ontology-based method for constructing a hazardous chemical knowledge base combines textual information of hazardous chemicals with label image data to comprehensively extract their semantic information, thereby ensuring that the chemical characteristics are fully identified and analyzed. Although textual information covers many important characteristics of chemicals, it is often difficult to fully present all their hazardous features. The introduction of label image data fills the gaps in textual information that are difficult to fully cover. The combined analysis of text and images ensures that the hazardous characteristics of hazardous chemicals are fully assessed, thereby significantly improving the accuracy and completeness of the hazardous chemical knowledge base.
[0018] (2) The ontology-based method for constructing a hazardous chemical knowledge base uses a pre-trained semantic extraction model to extract text information and labeled image data, thereby ensuring that the characteristics of chemicals are fully identified. Based on the sub-network division of the semantic extraction model, the characteristics of the corresponding chemicals are fully extracted from the corresponding data. Then, the characteristics are integrated and filled in by the fusion output sub-network, thereby ensuring the complementarity between text and image information, ensuring the supplementation of missing information, and enabling the characteristics of hazardous chemicals to be fully characterized, thereby improving the construction efficiency and knowledge integrity of the chemical knowledge base.
[0019] (3) The ontology-based hazardous chemical knowledge base construction method analyzes the attribute set of each hazardous chemical to generate a comprehensive hazard characteristic value of each hazardous chemical and sorts them in descending order, thereby distinguishing the risk levels of different chemicals. This allows users to retrieve information on high-risk chemicals more intuitively and efficiently. The sorted chemical structure provides a systematic query framework, which facilitates quick access to all characteristics and hazard assessment information related to a specific chemical in practical applications. This enhances the data organization of the knowledge base and improves the intelligence level of querying and management in practical applications.
[0020] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0021] Figure 1 This is a flowchart of the ontology-based hazardous chemicals knowledge base construction method of the present invention.
[0022] Figure 2 This is a schematic diagram of the hazard feature set sequence data in the ontology-based hazardous chemical knowledge base construction method of the present invention.
[0023] Figure 3 This is a flowchart illustrating the specific steps involved in analyzing the semantic triplet set of each hazardous chemical in the ontology-based hazardous chemical knowledge base construction method of this invention.
[0024] Figure 4 This is a flowchart illustrating the specific steps involved in analyzing the image semantic representation vector of each hazardous chemical in the ontology-based hazardous chemical knowledge base construction method of this invention. Detailed Implementation
[0025] Please see Figure 1 This invention provides a technical solution: a method for constructing a hazardous chemicals knowledge base based on ontology, comprising the following steps: acquiring textual information (including but not limited to material safety data sheets, accident case reports, and operating procedure descriptions, such as unstructured content including chemical names, basic descriptions, physicochemical parameters, usage restrictions, storage conditions, reactivity descriptions, and accident records) and tag image data of several hazardous chemicals, and inputting them into a pre-trained semantic extraction model to analyze the semantic triplet set of each hazardous chemical; based on the semantic triplet set of each hazardous chemical, extracting the attribute set of each hazardous chemical (i.e., extracting all attribute triplets from the semantic triplet set to obtain the attribute set of each hazardous chemical, which contains all attributes of the hazardous chemical and their corresponding values); and analyzing the hazard feature set of each hazardous chemical, including fire and explosion hazard feature values, diffusion and pollution risk feature values, and reaction stability feature values; based on the hazard feature set of each hazardous chemical, analyzing the comprehensive hazard feature value of each hazardous chemical; and constructing a hazardous chemicals knowledge base based on the comprehensive hazard feature value and the semantic triplet set of each hazardous chemical.
[0026] The specific formula for calculating the comprehensive hazard characteristic value of a certain hazardous chemical is as follows:
[0027] ;in, This represents the comprehensive hazard characteristic value of a certain hazardous chemical. This refers to the fire and explosion hazard characteristic value of a certain hazardous chemical. The explosion coefficient is stored in the database. This represents the characteristic value of the diffusion and pollution risk of a certain hazardous chemical. The pollution coefficient is stored in the database. This is a characteristic value for the reactivity of a certain hazardous chemical. The stability coefficients are stored in the database. This is the adjustment coefficient stored in the database, and in this implementation example, it is set to 3.000. The interaction coefficients are stored in the database. This is the collaboration coefficient stored in the database, and in this implementation example, it is set to 0.100 (to avoid a denominator of 0).
[0028] It needs to be explained that the explosion coefficient is stored in the database. The steps are as follows: Read the fire and explosion hazard characteristic value, diffusion and pollution risk characteristic value, and reaction stability characteristic value for each hazardous chemical, and perform mean processing to obtain the mean values of the fire and explosion hazard characteristic, diffusion and pollution risk characteristic, and reaction stability characteristic. Then, sum these values to obtain the hazard sum value. Compare the mean fire and explosion hazard characteristic value with the hazard sum value, and use the result as the explosion coefficient. ;
[0029] Pollution coefficients stored in the database The steps for obtaining the values are as follows: Read the mean values of fire and explosion hazard characteristics, diffusion pollution risk characteristics, and reaction stability characteristics, and sum them to obtain the hazard sum value. Then, calculate the ratio between the mean diffusion pollution risk characteristics and the hazard sum value, and use the result as the pollution coefficient. ;
[0030] Stability coefficients stored in the database The steps for obtaining the values are as follows: Read the mean values of the disaster and explosion hazard characteristics, the mean values of the diffusion and pollution risk characteristics, and the mean values of the reaction stability characteristics, and sum them to obtain the hazard sum value. Then, calculate the ratio between the mean value of the reaction stability characteristics and the hazard sum value, and use the result as the stability coefficient. ;
[0031] Interaction coefficients stored in the database The steps are as follows: Read the fire and explosion hazard characteristic value, diffusion and pollution risk characteristic value, and reaction stability characteristic value for each hazardous chemical. Calculate the correlation value between any two characteristic values based on the Pearson coefficient, perform weighted processing, and use the result as the interaction coefficient. .
[0032] The following is an example of calculating the comprehensive hazard characteristic value of a certain hazardous chemical. The available data includes: fire and explosion hazard characteristic values, diffusion and pollution risk characteristic values, and reaction stability characteristic values of five hazardous chemicals, as shown in Table 1.
[0033] Table 1. Examples of Hazard Feature Set Sequence Data
[0034]
[0035] Explosion coefficient stored in the database Approximately 0.391;
[0036] Pollution coefficients stored in the database Approximately 0.363;
[0037] Stability coefficients stored in the database Approximately 0.246;
[0038] Adjustment coefficients stored in the database The value is: 3.000;
[0039] The interaction coefficient stored in the database is approximately 0.543;
[0040] The collaboration coefficient stored in the database is 0.010;
[0041] Substituting the data from Table 1 and the coefficients mentioned above into the specific formula for calculating the comprehensive hazard characteristic value of a certain hazardous chemical, we obtain:
[0042] The comprehensive hazard characteristic value of the first hazardous chemical is calculated as follows: (exp(0.391×0.743)+0.628^0.363+0*(1 / (1+0.287))^0.246) / 3.000)×(0.543×((0.743×0.528) / (0.287+0.010)))≈0.933;
[0043] The comprehensive hazard characteristic value of the second hazardous chemical is calculated as follows: (exp(0.391×0.432)+0.574^0.363+0*(1 / (1+0.316))^0.246) / 3.000)×(0.543×((0.432×0.574) / (0.316+0.010)))≈0.403;
[0044] The comprehensive hazard characteristic value of the third type of hazardous chemical = ((exp(0.391×0.561)+0.614^0.363+0*(1 / (1+0.354))^0.246) / 3.000)×(0.543×((0.561×0.614) / (0.354+0.010)))≈0.514;
[0045] The comprehensive hazard characteristic value of the fourth hazardous chemical is calculated as follows: (exp(0.391×0.687)+0.715^0.363+0*(1 / (1+0.238))^0.246) / 3.000)×(0.543×((0.687×0.715) / (0.238+0.010)))≈1.124;
[0046] The comprehensive hazard characteristic value of the fifth hazardous chemical is: (exp(0.391×0.527)+0.684^0.363+0*(1 / (1+0.412))^0.246) / 3.000)×(0.543×((0.527×0.684) / (0.412+0.010)))≈0.465.
[0047] The specific steps for constructing a hazardous chemicals knowledge base are as follows: First, sort the chemicals in descending order based on their comprehensive hazard characteristic values (ranking hazardous chemicals from highest to lowest comprehensive hazard characteristic value; the higher the comprehensive hazard characteristic value, the greater the hazard of the chemical), generating a chemical hazard ranking table. Then, based on the chemical hazard ranking table and the semantic triple set for each hazardous chemical, construct the hazardous chemicals knowledge base. Specifically, associate the comprehensive hazard characteristic value of each hazardous chemical in the generated chemical hazard ranking table with its corresponding semantic triple set. That is, store the comprehensive hazard characteristic value of each hazardous chemical together with its semantic triple set using a relational database (such as MySQL), thereby generating the hazardous chemicals knowledge base.
[0048] Specifically, such as Figure 2 As shown, the semantic extraction model includes a text recognition subnetwork, an image recognition subnetwork, and a fusion output subnetwork. The specific steps for analyzing the semantic triplet set of each hazardous chemical are as follows: In the text recognition subnetwork of the semantic extraction model, the text information of each hazardous chemical is received, and the structural semantic representation vector of each hazardous chemical is analyzed; In the image recognition subnetwork of the semantic extraction model, the label image data of each hazardous chemical is received, and the image semantic representation vector of each hazardous chemical is analyzed; In the fusion output subnetwork of the semantic extraction model, the structural semantic representation vector and the image semantic representation vector of each hazardous chemical are received, and the semantic triplet set of each hazardous chemical is analyzed.
[0049] The pre-training steps of the semantic extraction model are as follows: Construct a defined semantic dataset for hazardous chemicals. This dataset should include labeled text information and labeled image data, and accurately label each chemical characteristic (such as chemical name, physicochemical properties, risk behaviors, etc.) to ensure that the dataset covers different chemical categories (such as flammable, corrosive, volatile, etc.). Each data sample should provide complete semantic labels, including but not limited to: text data: containing chemical name, related properties, hazardous behaviors, etc. The labels should include chemical properties (such as flash point, density, lower explosive limit, etc.); image data: the labeled images should be manually labeled or automatically labeled to generate corresponding warning text areas (such as hazard signs, warning text, etc.) so as to serve as supervision signals for the model in subsequent training.
[0050] During the model structure initialization phase, an appropriate Transformer structure (such as BERT or RoBERTa) is used as the backbone network for text recognition and semantic analysis, and reasonable model hyperparameters are set. For the image recognition part, a convolutional neural network (CNN) can be used as the basic architecture for image preprocessing and text region recognition.
[0051] The semantic extraction model is trained, taking the text recognition sub-network training as an example: In the semantic encoding layer, the text information of each chemical is input and encoded. A word segmenter is used to segment the text, identifying the chemical's name, physicochemical properties, action verbs, etc. A pre-trained BERT or RoBERTa model is used to process the text. A self-attention mechanism is used to model the context dependency of each word in the text, generating a high-dimensional semantic representation vector. In the entity recognition layer, a bidirectional recurrent neural network (BiLSTM) is used to model the context of the text's semantic vectors, extracting the relationships between semantic entities and classifying and labeling entities. A conditional random field (CRF) is used to optimize the consistency of entity boundaries and the rationality of label distribution. In the text output layer, semantic structure combinatorial analysis is performed. The semantic entity vectors of each chemical text are used to construct subject-verb-object triples and other related structures, and structural filtering is performed to ensure that each triple conforms to grammatical rules and semantic consistency.
[0052] During training, cross-entropy loss is used to optimize the loss functions of the text and image sub-networks, and the entire model is trained using gradient descent algorithms (such as Adam). After each round of training, the model is evaluated using a validation set to ensure that the model can accurately extract semantic information from the text and warning content from the image. Multi-task joint training is adopted to fuse the features of text recognition and image recognition and train and optimize them together to improve the accuracy of comprehensive recognition.
[0053] We use a weighted combined loss function, combining various features of text and images, and use MSE, cross-entropy, IoU and other losses for error optimization to improve the model’s performance in semantic extraction tasks. During training, we use an early stopping mechanism to avoid overfitting and dynamically adjust the learning rate and optimizer parameters to ensure the model’s generalization ability in different tasks.
[0054] After each round of training, the model is evaluated using a validation set. Evaluation metrics such as precision, recall, and F1-score are calculated for the model in semantic recognition and image recognition tasks to ensure good performance of the model on different tasks. After multiple rounds of training and optimization, the trained model interface package is exported, which can automatically analyze the text and image data of each hazardous chemical, extract its semantic information, and construct a semantic triplet set.
[0055] In this implementation scheme, by integrating textual information and image data, the semantic features of hazardous chemicals can be comprehensively extracted. Secondly, during the training process, a multi-task joint training strategy is used, and optimization is performed through cross-entropy loss, MSE loss, etc. Conditional Random Field (CRF) is adopted to ensure the consistency of entity boundaries and the rationality of label distribution, thereby improving the accuracy of semantic recognition. At the same time, by combining textual and image features, the model can fully explore key information in the image that is not covered by the text, thereby supplementing the complete characteristics of hazardous chemicals. Finally, by optimizing the training strategy and loss function, overfitting is avoided while ensuring multi-task optimization, and the generalization ability of the model is improved. Thus, the trained model can efficiently analyze and extract textual and image data of chemicals, ensuring the comprehensiveness of the hazardous chemical knowledge base.
[0056] Specifically, the text recognition sub-network includes a semantic encoding layer, an entity recognition layer, and a text output layer. The specific steps for analyzing the high-dimensional semantic vector of each hazardous chemical are as follows: In the semantic encoding layer of the text recognition sub-network, the text information of each hazardous chemical is encoded. Specifically, this involves: cleaning and standardizing the encoding of Chinese characters in the text information (e.g., converting full-width characters and removing illegal symbols); performing sentence segmentation and segmentation annotation to distinguish independent semantic units; and using a dedicated word segmenter based on industry-specific vocabulary and public corpora to segment the text, identifying chemical nouns, units, parameter phrases, and lines. The word segmentation results are fed into a Chinese semantic understanding model (such as BERT or RoBERTa architecture) based on a multi-layer Transformer structure. This model uses a multi-layer self-attention mechanism to model the context dependency of each word in the text, and combines the semantic relationship between the words before and after the word in the sentence to generate a corresponding high-dimensional semantic representation vector. This semantic representation vector preserves the semantic meaning, structural role, and association characteristics of the word in the context, and outputs a semantic embedding result sequence composed of several high-dimensional semantic vectors.
[0057] In the entity recognition layer of the text recognition subnetwork, based on the encoded text information of each hazardous chemical, the semantic entity vector of each hazardous chemical is analyzed. Specifically, the semantic embedding result sequence is subjected to contextual modeling processing. That is, a bidirectional recurrent neural network is used to learn the contextual dependencies between each semantic vector to establish the sequence connection structure between words. In the modeling process, the semantic connection strength between any two adjacent words or phrases is extracted by combining the position index, semantic meaning, structural function, and association characteristics of each word in the text with other words (cosine similarity analysis is performed on the semantic meaning and structural function of any two adjacent words), and labeled. Record the connection type information (based on dependency analysis and syntax tree, extract the dependency type between two words, such as: subject-predicate relationship: such as (methane, trigger, explosion), describing the relationship between chemicals and events; modifier-headword relationship: such as (methane, has, flammability), describing the properties of chemicals; conditional relationship: such as (methane, at high temperature, spontaneous combustion), describing the reaction under specific conditions and marking it), including causal relationship, such as cause, trigger; attributive relationship, such as belong to, possess; conditional relationship, such as under… conditions, etc.; encode the semantic connection strength and connection type information as auxiliary features of the context structure into the context extension representation of each semantic vector;
[0058] Semantic entity recognition processing is performed on the modeled semantic vector sequence. This involves using a multi-label annotation mechanism (assigning multiple labels to the same input instance, such as a word or phrase, instead of selecting only one) to classify each semantic vector into a semantic category, determining its semantic role in the structural representation. Sequence constraint optimization is then performed using a Conditional Random Field (CRF) structure (considering the relationship between semantic connection strength and labels to ensure the final annotation results conform to language rules and semantic logic). This improves entity boundary consistency and label distribution rationality, ensuring the accurate attribution of each word in the semantic structure. Furthermore, during the recognition process, for semantic entities with numerical structures... The system extracts the corresponding numerical fields and unit identifiers of each hazardous chemical and embeds them into a vector representation structure. The final output is a semantic entity vector for each hazardous chemical, including: chemical name entities (e.g., chlorine, methane), physicochemical property entities (e.g., flash point, density, auto-ignition point, toxicity threshold, vapor diffusion gradient, conductivity change rate, gas-liquid interface energy consumption, thermal radiation absorption, reactivity, etc.), numerical attribute entities (e.g., density of 1.2 grams per cubic centimeter), risk behavior entities (e.g., leakage, corrosion, combustion), outcome event entities (e.g., poisoning, explosion, environmental pollution), and additional condition entities (e.g., high-temperature environment, confined space, near fire source, etc.).
[0059] In the text output layer of the text recognition subnetwork, based on the semantic entity vector of each hazardous chemical, the structural semantic representation vector of each hazardous chemical is analyzed. Specifically, structural combination analysis is performed on the semantic entity vector of each hazardous chemical (specifically, based on the relative order of words in the text, the semantic connection strength and connection type information between semantic entities, a semantic structure combination candidate set is constructed; in the semantic structure combination candidate set, the chemical name entity is used as the subject entity, the risk behavior entity or semantic conjunction is used as the semantic relation component, and the physicochemical attribute entity, numerical attribute entity, result event entity or additional condition entity is used as the object entity, the three together constitute the structural semantic expression unit), and output in the structural format of subject entity, semantic relation, and object entity;
[0060] Furthermore, during the structural combination analysis process, based on the semantic connection strength and connection type information between entities, a preliminary set of semantic structure candidates is constructed, namely multiple potential subject-verb-object semantic structure units, used to express the relationship between the attributes, behaviors, and consequences of hazardous chemicals in a specific semantic context. Structural filtering is then performed, that is, based on whether the semantic connection strength of the three types of semantic entities is higher than a preset threshold, whether the entities appear consecutively in the same syntactic block (based on syntactic trees or dependency parsing to confirm the dependency relationship between words, ensuring that subject-verb-object entities are at the same level and in the same structure, for example, in chlorine gas has a flash point, chlorine gas, has, and *flash point are all in the same syntactic block), to determine whether the structure has expressive completeness. If the structure is incomplete, type-conflicting, or logically inconsistent, it is eliminated.
[0061] The final output is a set of structural semantic representation vectors for each hazardous chemical. The structural semantic representation vector is a combination of subject entity vectors, semantic relation vectors, and object entity vectors. For example, the subject entity mainly includes the chemical name (such as chlorine, methane, etc.), the semantic relation can include verb phrases (such as having, causing, being in) or causal conjunctions (such as due to, causing), and the object entity includes the corresponding physicochemical property value (such as flash point of −34℃, lower explosive limit of 5%), event result (such as poisoning, explosion) or applicable conditions (such as in a confined space).
[0062] In this implementation scheme, meticulous text information processing and structured analysis enable the efficient extraction of various semantic information of hazardous chemicals and the establishment of structured semantic representation vectors. Secondly, the semantic encoding layer of the text recognition sub-network cleans and standardizes text information, performs word segmentation and encoding, ensuring the quality and consistency of the text data. Next, the entity recognition layer models semantic relationships using a bidirectional recurrent neural network, extracts key entities, and optimizes entity boundaries through conditional random fields to improve the accuracy and consistency of annotation. The text output layer generates structured semantic representation vectors through structural combination analysis, expressing various relationships between chemicals, such as the association between chemicals and events, conditions, etc. Finally, by filtering combinations that do not conform to the rules, the final generated semantic structure will be more consistent with language and semantic logic, avoiding erroneous or unreasonable combinations. This allows the model to accurately construct structured semantic representation vectors for chemicals, thereby significantly improving the accuracy of the hazardous chemical knowledge base.
[0063] Specifically, the label image data consists of the pixel value and two-dimensional coordinates of each pixel in the label image.
[0064] The image recognition sub-network includes a preprocessing layer, a warning recognition layer, and an image output layer. The specific steps for analyzing the image semantic representation vector of each hazardous chemical are as follows: In the preprocessing layer of the image recognition sub-network, the label image data of each hazardous chemical is preprocessed, specifically: the received label image data is subjected to size normalization (e.g., the image is uniformly scaled to a set resolution, such as 256×256 pixels), then color channel normalization is performed (specifically, the RGB three-channel pixel values of the image are compressed to the range of [0, 1] and color distribution is balanced), and image sharpening and contrast enhancement are performed (specifically, the Laplacian edge enhancement and histogram equalization algorithms are applied to improve the clarity of the text edges and the overall visual contrast in the image) to enhance the boundary features of the subsequent warning text. The enhanced label image is subjected to noise removal (e.g., using a median filter to eliminate random noise in the image background) and binarization (using an adaptive thresholding method to convert the image to a black and white image to highlight the outline of the text area).
[0065] In the warning recognition layer of the image recognition subnetwork, based on the preprocessed label image data of each hazardous chemical, the warning semantic entity vector of each hazardous chemical is analyzed. Specifically, connected component analysis (e.g., using 8-neighborhood) is performed on the binarized label image to obtain several candidate text regions. The aspect ratio of each candidate text region is extracted (i.e., for each candidate text region, the minimum and maximum row and column coordinates of all pixels within it are recorded, and the rectangular area formed by the minimum and maximum row and column points is defined as the minimum bounding rectangle of the candidate region, i.e., its region bounding box. Then, the pixel width in the horizontal direction and the pixel height in the vertical direction of the region bounding box are analyzed, i.e., the length and width of the region bounding box, and then the ratio is processed). Boundary compactness is calculated by determining the perimeter (sum of length and width of the bounding box of the region × 2) and area (length × width of the bounding box of the region), and then analyzing (perimeter^2) / (4 × pi × area). Pixel density is calculated as the ratio of the total number of pixels in each candidate text region to the total number of pixels in the label image. It is then determined whether the aspect ratio of each candidate text region is within a preset aspect ratio range, whether the boundary compactness is higher than a preset boundary compactness threshold, and whether the pixel density is higher than a preset pixel density threshold. If the aspect ratio of the candidate text region is within the preset aspect ratio range, and the boundary compactness and pixel density are both higher than the preset boundary compactness threshold, then the candidate text region is marked as a text region.
[0066] For each text region, character recognition processing is performed (i.e., OCR based on CRNN or Transformer architecture is used to recognize characters in each text region of the image, and the recognition result of each text region is output to obtain the corresponding character sequence). This yields the character sequence (i.e., the warning text) for each text region. Semantic category classification is then performed (the recognized warning text is compared with a dictionary in the chemical ontology knowledge base, mapping it to chemical name entities, physicochemical property entities, risk behavior entities, outcome event entities, and additional condition entities, etc.). Finally, the warning semantic entity vector for each valid text region is output, including: chemical name entities (e.g., chlorine, methane, etc.); physicochemical property entities (e.g., flash point, vapor pressure, lower explosive limit, etc.); numerical property entities (e.g., flash point is −34℃, etc.); risk behavior entities (e.g., leakage, combustion, etc.); outcome event entities (e.g., poisoning, explosion, etc.); and additional condition entities (e.g., high temperature environment, confined space, etc.).
[0067] In the image output layer of the image recognition sub-network, based on the warning semantic entity vector of each hazardous chemical, the image semantic representation vector of each hazardous chemical is analyzed. Specifically, structural combination processing is performed on the warning semantic entity vector of each hazardous chemical, that is, constructing an image semantic structure candidate set (including several image semantic candidate structures). Based on the warning text, the dependency relationship and connection information between each entity are analyzed, and the combination relationship between different warning semantic entities is established. Dependency relationship analysis is performed by identifying the association mode between different entities through syntactic analysis or syntactic tree structure. For example, the word "trigger" indicates a causal relationship. If there is a triggering relationship between chemical A and explosion in the image, then the text combination should be regarded as a causal structure. Connection information analysis is performed by identifying and processing semantic conjunctions (such as having, triggering, causing, etc.). These conjunctions indicate the relationship type between entities, which helps to construct a more complete semantic structure. In the image semantic structure candidate set, the chemical name entity is used as the subject entity, the risk behavior entity is used as the semantic relationship component, and the physical and chemical attribute entity, numerical attribute entity, result event entity or additional condition entity is used as the object entity. The three together constitute the image semantic candidate structure.
[0068] For each image semantic candidate structure in the image semantic structure candidate set, a structure filtering process is performed. Specifically, based on the components of each candidate structure (i.e., subject entity, semantic relation, and object entity), it is determined whether it conforms to the preset semantic combination rules (such as type conflict, i.e., the subject and object types do not match, semantic inconsistency, i.e., it does not conform to the conventional causal relationship). If the image semantic candidate structure does not conform to the preset rules, it is eliminated; otherwise, it is marked as an image semantic valid structure. Image semantic valid structures that conform to the rules are retained. Each image semantic valid structure (including subject entity, semantic relation, and object entity) is converted into a vector representation and concatenated to form the final image semantic representation vector. For example, if methane is the subject and the semantic relation is flammability as the object, the final combination is: subject entity vector: methane, semantic relation vector: , object entity vector: flammability.
[0069] In this implementation scheme, the accuracy and completeness of semantic extraction of hazardous chemicals are significantly improved through in-depth analysis of hazardous chemical label image data. Secondly, in the preprocessing layer of the image recognition sub-network, the warning text and label information in the image are enhanced by processing the label image data, thereby ensuring the clarity and accuracy of subsequent warning information. In the warning recognition layer, OCR technology is used to recognize characters in each text region, and the recognized warning text is compared with the chemical ontology knowledge base to accurately map it to the corresponding semantic entities, thereby ensuring that the key information of the chemical can be extracted and converted into meaningful semantic vectors. Finally, the image output layer combines the warning text information with the relevant attributes of the chemical, and through structural combination analysis, establishes an effective semantic structure based on syntax trees and semantic connection strength, thereby ensuring the accuracy of the effective semantic structure of each image. This provides comprehensive semantic data for the construction of the hazardous chemical knowledge base, thereby improving the accuracy of the knowledge base.
[0070] Specifically, the fusion output subnetwork includes a filling fusion layer and an output layer. The specific content of the fusion output subnetwork is as follows: In the filling fusion layer of the fusion output subnetwork, based on the structural semantic representation vector and image semantic representation vector of each hazardous chemical, the fusion semantic representation vector of each hazardous chemical is analyzed. Specifically, the structural semantic representation vector and the image semantic representation vector are fused, that is, the structural semantic representation vector is dominant, and the image semantic representation vector is used as supplementary information to fill in the gaps (that is, if the image semantic representation vector contains entities that do not appear in the text information, then these missing entities are added to the structural semantic representation vector). In a vector representation, two vectors are merged by concatenation to form a complete representation vector containing both textual and image semantic structures. For example, if the structural semantic representation vector contains "methane is flammable" and the image semantic representation vector contains "methane causes an explosion", then the merged vector is "methane is flammable and causes an explosion". This merged semantic representation vector includes multiple semantic structures. It should be noted that for chemical name entities, physicochemical property entities, and numerical property entities, several attribute structures are included in the semantic structure, such as chlorine, flash point, −34℃; chlorine, density, 1.2 grams per cubic centimeter, etc.
[0071] In the output layer of the fusion output sub-network, the fusion semantic representation vector of each hazardous chemical is processed to obtain a set of semantic triples for each hazardous chemical. Specifically, the fusion semantic representation vector of each hazardous chemical is standardized to ensure that all entities (including text, images, and numerical attributes) are compared and analyzed on the same scale. For example, L2 normalization is used to adjust each dimension of the fusion vector to a uniform scale. Then, the set of semantic triples for each hazardous chemical is output, including several semantic structure triples and attribute triples. The semantic structure triples are specifically represented in the form of subject entity, semantic relation, and object entity, such as (methane, has, flammability), (methane, initiates, explosion). The attribute triples are specifically represented in the form of subject entity, attribute type, and attribute value, such as (chlorine, flash point, −34℃), (chlorine, density, 1.2 grams per cubic centimeter), etc., representing the chemical's attributes and corresponding numerical values.
[0072] In this implementation scheme, the comprehensiveness of the hazardous chemicals knowledge base is effectively improved by fusing textual and image semantic information. Secondly, by filling the fusion layer and combining the structural semantic representation vector with the image semantic representation vector, it is ensured that the uncovered textual data contained in the image can be supplemented into the semantic structure of the text, thereby enabling the complete hazard information of chemicals to be fully represented, thus improving the richness and accuracy of semantic information. Finally, the generated semantic triple set effectively expresses the attributes, behaviors and related events of chemicals, thereby ensuring that the characteristics of hazardous chemicals are fully analyzed, thus improving the efficiency and accuracy of knowledge base construction.
[0073] Specifically, the attribute set includes flash point, auto-ignition point, toxicity threshold, vapor diffusion gradient, conductivity change rate, gas-liquid interface energy consumption, thermal radiation absorption, reactivity, and gas-liquid interface tension.
[0074] The flash point is the lowest temperature at which a chemical can be ignited under specific conditions (such as standard atmospheric pressure).
[0075] The auto-ignition point is the lowest temperature at which a chemical can spontaneously ignite without an external ignition source. Even without a spark or open flame, a chemical can spontaneously ignite due to excessively high ambient temperatures.
[0076] The toxicity threshold refers to the highest concentration of a chemical that, under specific conditions (such as concentration and time), will not cause significant health effects on humans or the environment upon exposure. The lower the toxicity threshold, the greater the harm of the chemical to humans or the environment.
[0077] The vapor diffusion gradient value is the rate and direction of chemical diffusion from different concentration areas in gaseous form, that is, diffusion from high concentration areas to low concentration areas. When a chemical leaks, if it has a high vapor diffusion gradient value, the chemical will rapidly diffuse into the surrounding environment. This means that the chemical can spread over long distances, increasing the risk of fire, explosion, or environmental pollution.
[0078] The rate of change of conductivity is the speed at which the conductivity (the ability of a substance to conduct electricity) of a chemical changes when it leaks or volatilizes. A large rate of change of conductivity indicates that the chemical may be undergoing a chemical reaction in the environment or interacting with the surrounding environment.
[0079] The energy consumption value of the gas-liquid interface is the energy consumed at the gas-liquid interface when a chemical interacts with the environment. It reflects the chemical's ability to volatilize and diffuse. High-energy-consuming chemicals may volatilize and diffuse into the air rapidly when leaked, which will accelerate their diffusion rate in the environment and thus increase the risk of explosion or fire.
[0080] Thermal absorptivity (TA) is the ability of a chemical to absorb radiant energy from a heat source (such as a fire or explosion). It reflects the degree of physical or chemical change that a chemical undergoes when heated in a high-temperature environment. A high TTA means that a chemical is more susceptible to the effects of heat radiation in accidents such as fires, potentially leading to more intense reactions such as spontaneous combustion or explosion.
[0081] Reactivity is the rate at which a chemical reacts with other substances, i.e., the chemical reaction rate. Chemicals with high reactivity are more likely to react during storage, transportation, or use, potentially leading to explosions, fires, or toxic gas leaks.
[0082] The gas-liquid interfacial tension is the intermolecular force between a gas and a liquid, reflecting the degree of contact and stability between gas molecules and liquid. A higher gas-liquid interfacial tension indicates that the chemical molecules are more strongly bound together at the gas-liquid interface. Chemicals with high gas-liquid interfacial tension are less likely to volatilize and diffuse when leaked, while chemicals with low gas-liquid interfacial tension are more likely to evaporate and diffuse into the air, posing a risk of fire or explosion.
[0083] The specific steps for analyzing the hazard characteristic set of each hazardous chemical are as follows: Based on the flash point value and auto-ignition point value of each hazardous chemical, analyze the fire and explosion hazard characteristic value of each hazardous chemical. Specifically, standardize the flash point value, auto-ignition point value, and reactivity value of each hazardous chemical. Based on the standardization results, perform ratio processing, i.e., 1 / (1+weighted processing result), to obtain the fire and explosion hazard characteristic value of the corresponding hazardous chemical.
[0084] Based on the conductivity change rate, vapor diffusion gradient, toxicity threshold, and gas-liquid interface energy consumption value of each hazardous chemical, the diffusion pollution risk characteristic value of each hazardous chemical is analyzed. Specifically, the conductivity change rate, vapor diffusion gradient, toxicity threshold, and gas-liquid interface energy consumption value of each hazardous chemical are standardized, and the results of the standardization are weighted to obtain the diffusion pollution risk characteristic value of the corresponding hazardous chemical.
[0085] Based on the thermal radiation absorption value, reactivity value, and gas-liquid interfacial tension value of each hazardous chemical, the reaction stability characteristic value of each hazardous chemical is analyzed. Specifically, the thermal radiation absorption value, reactivity value, and gas-liquid interfacial tension value of each hazardous chemical are standardized to obtain the standardized thermal radiation absorption value, reactivity value, and gas-liquid interfacial tension value of each hazardous chemical, and then weighted to obtain the reaction stability characteristic value of the corresponding hazardous chemical.
[0086] The specific formula for calculating the reaction stability characteristic value of a certain hazardous chemical is as follows:
[0087] ;in, This is a characteristic value for the reactivity of a certain hazardous chemical. This refers to the thermal radiation absorption value of a certain hazardous chemical after standardization. The absorption coefficient is stored in the database. This refers to the reactivity value of a certain hazardous chemical after standardization. These are the activity coefficients stored in the database. This refers to the gas-liquid interfacial tension value of a certain hazardous chemical after standardization. The tension coefficient is stored in the database. .
[0088] It needs to be explained that the absorption coefficient is stored in the database. Activity coefficient Tension coefficient The acquisition steps are as follows: Read the standardized thermal radiation absorption value, reactivity value, and gas-liquid interfacial tension value for each hazardous chemical, and perform averaging to obtain the average thermal radiation absorption value, average reactivity value, and average gas-liquid interfacial tension value. Then, sum these values to obtain the reaction stability sum value. Ratios of the average thermal radiation absorption value, average reactivity value, and average gas-liquid interfacial tension value to the reaction stability sum value are used as the absorption coefficient. Activity coefficient Tension coefficient .
[0089] This implementation plan analyzes the hazard characteristics of hazardous chemicals by combining different attributes, thereby comprehensively assessing the degree of hazard. Secondly, by standardizing and weighting the various attributes of hazardous chemicals, it effectively quantifies various hazard characteristics. Based on the relative importance of different hazard characteristics, it rationally allocates the impact of each characteristic on the hazard assessment results, thus achieving a comprehensive and accurate hazard assessment of hazardous chemicals. Finally, through refined processing, it can effectively ensure that various hazard characteristics of hazardous chemicals are comprehensively assessed, thereby improving the accuracy of the entire knowledge base.
[0090] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0091] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for constructing a hazardous chemicals knowledge base based on ontology, characterized in that, Includes the following steps: The text information and label image data of several hazardous chemicals are obtained, specifically the pixel value and two-dimensional coordinates of each pixel in the label image. These are then input into a pre-trained semantic extraction model, which includes a text recognition sub-network, an image recognition sub-network, and a fusion output sub-network. The model analyzes the semantic triplet set for each hazardous chemical, specifically: In the text recognition subnetwork, structural semantic representation vectors are extracted from text information; In the image recognition sub-network, text region recognition and character recognition are performed on the label image data to obtain warning text, and the image semantic representation vector is identified based on the warning text; In the fusion output sub-network, the structural semantic representation vector is the main component. When there are semantic entities or semantic relations in the image semantic representation vector that are not included in the structural semantic representation vector, the semantic entities or semantic relations are added to the structural semantic representation vector to generate the fusion semantic representation vector. The corresponding set of semantic triples is then analyzed based on the fusion semantic representation vector. Based on the semantic triple set of each hazardous chemical, the attribute set of each hazardous chemical is extracted, and the hazard characteristic set of each hazardous chemical is analyzed, including fire and explosion hazard characteristic value, diffusion and pollution risk characteristic value, and reaction stability characteristic value. Based on the hazard characteristic set of each hazardous chemical, the comprehensive hazard characteristic value of each hazardous chemical is analyzed, and the specific formula for calculating the comprehensive hazard characteristic value of a certain hazardous chemical is as follows: ; in, , , , The values are, in order, the comprehensive hazard characteristic value, the fire and explosion hazard characteristic value, the diffusion and pollution risk characteristic value, and the reaction stability characteristic value of a certain hazardous chemical. , , , , The coefficients stored in the database are, in order: explosion coefficient, pollution coefficient, stability coefficient, regulation coefficient, interaction coefficient, and synergy coefficient. A hazardous chemicals knowledge base is constructed based on the comprehensive hazard characteristic value and semantic triple set of each hazardous chemical, specifically as follows: Based on the comprehensive hazard characteristic values, the chemicals are sorted in descending order to generate a hazard ranking table. Then, a hazard knowledge base is constructed by combining the semantic triple set of each hazardous chemical.
2. The method for constructing a hazardous chemicals knowledge base based on ontology according to claim 1, characterized in that, The text recognition sub-network consists of a semantic encoding layer, an entity recognition layer, and a text output layer. The specific steps for analyzing the high-dimensional semantic vector of each hazardous chemical are as follows: In the semantic coding layer of the text recognition subnetwork, the text information of each hazardous chemical is encoded. In the entity recognition layer of the text recognition subnetwork, the semantic entity vector of each hazardous chemical is analyzed based on the text information of each hazardous chemical after encoding processing; In the text output layer of the text recognition subnetwork, the structural semantic representation vector of each hazardous chemical is analyzed based on the semantic entity vector of each hazardous chemical.
3. The method for constructing a hazardous chemicals knowledge base based on ontology according to claim 1, characterized in that, The image recognition subnetwork includes a preprocessing layer, a warning recognition layer, and an image output layer. The specific steps for analyzing the image semantic representation vector of each hazardous chemical are as follows: In the preprocessing layer of the image recognition subnetwork, the label image data of each hazardous chemical is preprocessed; In the warning recognition layer of the image recognition subnetwork, the warning semantic entity vector of each hazardous chemical is analyzed based on the preprocessed label image data of each hazardous chemical. In the image output layer of the image recognition subnetwork, the image semantic representation vector of each hazardous chemical is analyzed based on the warning semantic entity vector of each hazardous chemical.
4. The method for constructing a hazardous chemicals knowledge base based on ontology according to claim 1, characterized in that, The attribute set includes flash point, auto-ignition point, toxicity threshold, vapor diffusion gradient, conductivity change rate, gas-liquid interface energy consumption, thermal radiation absorption, reactivity, and gas-liquid interface tension.
5. The method for constructing a hazardous chemicals knowledge base based on ontology according to claim 4, characterized in that, The specific steps for analyzing the hazard profile set of each hazardous chemical are as follows: Based on the flash point and auto-ignition point of each hazardous chemical, the fire and explosion hazard characteristics of each hazardous chemical are analyzed. Based on the conductivity change rate, vapor diffusion gradient, toxicity threshold, and gas-liquid interface energy consumption of each hazardous chemical, the diffusion pollution risk characteristics of each hazardous chemical are analyzed. Based on the thermal radiation absorption value, reactivity value, and gas-liquid interfacial tension value of each hazardous chemical, the reaction stability characteristics of each hazardous chemical are analyzed.