Transform-based security data table hazard entity identification method and system
By employing a Transformer-based preprocessing and feature fusion method, combined with knowledge graph construction, the unstructured nature of SDS texts is addressed, enabling efficient and reliable identification of hazardous entities, reducing manual review costs, and adapting to the differences in wording among various suppliers.
Patent Information
- Application Number
- CN202511913208.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-20
AI Technical Summary
Existing Security Data Sheets (SDS) are unstructured or semi-structured due to differences in vendor-defined formats and wording. This makes it difficult to efficiently and automatically identify key hazard entities, and manual review is costly and poorly adaptable, making it difficult to meet the needs of large-scale and rapid decision-making.
We employ a Transformer-based preprocessing, feature extraction and fusion, and knowledge graph construction method. Through BIO annotation, H encoding, Transformer model, and attention mechanism, we accurately identify hazardous entities in SDS, including adverse effects, adverse reaction sites, carcinogenic descriptions, non-carcinogenic claims, harmful effects, and toxic doses, and construct a three-dimensional knowledge graph.
It improves the efficiency and reliability of hazardous entity identification, reduces compliance verification costs, is highly adaptable, can quickly process large-scale SDS text libraries, and complies with OSHA standards.
Smart Images

Figure CN121706782A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of large language analysis, in particular to a safety data sheet hazard entity identification method and system based on a Transformer. BACKGROUND
[0002] Safety data sheets (SDS) are the authoritative tools for communicating hazard information and protection requirements in the chemical supply chain. The SDS provided by the supplier with each batch of goods enables downstream environmental health and safety departments, operations and logistics teams, regulatory agencies, and emergency response personnel to perform standard operating procedures, complete transportation classification and compliance work, and carry out emergency preparedness and response. Although the chapter layout of SDS is limited by templates, the differences in the specific expression style of suppliers result in the overall presentation of the text being unstructured or semi-structured. With the proliferation of product formulations, downstream environmental health and safety departments, operations and logistics teams, regulatory agencies, and emergency response personnel have to manually review a continuously expanding and structurally chaotic text library, which is increasingly inefficient, hindering reliable extraction of key facts and driving up compliance verification costs.
[0003] With the support of the United Nations Economic and Social Council Subcommittee on the Transport of Dangerous Goods, the United States Occupational Safety and Health Administration (OSHA), the Canadian Health Department, and other agencies, safety data sheets (SDS) have become the core tool for communicating chemical hazards. In terms of information management, the CAS registry has included over 290 million publicly available substances, making knowledge management and retrieval centered on SDS a rigid requirement. In terms of compliance, OSHA Hazard Communication Standard and Appendix D specify a 16-section framework and minimum information elements, but do not explicitly specify the format or content, resulting in differences in expression between suppliers.
[0004] However, the current SDS preparation specifications give suppliers the autonomy to format and phrase, resulting in problems such as heterogeneous templates, inconsistent terminology, missing data fields, and cross-document differences, which result in SDS actually presenting unstructured or semi-structured characteristics. Rule-based or dictionary-based processes and traditional sequence labeling methods have high maintenance costs and weak generalization capabilities across suppliers, long documents, and mixed text and image formats. Extraction systems that rely on manual review and high-frequency maintenance are costly, have poor adaptability, and are difficult to meet the large-scale management and rapid decision-making needs of complex long document scenarios. SUMMARY
[0005] The purpose of the present application is to provide a safety data sheet hazard entity identification method and system based on a Transformer, which accurately identifies hazard entities in safety data sheets through Transformer-based preprocessing, feature extraction and fusion, and knowledge graph construction, and improves the efficiency and reliability of hazard entity identification.
[0006] To achieve the above object, the present application provides the following scheme: A safety data table hazard entity identification method based on a Transformer, comprising the following steps: The original heterogeneous data of the safety data table is preprocessed, and the hazard entities in the preprocessed original heterogeneous data are labeled by a BIO labeling method to obtain a standardized labeled text sample; the hazard entities include adverse effects, adverse reaction sites, carcinogenicity descriptions, non-carcinogenicity statements, harmful effects, and toxic doses; Based on the determined H code of the chemical components in the standardized labeled text sample, a toxicity prior dictionary is constructed by the H code and encoded into a toxicity prior vector; The context semantic enhancement vector in the standardized labeled text sample is extracted by a Transformer model, and a component-hazard correlation vector is constructed according to the semantic similarity between the chemical components and the hazard entities; A bidirectional encoder is constructed according to an attention mechanism, and the toxicity prior vector, the context semantic enhancement vector, and the component-hazard correlation vector are fused by the bidirectional encoder to obtain a fusion feature; The fusion feature is preliminarily identified by a Transformer model to obtain a BIO label sequence; A three-dimensional knowledge graph is constructed according to the component-hazard correlation vector, the CAS number of the chemical components, and the BIO label sequence.
[0007] Optionally, the original heterogeneous data of the safety data table is preprocessed, and the hazard entities in the preprocessed original heterogeneous data are labeled by a BIO labeling method to obtain a standardized labeled text sample, comprising: Based on the chapter framework specified by OSHA, the core chapter keywords of the original heterogeneous data are extracted, the original heterogeneous data is classified into templates by a K-means clustering algorithm to obtain typical templates, and the core chapter keywords are mapped to standard field names according to the typical templates to obtain a unified text; The unified text is subjected to meaningless data cleaning, format standardization, and term normalization to obtain an initial sample; The initial sample is subjected to missing information completion and redundant information filtering to obtain a data sample; Labeling rules are constructed according to the hazard entities, and the data sample is automatically labeled by the labeling rules and the fuzzy boundary entities are corrected to obtain a standardized labeled text sample.
[0008] Optionally, based on the determined H code of the chemical components in the standardized labeled text sample, a toxicity prior dictionary is constructed by the H code and encoded into a toxicity prior vector, comprising: H code is obtained by querying the CAS number of the chemical composition; An toxicity prior dictionary is constructed according to the H code, and a dictionary slot is determined; The slot value of the dictionary slot is determined according to the category description of the chemical composition; An initial vector is constructed according to the dictionary slot and the slot value, and the initial vector is blocked and spliced according to the slot order to obtain a toxicity prior vector.
[0009] Optionally, a context semantic enhancement vector is extracted from the standardized annotated text sample by a Transformer model, and a cost and hazard correlation vector is constructed according to the semantic similarity between the chemical composition and the hazard entity, including: The basic Transformer model is fine-tuned in the field by the standardized annotated text sample and the H code to obtain a specific model; The standardized annotated text sample is segmented into sentence units, and the sentence units are subjected to mean pooling by the specific model to obtain a sentence-level semantic vector; The context window text of the hazard entity is extracted, and the entity center hidden state of the context window text is extracted by the specific model to obtain an entity-level semantic vector; The sentence-level semantic vector and the entity-level semantic vector are weighted and summed according to elements, and are multiplied element by element with the toxicity prior vector to obtain a context semantic enhancement vector; The CAS number and the hazard entity are both embedded vectors, and the cosine similarity between the embedded vectors is calculated; The H code and the hazard entity are semantically matched by the specific model to obtain a matching score; Based on the pre-constructed dose and hazard intensity mapping rule, the association score of the toxicity dose and the hazard severity is calculated by the specific model, and the cosine similarity, the matching score and the association score are weighted and summed to obtain the semantic similarity; The semantic similarity of the hazard entity is calculated respectively, and a similarity matrix is constructed according to the plurality of semantic similarities; The similarity matrix is subjected to matrix compression and global maximum pooling processing to obtain a preliminary correlation feature; The preliminary correlation feature is weighted and summed with the preset differentiated weight, and the obtained weighted feature is subjected to linear transformation to obtain a cost and hazard correlation vector.
[0010] Optionally, a bidirectional encoder is constructed according to an attention mechanism, and the toxicity prior vector, the context semantic enhancement vector and the composition and hazard correlation vector are fused by the bidirectional encoder to obtain a fusion feature, including: The toxicity prior vector and the composition and hazard correlation vector are spliced to obtain a fusion vector; The fusion vector and the context semantic enhancement vector are fused by an attention mechanism to obtain a first feature; An average negative log-likelihood score of the standardized annotated text sample is calculated, the average negative log-likelihood score is taken as a sample difficulty score to grade the first feature, and the parameters of the bidirectional encoder are adjusted according to the grading result to obtain an optimized encoder; The fusion vector and the context semantic enhancement vector are fused by the optimized encoder to obtain a fusion feature.
[0011] Optionally, the calculation formula of the average negative log-likelihood score is: ; wherein, is a numerical stability term, is an effective token set, is a prediction probability.
[0012] Optionally, the focal loss function expression of the optimized encoder is: ; wherein, is a softmax probability, is a control focus intensity, is a category weight, and Ω is an effective supervision position set.
[0013] Optionally, the fusion feature is subjected to preliminary entity recognition by a Transformer model to obtain a BIO label sequence, including: The labels of the fusion feature are extracted by the Transformer model to obtain a preliminary entity set; The boundary of the preliminary entity set is checked, and the preliminary entity set is modified according to the checking result to obtain a modified result; When the overlap rate of the modified result is ≥80%, the entity boundary is updated and the boundary is checked again, and when the overlap rate of the modified result is <80%, the modified result is retained and false positive results that do not conform to the OSHA entity definition are removed to obtain the BIO label sequence.
[0014] A safety data table hazard entity recognition system based on a Transformer includes: A text processing module is configured to perform data preprocessing on original heterogeneous data of a safety data table, and label hazard entities in the preprocessed original heterogeneous data by a BIO labeling method to obtain a standardized annotated text sample. The hazard entities include adverse effects, adverse reaction sites, carcinogenicity descriptions, non-carcinogenicity statements, harmful effects, and toxic doses. A prior construction module is configured to determine H encoding based on chemical components in the standardized annotated text sample, construct a toxicity prior dictionary by the H encoding, and encode the toxicity prior dictionary into a toxicity prior vector; The feature extraction module is used to extract contextual semantic enhancement vectors from standardized labeled text samples using the Transformer model, and to construct component-hazard association vectors based on the semantic similarity between chemical components and hazardous entities. The feature fusion module is used to construct a bidirectional encoder based on the attention mechanism, and then fuse the toxicity prior vector, the contextual semantic enhancement vector, and the component-harm association vector in a two-way manner through the bidirectional encoder to obtain the fused features. The entity recognition module is used to perform preliminary entity recognition on the fused features through the Transformer model to obtain the BIO tag sequence. The knowledge graph construction module is used to construct a three-dimensional knowledge graph based on the component-hazard association vector, the CAS number of the chemical component, and the BIO tag sequence.
[0015] According to specific embodiments provided by the present invention, the following technical effects are disclosed: The present invention provides a method and system for identifying hazardous entities in a safety data table based on Transformer. The method includes: preprocessing the original heterogeneous data of the safety data table, and annotating the hazardous entities in the preprocessed original heterogeneous data using the BIO annotation method to obtain standardized annotated text samples; the hazardous entities include: adverse effects, adverse reaction sites, carcinogenic descriptions, non-carcinogenic claims, harmful effects, and toxic doses; based on the determination of H-encoding of chemical components in the standardized annotated text samples, constructing a toxicity prior dictionary through H-encoding and encoding the toxicity prior dictionary into a toxicity prior vector; extracting contextual semantic enhancement vectors from the standardized annotated text samples using a Transformer model, and constructing a component-hazard association vector based on the semantic similarity between chemical components and hazardous entities; constructing a bidirectional encoder based on an attention mechanism, and fusing the toxicity prior vector, contextual semantic enhancement vector, and component-hazard association vector through the bidirectional encoder to obtain fused features; performing preliminary entity recognition on the fused features using a Transformer model to obtain a BIO tag sequence; and constructing a three-dimensional knowledge graph based on the component-hazard association vector, the CAS number of the chemical components, and the BIO tag sequence. This method accurately identifies hazardous entities in the security data table by using Transformer preprocessing, feature extraction and fusion, and knowledge graph construction, and improves the efficiency and reliability of hazardous entity identification. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of the safety data table hazard entity identification method of the present invention; Figure 2 This is a statistical result diagram of the labeled entities in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the safety data table hazard entity identification system of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] like Figure 1 As shown, this invention provides a method for identifying hazardous entities using a security data table based on Transformer, comprising the following steps: Step 100: Perform data preprocessing on the raw heterogeneous data of the safety data sheet, and use the BIO annotation method to annotate the hazard entities in the preprocessed raw heterogeneous data to obtain standardized annotation text samples; hazard entities include: adverse effects, adverse reaction sites, carcinogenicity descriptions, non-carcinogenicity claims, harmful effects, and toxicity doses; Step 200: Based on the determination of chemical components in standardized annotated text samples, H-encoding is used to construct a toxicity prior dictionary and encode the toxicity prior dictionary into a toxicity prior vector; Step 300: Extract contextual semantic enhancement vectors from standardized labeled text samples using the Transformer model, and construct component-hazard association vectors based on the semantic similarity between chemical components and hazardous entities; Step 400: Construct a bidirectional encoder based on the attention mechanism, and fuse the toxicity prior vector, the contextual semantic enhancement vector, and the component-harm association vector in a dual-path manner through the bidirectional encoder to obtain the fused features; Step 500: Perform preliminary entity recognition on the fused features using the Transformer model to obtain the BIO tag sequence; Step 600: Construct a three-dimensional knowledge graph based on the component-hazard association vector, the CAS number of the chemical component, and the BIO tag sequence.
[0021] Preferably, the original heterogeneous data of the safety data table is preprocessed, and the hazardous entities in the preprocessed original heterogeneous data are labeled using the BIO annotation method to obtain standardized labeled text samples, including: Based on the chapter framework specified by OSHA, the core chapter keywords of the original heterogeneous data are extracted. The original heterogeneous data is then classified into templates using the K-means clustering algorithm to obtain typical templates. Based on the typical templates, the core chapter keywords are mapped to standard field names to obtain unified text. The unified text is cleaned of meaningless data, its format is standardized, and its terminology is normalized to obtain the initial sample; The initial sample is filled with missing information and filtered for redundant information to obtain a data sample. Based on the harmful entities, annotation rules are constructed, and the data samples are automatically annotated and the fuzzy boundary entities are corrected by rule-driven annotation rules to obtain standardized annotated text samples.
[0022] In the specific implementation process, step 100 first defines the scope of core chapters such as hazard identification, ingredient information, and first aid measures based on the 16-chapter framework specified in Appendix D of the OSHA Hazard Communication Standard. Then, it calculates the importance score of each word in the original heterogeneous data using the TextRank algorithm. The calculation formula is as follows: Where v is the target word, Let Out(v) be the set of words pointing to v, d be the damping coefficient (with a value of 0.85), and |Out(u)| be the out-degree of word u. Words ranking in the top 20% and highly relevant to the core chapter's theme are extracted as core chapter keywords. Subsequently, K-means clustering is used to classify the original heterogeneous data into templates. First, the core chapter keywords are converted into TF-IDF feature vectors. The optimal number of clusters k is determined using the elbow rule. Then, the K-means iterative process is executed with this k value to obtain typical templates. A mapping table between each typical template and OSHA standard field names is then established. The cosine similarity between the core chapter keywords and the standard field names is calculated. Core chapter keywords with a similarity ≥ 0.8 are mapped to their corresponding standard field names to obtain unified text. Next, meaningless data cleaning was performed on the unified text, removing useless special characters such as parentheses and exclamation marks, eliminating redundant spaces and special characters, standardizing Chinese and English symbols, and standardizing the format using regular expression matching to unify dosage units and date formats. Terminology normalization was based on the GHS hazard terminology dictionary, using precise string matching and semantic alignment to unify synonyms such as "toxic," "highly toxic," and "harmful" into the relevant standard expression "ToxicityDescription," obtaining the initial sample. Subsequently, missing information was filled in on the initial sample by querying the CAS registry to fill in missing chemical component CAS numbers and querying the GHS hazard database to fill in missing H codes and related hazard description information. Redundant information filtering was based on the cosine similarity of TF-IDF vectors, setting a similarity threshold of 0.9, and removing duplicate text segments with similarity ≥ 0.9 to obtain the data sample. Subsequently, based on the six hazard entities defined by OSHA, including: Symptom, Target Site, Cracinogenicity, Non-Cracinogenicity, Toxicity Description, and Toxicological Dose, as shown in Table 1, the statistical results are as follows: Figure 2 As shown.
[0023] Table 1. Detailed Description of Hazardous Entities
[0024] Construct a labeling rule base, including entity boundary rules (such as labeling symptom descriptions following trigger words like "appears / causes / leads to" as adverse effects) and terminology matching rules (such as those containing "LD"). 50 “LC” 50The data samples were labeled with various rules, including: "equal dose parameters" (e.g., toxic doses), "carcinogenic" or "carcinogenicity" (e.g., descriptions of body parts following verbs like "acts on / affects / damages") (e.g., adverse reaction sites). A BIO (Browser-Induced Isolation) scheme was used for rule-driven automatic labeling of the data samples. "B" marked the start position of each entity, "I" marked the middle and end positions, and "O" marked non-entity parts. A sliding window with a size of three lexical units was used to traverse potentially ambiguous boundary regions. Semantic relevance scores between lexical units and corresponding entity categories within the window were calculated using Word2Vec word vectors fine-tuned based on the SDS domain corpus. A confidence threshold of 0.7 was set, and entity boundaries below the threshold were adjusted. Furthermore, the GHS hazard terminology database and contextual semantic logic were combined to correct for "entity-non-entity ambiguity" and "confusion between different entity categories," ultimately resulting in standardized labeled text samples.
[0025] Preferably, based on the H-encoding of chemical components determined in standardized annotated text samples, a toxicity prior dictionary is constructed using the H-encoding, and the toxicity prior dictionary is encoded into a toxicity prior vector, including: The H code can be obtained by looking up the CAS number of the chemical composition; Construct a toxicity prior dictionary based on H-codes and determine dictionary slots; The slot value of the dictionary slot is determined by the category description of the chemical composition; An initial vector is constructed based on the dictionary slots and slot values, and the initial vector is then segmented and concatenated according to the slot order to obtain the toxicity prior vector.
[0026] In the specific implementation process, step 200, for the feature data of the ToxicityDescription category, queries the H-code defined by GHS from the CAS number of the chemical product's components. Each H-code details the toxic hazard description of the chemical component. A dictionary of 12 slots is defined from the toxic hazard description of the product's chemical components, with each slot taking an enumerated value. The slot names and enumerated values of the dictionary are shown in Table 2. By injecting additional knowledge, the model can better capture the contextual information of the label, providing additional semantic guidance to the model, thereby improving the model's predictive ability and generalization performance. For samples lacking H-codes, all 12 slot values are set to (no / none) as default while maintaining the same dimension. The slot values are obtained by extracting the textual description of the toxic hazard from the H-code corresponding to the chemical component of each chemical product through rule extraction. The text-to-dictionary rule extraction uses anchor points + hazard category light rules for parsing to obtain a structured dictionary. For example, for various acute toxicities caused by exposure, keywords (oral / dermal / inhal, etc.) are first captured as anchor points to determine the slots, then the category description is captured, and finally the slot value is obtained.
[0027] Table 2. Definitions of Toxicity Priors (from the dictionary)
[0028] To incorporate the toxicity and hazard information of the components as prior knowledge into the model, the hazard prior for each sample is extracted into a dictionary consisting of 12 discrete slots. For the first Each slot Pre-fixed ordered value table The first one This is the default state. The given value for this slot in the dictionary. Define its index If missing, then Then, a one-hot vector is constructed for each slot. The samples are then segmented and stitched together according to a fixed slot order to obtain a 44-dimensional sparse vector, i.e., the toxicity prior vector, expressed as: .
[0029] Preferably, the contextual semantic enhancement vector in the standardized labeled text samples is extracted using the Transformer model, and a cost-hazard association vector is constructed based on the semantic similarity between chemical components and hazardous entities, including: By using standardized labeled text samples and H-coding, the basic Transformer model is fine-tuned for a specific domain to obtain a specific model; Standardized labeled text samples are segmented into sentence units, and the sentence units are subjected to mean pooling using a specific model to obtain sentence-level semantic vectors; Extract the context window text of the harmful entity, and extract the entity center hidden state of the context window text through a specific model to obtain the entity-level semantic vector; The sentence-level semantic vector and the entity-level semantic vector are summed element-wise with weights, and then multiplied element-wise with the toxicity prior vector to obtain the context semantic enhancement vector. Both the CAS number and the harmful entity are embedded vectors, and the cosine similarity between the embedded vectors is calculated. A specific model is used to semantically match H-encodings with harmful entities to obtain a matching score; Based on pre-built dose-hazard intensity mapping rules, the association score between toxic dose and hazard severity is calculated through a specific model, and the semantic similarity is obtained by weighted summation of cosine similarity, matching score and association score. Calculate the semantic similarity of the harmful entities separately, and construct a similarity matrix based on the multiple semantic similarities; The similarity matrix is compressed and global max pooling is performed to obtain preliminary association features; The initial correlation features are weighted and summed with the preset differential weights, and the resulting weighted features are linearly transformed to obtain the cost-harm correlation vector.
[0030] In the specific process, step 300 selects BERT-base as the base Transformer model, uses standardized labeled text samples and H-coded toxicity descriptions corresponding to each chemical component as the domain fine-tuning dataset, sets the fine-tuning task as domain semantic alignment and entity boundary prediction, adopts the AdamW optimizer, and sets the learning rate to 3e. -5The batch size is 24, and the training runs for 6 epochs. Gradient descent is used to minimize cross-entropy loss, allowing the model to adapt to the SDS domain terminology and semantic features, resulting in a specific model to optimize the model's semantic understanding of chemical safety terminology and hazard description sentences. Subsequently, standardized labeled text samples are segmented into independent sentence units according to Chinese punctuation marks. The specific model is used to perform mean pooling on all word embedding vectors of each sentence unit to obtain sentence-level semantic vectors. Simultaneously, for each hazard entity, three words before and after the entity's initial word are extracted into a context window. This window text is encoded using the specific model, and the last Transformer hidden state corresponding to the entity's central word is extracted as the entity-level semantic vector. The sentence-level semantic vector and the entity-level semantic vector are element-wise weighted and summed with a preset weight of 0.4 for the sentence level and 0.6 for the entity level. This sum is then multiplied element-wise with the toxicity prior vector to obtain a context-enhanced semantic vector, achieving the fusion of domain semantics and toxicity prior knowledge. Next, the CAS number and the hazard entity are converted into 768-dimensional embedding vectors through the embedding layer of a specific model, and their cosine similarity is calculated. Semantic matching is then performed on the toxicity description text corresponding to the H-encoding and the hazard entity using the specific model, outputting a matching score in the 0-1 range; a higher score indicates a stronger semantic association. Finally, based on a pre-constructed dose-hazard intensity mapping rule, the rule explicitly defines LD... 50 LC 50 The relationship between isodose parameters and the severity of harm, such as LD. 50 <50mg / kg corresponds to a hazard severity level 5, LD50 50A dose >5000 mg / kg corresponds to a hazard severity level of 1. A specific model queries the corresponding hazard severity level based on the toxic dose value and converts it into an association score within the 0-1 range. Cosine similarity, matching score, and association score are weighted and summed with weights of 0.3, 0.4, and 0.3 respectively to obtain the semantic similarity between a single chemical component and a single hazard entity. Finally, for all chemical components and all hazard entities, semantic similarity is calculated separately, and an M (number of chemical components) × 6 (number of hazard entities) dimensional similarity matrix is constructed. This matrix is compressed into an M × 32 dimensional feature matrix using a 1 × 1 convolution kernel, and then global max pooling is applied to obtain a preliminary 32-dimensional association feature. An entity type attention weight is introduced, assigning differentiated weights to six types of hazardous entities: ToxicityDescription (0.25), ToxicologicalDose (0.2), Carcinogenicity and NonCarcinogenicity (0.15 each), and Symptom and TargetSite (0.125 each). The initial association features are multiplied element-wise with the attention weight vector, and then linearly transformed through a fully connected layer to obtain the component-hazard association vector. This ensures that its dimension is consistent with the toxicity prior vector, facilitating subsequent feature fusion.
[0031] Preferably, a bidirectional encoder is constructed based on an attention mechanism, and the toxicity prior vector, contextual semantic enhancement vector, and component-harm association vector are fused in a dual-path manner through the bidirectional encoder to obtain fused features, including: The toxicity prior vector and component are concatenated with the hazard association vector to obtain the fusion vector; The first feature is obtained by fusing the fused vector and the contextual semantic enhancement vector through an attention mechanism. The average negative log-likelihood score of the standardized annotated text samples is calculated. The average negative log-likelihood score is used as the sample difficulty score to classify the difficulty of the first feature. Based on the classification results, the parameters of the bidirectional encoder are adjusted to obtain the optimized encoder. By optimizing the encoder, the fused vector and the contextual semantic enhancement vector are fused to obtain fused features.
[0032] In the specific implementation process, step 400 assumes that the context semantic enhancement vector, after linear transformation, is... , Let H be the sequence length, H be the hidden layer dimension, and the fusion vector be... The expression for the fusion process includes: ; ; ; in, , , , , All of these are trainable parameters. The scaling factor for the attention head dimension. This is the first characteristic.
[0033] Furthermore, the formula for calculating the mean negative log-likelihood score is: ; in, For numerically stable terms, For a valid set of tokens, To predict probabilities. Using all samples... Based on the formed unimodal distribution, five levels—very simple, easy, medium, hard, and very hard—are obtained by segmenting at quantile points. After the training phase of each level, the average negative log-likelihood score is recalculated, and the bidirectional Transformer encoder parameters are adaptively adjusted according to the average negative log-likelihood score of the current level to obtain the optimized encoder. Specifically: when When <0.2, the number of rounds is 5, the batch size is 32, and the learning rate is 5e-5; when 0.2≤ When <0.4, the number of rounds is 6, the batch size is 24, and the learning rate is 3e-5; when 0.4 ≤ When <0.6, the number of rounds is 7, the batch size is 16, and the learning rate is 2e-5; when 0.6 ≤ When <0.8, the number of rounds is 8, the batch size is 12, and the learning rate is 1e-5; when When the value is ≥0.8, the number of rounds is 5, the batch size is 8, and the learning rate is 5e-6.
[0034] Furthermore, the recall of ToxicityDescription is further improved by suppressing the dominant effect of simple and high-frequency samples (especially "O" class) on the gradient in SDS sequence annotation. Additionally, a focus loss is introduced as the imbalance-aware objective function based on cross-entropy, and the optimized focus loss function expression of the encoder is: ; in, The softmax probability of the true class. To control the focusing intensity, Let Ω represent the class weights and Ω be the set of effective supervised locations. This loss function adaptively reduces the weights of correctly classified, high-confidence samples and enhances the gradient contributions of hard-classified and minority cases. The focus loss improves the learning efficiency of long-span toxicity descriptions and enhances the overall convergence stability of SDS hazard named entity recognition (NER).
[0035] Furthermore, the model's performance was evaluated using accuracy (P), recall (R), and F1 score, calculated using the following formulas: ; ; ; In this context, TP (True Positive) represents a true positive sample, which is a positive sample correctly predicted by the model, meaning that the true value is positive and the model also predicts it as positive; FP (False Positive) represents a false positive sample, which is a positive sample incorrectly predicted by the model, meaning that the true value is negative but the model predicts it as positive; and FN (False Negative) represents a false negative sample, which is a negative sample incorrectly predicted by the model, meaning that the true value is positive but the model predicts it as negative.
[0036] Precision is the ratio of identified entities to actual entities, while recall is the ratio of identified entities to predicted entities. F1 combines the values of precision and recall into a weighted harmonic mean to more accurately represent model performance.
[0037] Preferably, preliminary entity recognition is performed on the fused features using a Transformer model to obtain a BIO tag sequence, including: The labels of the fused features are extracted using the Transformer model to obtain a preliminary entity set; Boundary checks are performed on the initial entity set, and text corrections are made to the initial entity set based on the check results to obtain the corrected results; When the overlap rate of the corrected results is ≥80%, the entity boundary is updated and boundary verification is performed again. When the overlap rate of the corrected results is <80%, the corrected results are retained and false positive results that do not conform to the OSHA entity definition are removed, resulting in the BIO tag sequence.
[0038] In the specific implementation process, step 500 first inputs the fused features into a specific Transformer model that has been fine-tuned in the SDS domain. The model maps the fused features to the logits vector of each word corresponding to 13 types of BIO tags (B / I tags and O tags of 6 types of entities) through a fully connected layer, and then passes the softmax function. Convert to a label probability distribution, where Let be the logits value of the i-th tag and j be the tag category index. The tag with the highest probability for each lexical element is taken as the predicted tag for that lexical element. These lexical elements are concatenated in lexical order to form a preliminary tag sequence. Then, according to the BIO tagging rules (i.e., "B" is the entity's starting lexical element, "I" is the subsequent lexical element within the entity, and "O" is a non-entity lexical element), text fragments corresponding to consecutive "B-XXX" and "I-XXX" tag sequences are extracted, resulting in a preliminary entity set containing entity type, starting lexical element position, ending lexical element position, and entity text. Boundary validation is then performed. Based on the OSHA-defined six categories of hazardous entity terms and a domain dictionary, the cosine similarity between each entity text in the preliminary entity set and the corresponding category term in the dictionary is calculated. Entities with a similarity < 0.7 are considered to have ambiguous boundaries. A new window is then formed by expanding one lexical element before and after the original entity boundary, and the similarity between the text within the window and the domain term is recalculated. This expansion is repeated until the similarity is ≥ 0.7 or the maximum window length does not exceed twice the original entity length, thus completing text correction and obtaining the corrected entity set. Next, the overlap rate of any two entities of the same type in the corrected entity set is calculated using the following formula: ; in The number of reduplicated word elements. Let A be the number of word elements in entity A. Let B be the number of word elements, when When the percentage is ≥80%, it is determined to be a boundary misjudgment of the same entity. The two entities are merged, and the entity boundary is updated by taking the minimum start word position and the maximum end word position. Then, the boundary is re-validated on the updated entity. When the accuracy is less than 80%, all corrected entities are retained. At the same time, semantic matching is performed by referring to the OSHA entity definition standard and querying the OSHA official technical specification terminology database. False positive results that fail to match are eliminated. Finally, the remaining entities are mapped to the text according to the word position order to form a complete BIO tag sequence.
[0039] In the specific implementation process, step 600 first constructs a chemical component node using the CAS number of the chemical component as a unique identifier. The attributes include the CAS number, the specific values of the 12 slots in the GHS hazard statement obtained through H-coding, and the toxicity prior dictionary. Hazard entity nodes are constructed based on hazard entities extracted from BIO tag sequences. The attributes include entity type, entity text, the starting and ending lexical positions corresponding to the BIO tag, and the sentence unit index to which the entity belongs. Association attribute nodes are constructed based on the core features of the component-hazard association vector. The attributes include the weighted sum of semantic similarity, cosine similarity value, H-coding matching score, and the association score between dosage and hazard severity. Subsequently, the relationships between nodes are constructed. By parsing the component and hazard association vectors, edges are established between chemical component nodes and hazard entity nodes. The edge attributes are assigned the specific feature values of the component-entity pair in the association vector. Based on the contextual association of entities in the BIO tag sequence, edges are established between hazard entity nodes, and the edge attributes include the semantic similarity of the context window text. Relying on the mapping relationship between CAS number and H encoding, edges are established between chemical component nodes and associated attribute nodes. The edge attributes are the preliminary association feature values after similarity matrix compression, and are stored using the Neo4j graph database. Then, a 3D layout design is performed, allocating node positions according to the Z-axis layering rule: chemical component nodes are located at Z=0 layer, hazard entity nodes at Z=1 layer, and associated attribute nodes at Z=2 layer. The XY plane uses the ForceAtlas2 algorithm for layout, and the node spacing is adjusted according to the association strength of the edges; nodes with higher association strength are closer together. Finally, the completeness and accuracy of the knowledge graph are verified. All chemical component nodes are traversed to check whether each node is associated with the corresponding hazard entity node and whether the edge attributes are consistent with the component and hazard association vector. The type of hazard entity node is verified to be completely matched with the entity type in the BIO tag sequence. Weak associations with an association strength of <0.3 are removed by querying the OSHA entity definition standard to ensure that the three-dimensional knowledge graph can accurately present the three-dimensional relationship between chemical components, hazard entities and their association features.
[0040] like Figure 3 As shown, the present invention also provides a security data table-based hazard entity identification system based on Transformer, comprising: The text processing module is used to preprocess the raw heterogeneous data of the safety data table, and to annotate the hazard entities in the preprocessed raw heterogeneous data using the BIO annotation method to obtain standardized annotated text samples; hazard entities include: adverse effects, adverse reaction sites, carcinogenicity descriptions, non-carcinogenic claims, harmful effects, and toxicity doses; The prior construction module is used to determine the H-code of chemical components in standardized labeled text samples, construct a toxicity prior dictionary through the H-code, and encode the toxicity prior dictionary into a toxicity prior vector. The feature extraction module is used to extract contextual semantic enhancement vectors from standardized labeled text samples using the Transformer model, and to construct component-hazard association vectors based on the semantic similarity between chemical components and hazardous entities. The feature fusion module is used to construct a bidirectional encoder based on the attention mechanism, and then fuse the toxicity prior vector, the contextual semantic enhancement vector, and the component-harm association vector in a two-way manner through the bidirectional encoder to obtain the fused features. The entity recognition module is used to perform preliminary entity recognition on the fused features through the Transformer model to obtain the BIO tag sequence. The knowledge graph construction module is used to construct a three-dimensional knowledge graph based on the component-hazard association vector, the CAS number of the chemical component, and the BIO tag sequence.
[0041] The beneficial effects of this invention are as follows: 1) By incorporating prior knowledge of toxicity corresponding to the chemical component H encoding, and combining the weighted fusion of sentence-level and entity-level semantic enhancement vectors, a deep integration of domain priors and contextual semantics is achieved. Furthermore, the focus loss function is used to suppress high-frequency class interference, solving the problems of class imbalance, low annotation resources, and long-span complex entity recognition in SDS data, thereby improving the accuracy and reliability of hazardous entity recognition. 2) K-means clustering was used to classify heterogeneous templates, implement rule-driven automatic labeling and fuzzy boundary correction, replacing the traditional manual review and high-frequency maintenance mode. The standardized preprocessing process quickly converts unstructured / semi-structured SDS data into machine-readable format, which greatly reduces compliance verification costs and information extraction time, adapts to the governance needs of large-scale SDS text library, and improves SDS data processing efficiency and automation level. 3) Based on the OSHA standard, annotation rules and entity definitions were constructed to ensure that the recognition results meet industry compliance requirements. Through TF-IDF feature conversion, cosine similarity matching and terminology normalization, the problems of expression differences, terminology inconsistencies and template heterogeneity of SDS from different suppliers were solved, which improved the adaptability of the model in cross-supplier, long document and mixed text and image layout scenarios, and enhanced the model’s cross-scenario generalization ability. 4) A three-dimensional knowledge graph integrating chemical components (CAS number), hazardous entities (BIO tag sequence) and associated features (component-hazard association vector) was constructed. The three-dimensional layout was performed using Neo4j storage and the ForceAtlas2 algorithm, which intuitively presented the three-dimensional relationship between components, hazards and associated attributes. At the same time, the slot design of the toxicity prior dictionary made the model decision-making process traceable, which facilitated users to verify and query key hazard information, and realized the interpretability and visualization of the identification results. 5) By using similarity matrix compression, global max pooling, and differentiated weight allocation, the focus on long-tail entities such as non-carcinogenic claims and specific target organ toxicity is enhanced. Furthermore, the attention mechanism-driven dual-path feature fusion captures the weak semantic association between chemical components and hazardous entities, reducing the identification omissions caused by missing data fields and cross-document differences, and improving the ability to capture long-tail entities and weakly related information.
[0042] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0043] Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. Furthermore, those skilled in the art will recognize that, based on the ideas of this invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A method for identifying hazardous entities using a security data table based on Transformer, characterized in that, Includes the following steps: The raw heterogeneous data of the safety data sheet is preprocessed, and the hazard entities in the preprocessed raw heterogeneous data are labeled using the BIO labeling method to obtain standardized labeled text samples; the hazard entities include: adverse effects, adverse reaction sites, carcinogenicity descriptions, non-carcinogenicity claims, harmful effects, and toxicity doses; H-codes are determined based on the chemical composition in the standardized annotated text samples. A toxicity prior dictionary is constructed using the H-codes and encoded as a toxicity prior vector. The Transformer model is used to extract the contextual semantic enhancement vector from the standardized labeled text samples, and the component-hazard association vector is constructed based on the semantic similarity between the chemical component and the hazardous entity. A bidirectional encoder is constructed based on an attention mechanism, and the toxicity prior vector, the contextual semantic enhancement vector, and the component-harm association vector are fused in two ways through the bidirectional encoder to obtain fused features; The fused features are initially identified using the Transformer model to obtain a BIO tag sequence. A three-dimensional knowledge graph is constructed based on the component-hazard association vector, the CAS number of the chemical component, and the BIO tag sequence.
2. The method for identifying hazardous entities based on a Transformer-based security data table according to claim 1, characterized in that, The raw heterogeneous data of the safety data table is preprocessed, and the hazard entities in the preprocessed raw heterogeneous data are labeled using the BIO annotation method to obtain standardized annotation text samples, including: Based on the chapter framework specified by OSHA, the core chapter keywords of the original heterogeneous data are extracted. The original heterogeneous data is then classified into templates using the K-means clustering algorithm to obtain typical templates. The core chapter keywords are then mapped to standard field names based on the typical templates to obtain unified text. The unified text is subjected to meaningless data cleaning, format standardization, and terminology normalization to obtain an initial sample; The initial sample is filled with missing information and filtered for redundant information to obtain a data sample; Based on the hazardous entities, annotation rules are constructed, and the data samples are automatically annotated and fuzzy boundary entity corrected using the annotation rules to obtain the standardized annotated text samples.
3. The method for identifying hazardous entities based on a Transformer-based security data table according to claim 1, characterized in that, Based on the chemical composition in the standardized annotated text samples, H-codes are determined. A toxicity prior dictionary is constructed using these H-codes, and the toxicity prior dictionary is encoded into a toxicity prior vector, including: The H code can be obtained by looking up the CAS number of the chemical component; The toxicity prior dictionary is constructed based on the H encoding, and the dictionary slots are determined. The slot value of the dictionary slot is determined by the category description of the chemical component; An initial vector is constructed based on the dictionary slots and slot values, and the initial vector is segmented and concatenated according to the slot order to obtain the toxicity prior vector.
4. The method for identifying hazardous entities based on a Transformer-based security data table according to claim 1, characterized in that, The Transformer model is used to extract contextual semantic enhancement vectors from the standardized labeled text samples, and cost-hazard association vectors are constructed based on the semantic similarity between the chemical components and the hazardous entities, including: By using the standardized annotated text samples and the H-encoding, the basic Transformer model is fine-tuned for a specific domain to obtain a particular model. The standardized labeled text samples are segmented into sentence units, and the sentence units are subjected to mean pooling using the specific model to obtain sentence-level semantic vectors. Extract the context window text of the harmful entity, and extract the entity center hidden state of the context window text through the specific model to obtain the entity-level semantic vector; The sentence-level semantic vector and the entity-level semantic vector are summed element-wise with weights, and then multiplied element-wise with the toxicity prior vector to obtain the context semantic enhancement vector. Both the CAS number and the harmful entity are embedded vectors, and the cosine similarity between the embedded vectors is calculated. The H-encoding is semantically matched with the harmful entity using the specific model to obtain a matching score. Based on the pre-constructed dose-harm intensity mapping rules, the association score between the toxic dose and the severity of harm is calculated through the specific model, and the cosine similarity, the matching score and the association score are weighted and summed to obtain the semantic similarity; Calculate the semantic similarity of the harmful entities respectively, and construct a similarity matrix based on the multiple semantic similarities; The similarity matrix is subjected to matrix compression and global max pooling to obtain preliminary association features; The preliminary correlation features are weighted and summed with preset differential weights, and the resulting weighted features are linearly transformed to obtain the cost-hazard correlation vector.
5. The method for identifying hazardous entities based on a Transformer-based security data table according to claim 1, characterized in that, A bidirectional encoder is constructed based on an attention mechanism, and the toxicity prior vector, the contextual semantic enhancement vector, and the component-harm association vector are fused in a dual-path manner through the bidirectional encoder to obtain fused features, including: The toxicity prior vector and the component-hazard association vector are concatenated to obtain a fusion vector; The first feature is obtained by fusing the fusion vector and the context semantic enhancement vector through the attention mechanism. Calculate the average negative log-likelihood score of the standardized annotated text samples, use the average negative log-likelihood score as the sample difficulty score to classify the difficulty of the first feature, and adjust the parameters of the bidirectional encoder according to the classification result to obtain an optimized encoder. The optimized encoder performs feature fusion on the fusion vector and the contextual semantic enhancement vector to obtain the fused feature.
6. The method for identifying hazardous entities based on a Transformer-based security data table according to claim 5, characterized in that, The formula for calculating the average negative log-likelihood score is as follows: ;in, For numerically stable terms, For a valid set of tokens, To predict probabilities.
7. The method for identifying hazardous entities based on a Transformer-based security data table according to claim 5, characterized in that, The focus loss function expression of the optimized encoder is: ;in, For softmax probability, To control the focusing intensity, Ω represents the category weights, and Ω represents the set of effective supervised locations.
8. The method for identifying hazardous entities based on a Transformer-based security data table according to claim 1, characterized in that, The fused features are initially identified using the Transformer model to obtain a BIO tag sequence, including: The labels of the fused features are extracted using the Transformer model to obtain a preliminary entity set; Boundary verification is performed on the preliminary entity set, and text correction is performed on the preliminary entity set based on the verification results to obtain the correction result; When the overlap rate of the corrected results is ≥80%, the entity boundary is updated and boundary verification is performed again. When the overlap rate of the corrected results is <80%, the corrected results are retained and false positive results that do not conform to the OSHA entity definition are removed, thus obtaining the BIO tag sequence.
9. A security data table-based hazard entity identification system based on Transformer, characterized in that, include: The text processing module is used to preprocess the raw heterogeneous data of the safety data table, and to annotate the hazard entities in the preprocessed raw heterogeneous data using the BIO annotation method to obtain standardized annotated text samples; the hazard entities include: adverse effects, adverse reaction sites, carcinogenicity descriptions, non-carcinogenic claims, harmful effects, and toxicity doses; The prior construction module is used to construct a toxicity prior dictionary based on the determined H-code of the chemical components in the standardized labeled text sample and encode the toxicity prior dictionary into a toxicity prior vector. The feature extraction module is used to extract the contextual semantic enhancement vector from the standardized labeled text sample using the Transformer model, and to construct the component-hazard association vector based on the semantic similarity between the chemical component and the hazardous entity. The feature fusion module is used to construct a bidirectional encoder based on an attention mechanism, and to fuse the toxicity prior vector, the context semantic enhancement vector, and the component-harm association vector in a dual-path manner through the bidirectional encoder to obtain fused features; An entity recognition module is used to perform preliminary entity recognition on the fused features through the Transformer model to obtain a BIO tag sequence; The graph construction module is used to construct a three-dimensional knowledge graph based on the component-hazard association vector, the CAS number of the chemical component, and the BIO tag sequence.
Citation Information
Cited By
Vehicle fault named entity recognition method, device, equipment, medium and product
CN122221859A