A method, medium, and system for establishing a grid-based hazardous chemicals knowledge base
By using multi-source data fusion and grid processing technology, a grid-based hazardous chemicals knowledge base was constructed, which solved the problems of traditional knowledge bases relying on manual operation and insufficient knowledge correlation, and realized efficient and intelligent chemical safety management and risk identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSPECTION & QUARANTINE TECH CENT SHANDONG ENTRY EXIT INSPECTION & QUARANTINE BUREAU
- Filing Date
- 2025-06-06
- Publication Date
- 2026-05-26
AI Technical Summary
The existing hazardous chemicals knowledge base relies heavily on manual operations, and the expression of interrelationships between knowledge points is insufficient, resulting in inadequate accuracy and comprehensiveness in risk assessment and prevention.
A data matrix is formed by multi-source data fusion technology, an accuracy assessment model and a knowledge correlation matrix are constructed, a gridded knowledge mapping structure is established, conflicts are handled by Kruskal's algorithm, defective knowledge is handled by matrix completion algorithm, a knowledge reasoning mechanism and an adaptive optimization algorithm are designed to dynamically adjust the grid structure, and a gridded hazardous chemicals knowledge base is formed.
It has achieved efficient collection and integration of hazardous chemical data, significantly reduced manual operations, and solved the problem of insufficient expression of the correlation between knowledge through a multi-dimensional knowledge grid structure and automated processing mechanism, thereby improving the intelligence and comprehensiveness of chemical safety management and enhancing risk identification and prevention capabilities.
Smart Images

Figure CN120633795B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of knowledge base establishment methods, specifically, it relates to a gridded hazardous chemicals knowledge base establishment method, medium, and system. Background Technology
[0002] A hazardous chemicals knowledge base is a crucial infrastructure in the field of chemical safety management and risk control, providing knowledge support for safe production, accident prevention, and emergency response. Traditionally, the construction of hazardous chemicals knowledge bases primarily relies on manual collection and organization. Professionals consult relevant literature, manuals, experimental data, and accident reports to extract information such as the physicochemical properties, hazardous characteristics, storage and transportation requirements, and emergency response methods of chemicals. This information is then organized and categorized according to a specific structure to form a structured knowledge system. This expert-based knowledge base construction method has played a certain role in practical applications.
[0003] However, there are significant shortcomings in the construction of traditional hazardous chemical knowledge bases. First, the manual collection and organization of data consumes a large amount of human resources. With the continuous increase in the types of chemicals and the ongoing updates to safety standards, it is difficult to update and maintain the knowledge base in a timely manner by relying on manual operations. Second, existing knowledge bases generally use simple classification or hierarchical structures to organize knowledge, which makes it difficult to express the complex interactions and correlations of hazardous properties between chemicals. For example, the synergistic hazardous effects produced by mixing certain chemicals and the changes in hazards under specific environmental conditions are difficult to express and mine effectively.
[0004] The current technical problem of constructing a hazardous chemicals knowledge base, which relies heavily on manual operation and lacks sufficient expression of inter-knowledge relationships, severely restricts the accuracy and comprehensiveness of hazardous chemicals risk assessment and control. There is an urgent need for a new method that can reduce reliance on manual labor and enhance the ability to express knowledge relationships to overcome this technical bottleneck. In other words, existing technologies often rely heavily on manual operation when constructing hazardous chemicals knowledge bases, resulting in insufficient expression of inter-knowledge relationships. Summary of the Invention
[0005] In view of this, the present invention provides a method, medium and system for establishing a gridded hazardous chemical knowledge base, which can solve the technical problems in the prior art where the construction of a hazardous chemical knowledge base often relies on a large amount of manual operation and there is insufficient expression of the correlation between knowledge.
[0006] The present invention is implemented as follows: The first aspect of the present invention provides a method for establishing a gridded hazardous chemicals knowledge base, comprising: collecting hazardous chemicals data to establish an initial knowledge system; using multi-source data fusion technology to form a data matrix; and filtering data sources based on a recognition index; constructing an accuracy evaluation model; marking redundant knowledge, conflicting knowledge, erroneous knowledge, and defective knowledge to form a knowledge correlation matrix; establishing a gridded knowledge mapping structure; calculating gridded centrality to determine the weight of knowledge points; using the Kruskal algorithm to construct a minimum spanning tree to handle knowledge conflicts; applying a matrix completion algorithm to handle defective knowledge; constructing a chemical risk association network; calculating a hazardous chemicals grid index to quantify the hazard level; designing a knowledge reasoning mechanism; and using a gridded knowledge compression function for dimensionality reduction; and implementing an adaptive optimization algorithm to dynamically adjust the grid structure to form a gridded hazardous chemicals knowledge base.
[0007] The accuracy assessment model includes a basic accuracy calculation equation, a weight adjustment equation, a conflict penalty equation, an experimental verification adjustment equation, and a final accuracy synthesis equation. The basic accuracy calculation equation is used to calculate the initial accuracy value of the knowledge points. The final accuracy synthesis equation is used to calculate the final accuracy index by integrating various factors.
[0008] In establishing a gridded knowledge mapping structure, knowledge grid units are divided according to the properties of hazardous chemicals, and grid size is set to divide the knowledge granularity, forming a multi-dimensional knowledge grid space. Grid size refers to the granularity of knowledge grid division. Smaller sizes provide more refined knowledge representation, while larger sizes provide a more macroscopic knowledge overview.
[0009] Among them, grid centrality is an indicator describing the importance of knowledge points in the entire grid structure. Knowledge points with high centrality usually have more connections and stronger influence. In calculating grid clustering, grid clustering is an indicator that measures the density of knowledge points in a certain area and is used to identify key knowledge areas and blank areas.
[0010] Among them, the knowledge conflict degree matrix is a two-dimensional array that records contradictory or inconsistent information in the knowledge base, used to identify knowledge that needs to be addressed in a key manner; the edge weights represent the intensity of conflict between knowledge points; in the application of gridded dispersion analysis of knowledge distribution characteristics, gridded dispersion is an index describing the uniformity of knowledge points distributed in the grid space.
[0011] Among them, the Hazardous Chemicals Grid Index is a composite index that comprehensively considers the hazardous characteristics of chemicals and the completeness of knowledge, and is used to quantitatively assess the risk level of chemicals; it establishes mapping relationships between grid nodes to improve the accuracy of risk identification.
[0012] Among them, the grid knowledge compression function is used to reduce the dimensionality of ultra-long knowledge in the hazardous chemicals knowledge base and retain key information. The inputs include knowledge vector representation, grid centrality weight, knowledge correlation coefficient, compression ratio parameter and information retention importance threshold. The output is the compressed knowledge representation vector and information loss assessment index.
[0013] The chemical knowledge grid representation model has a multi-layer bidirectional transformer network architecture, which includes a hazardous chemical word embedding layer, a multi-head grid attention layer, a feedforward neural network layer, and a gridded representation output layer. The parameters of the gridded multi-head attention mechanism are jointly determined by the gridded centrality, the gridded size, and the knowledge association matrix.
[0014] A second aspect of the present invention provides a computer-readable storage medium storing program instructions, which, when executed in a computer, are used to perform the above-described method for establishing a gridded hazardous chemicals knowledge base.
[0015] A third aspect of the present invention provides a gridded hazardous chemicals knowledge base establishment system, comprising the aforementioned computer-readable storage medium, wherein the system is any one of a computer, a server, or a microcontroller, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.
[0016] This invention achieves efficient collection and integration of hazardous chemical data by introducing multi-source data fusion technology and automated data filtering mechanisms, significantly reducing reliance on manual operations. By constructing a multi-dimensional knowledge grid structure and a gridded representation model, it creatively solves the problem of insufficient expression of the correlations between hazardous chemical knowledge. Employing accuracy assessment models and knowledge correlation matrices, this invention achieves automated hierarchical analysis and quantification of correlation strength for knowledge points, enabling precise expression of complex relationships between knowledge. Through innovative indicators such as grid centrality and grid clustering, it captures the distribution characteristics and influence of knowledge points in the grid space, providing a systematic perspective for a comprehensive understanding of the hazardous characteristics of chemicals.
[0017] This method effectively solves the core technical problems of relying on a large amount of manual operation and insufficient expression of the correlation between knowledge in the traditional construction of hazardous chemical knowledge bases by constructing a gridded knowledge representation system and an automated processing mechanism. It provides more intelligent, systematic and comprehensive knowledge support for chemical safety management and greatly improves the ability to identify and control hazardous chemical risks. Attached Figure Description
[0018] Figure 1 This is a flowchart of the method of the present invention.
[0019] Figure 2 This is a comparison chart of the acceptance indices of various data sources in Example 2.
[0020] Figure 3 This is a comparison chart of the accuracy assessment of the knowledge points on hazardous chemicals in Example 2.
[0021] Figure 4 This is a grid-based centrality analysis diagram of key knowledge points in Example 2.
[0022] Figure 5 This is a distribution map of hazardous chemicals knowledge in the four-dimensional grid structure of Example 2.
[0023] Figure 6 This is the grid index and risk assessment diagram for hazardous chemicals in Example 2. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0025] like Figure 1 The diagram shown is a flowchart of a method for establishing a gridded hazardous chemicals knowledge base according to the first aspect of the present invention. This method includes the following steps:
[0026] S01. Collect data on hazardous chemicals to establish an initial knowledge system, use multi-source data fusion technology to form a data matrix, filter data sources according to the recognition index, calculate the recognition index matrix of each knowledge point and set a threshold to filter low-quality data.
[0027] S02. Construct an accuracy evaluation model, calculate the accuracy index of each knowledge point, perform hierarchical analysis on the knowledge points and mark redundant knowledge, conflicting knowledge, erroneous knowledge and defective knowledge, form a knowledge correlation matrix and quantify the correlation strength.
[0028] S03. Establish a gridded knowledge mapping structure, divide knowledge grid units according to the properties of hazardous chemicals, calculate grid centrality to determine the weight of knowledge points, set grid size to divide knowledge granularity, and form a multi-dimensional knowledge grid space.
[0029] S04. Use Kruskal's algorithm to construct a minimum spanning tree of knowledge to handle knowledge conflicts, construct a knowledge conflict degree matrix to measure the degree of conflict, apply a matrix completion algorithm to handle defective knowledge, calculate gridded clustering degree to identify knowledge-dense areas, and use edge weights to represent the intensity of conflict between knowledge points.
[0030] S05. Input the hazardous chemical attribute data and grid block information into a pre-trained chemical knowledge grid representation model to construct a chemical risk association network, calculate the hazardous chemical grid index to quantify the hazard level, apply grid dispersion analysis to analyze the knowledge distribution characteristics, establish mapping relationships between grid nodes, and improve the accuracy of risk identification.
[0031] S06. Design a knowledge reasoning mechanism that integrates logical rules and statistical methods. For ultra-long knowledge, use a grid knowledge compression function to reduce dimensionality, construct a hierarchical knowledge representation model to optimize the storage structure, and achieve knowledge simplification and key information retention.
[0032] S07. Implement an adaptive optimization algorithm based on a chemical knowledge grid representation model to dynamically adjust the grid structure, design a differentiated storage strategy according to the knowledge update frequency, establish a grid boundary fuzzy processing mechanism to solve the cross-grid knowledge mapping problem, realize intelligent adjustment of the grid structure, and obtain the final gridded hazardous chemical knowledge base.
[0033] The Acceptance Index is a quantitative indicator measuring the degree to which a piece of hazardous chemical knowledge is accepted by multiple authoritative data sources. Its value ranges from 0 to 1, with higher values indicating higher acceptance. The Accuracy Index is a quantitative value of the precision of a knowledge point, calculated through experimental verification and theoretical models, used to screen high-quality knowledge points. The Knowledge Relevance Matrix is a two-dimensional array representing the degree of interconnection between hazardous chemical knowledge points; each element represents the strength of the association between two knowledge points. The Knowledge Conflict Matrix is a two-dimensional array recording contradictory or inconsistent information in the knowledge base, used to identify conflicting knowledge that needs to be addressed. Grid Centrality describes the importance of a knowledge point within the overall grid structure; knowledge points with high centrality typically have more connections and stronger influence. Grid Size refers to the granularity of the knowledge grid division; smaller sizes provide finer knowledge representation, while larger sizes provide a more macroscopic overview. Grid Clustering measures the density of knowledge points within a region, used to identify key knowledge areas and blank areas. Grid Dispersion describes the uniformity of knowledge point distribution within the grid space; high dispersion indicates uniform knowledge distribution, while low dispersion indicates uneven distribution. The Hazardous Chemicals Grid Index is a composite index that comprehensively considers the hazardous characteristics of chemicals and the completeness of knowledge about them, and is used to quantitatively assess the risk level of chemicals.
[0034] The grid knowledge compression function is used to reduce the dimensionality of extremely long knowledge in the hazardous chemicals knowledge base and retain key information. The inputs include the knowledge vector representation obtained from S03, the grid centrality weight calculated from S03, the knowledge correlation coefficient extracted from the knowledge correlation matrix formed in S02, the manually set compression ratio parameter, and the information retention importance threshold determined based on knowledge importance analysis. The output is the compressed knowledge representation vector used for the construction of the hierarchical knowledge representation model in S06 and the information loss assessment index used for the knowledge base performance evaluation in S09.
[0035] The specific structure of the chemical knowledge grid representation model is a multi-layer bidirectional transformer network architecture, which includes a hazardous chemical word embedding layer, a multi-head grid attention layer, a feedforward neural network layer, and a gridded representation output layer. The gridded multi-head attention mechanism parameters are jointly determined by the gridded centrality calculated in S03, the gridded size set in S03, and the knowledge association matrix formed in S02. The number of grid attention heads matches the grid partitioning dimension, and each attention head captures the association patterns of different types of chemical hazardous characteristics.
[0036] The specific steps in establishing the training dataset during the training process of the chemical knowledge grid representation model include collecting international hazardous chemicals safety data sheets, hazardous chemical accident case reports, chemical structure database information, and hazardous chemical descriptions from professional journal literature; constructing triplet knowledge representations; generating knowledge graph node representations; designing a knowledge grid mapping matrix; constructing positive and negative sample pairs; forming a gridded knowledge training corpus; designing training task types; and dividing the training set, validation set, and test set.
[0037] The specific steps for training the chemical knowledge grid representation model include initializing model parameters, designing a knowledge mask recovery task to train grid representation capabilities, constructing a grid node prediction task to enhance the model's understanding of the grid structure, designing a hazardous chemical attribute association prediction task to improve risk identification capabilities, using a contrastive learning framework to optimize the gridded representation space, combining knowledge distillation technology to integrate expert knowledge, using a distributed training system to accelerate the training process, evaluating model performance through a validation set, employing an early stopping strategy to prevent overfitting, saving training weights and grid mapping parameters, and constructing a model deployment package for subsequent applications.
[0038] The accuracy evaluation model includes the basic accuracy calculation equation, the weight adjustment equation, the conflict penalty equation, the experimental verification adjustment equation, and the final accuracy synthesis equation.
[0039] The accuracy basic calculation equation is used to calculate the initial accuracy value of the knowledge point. The input includes the knowledge point recognition index matrix obtained from S01, the knowledge point data source reliability coefficient, the knowledge point release time decay factor, and the average expert score of the knowledge point. The output is the basic accuracy index of the knowledge point.
[0040] The weight adjustment equation is used to adjust the accuracy assessment weights according to the importance of knowledge points. The inputs include the frequency of knowledge point citation, the coverage of knowledge point application scenarios, the number of knowledge points associated with knowledge points, and the risk level coefficient of knowledge points. The output is the knowledge point weight adjustment coefficient.
[0041] The conflict penalty equation is used to reduce the accuracy score of conflicting knowledge points. The inputs include the number of conflicts between knowledge points, the severity coefficient of the conflict, the average accuracy of conflicting knowledge points, and the difficulty coefficient of conflict resolution. The output is the knowledge point conflict penalty value.
[0042] The experimental verification adjustment equation is used to correct the accuracy based on the experimental data. The inputs include the number of experimental verifications, the experimental data consistency coefficient, the experimental condition strictness score, the experimental method reliability coefficient, and the experimental result standard deviation. The output is the experimental verification adjustment coefficient.
[0043] The final accuracy synthesis equation is used to calculate the final accuracy index by integrating various factors. The inputs include the basic accuracy index of knowledge points, the weight adjustment coefficient of knowledge points, the conflict penalty value of knowledge points, and the experimental verification adjustment coefficient. The output is the final accuracy index of knowledge points, which is used for knowledge point hierarchical analysis and labeling in S02.
[0044] The specific implementation methods of the above steps are described in detail below.
[0045] The specific implementation of step S01 involves first collecting hazardous chemical data from multiple sources, including the National Hazardous Chemicals Database, International Chemical Safety Cards, and academic journal databases. Then, a hash table data structure is used to organize the multi-source data, and data fusion technology is employed to merge the same knowledge point from different sources into a unified representation, forming... 3D data matrix, where Indicates the number of knowledge points. This indicates the number of data sources. Then, based on factors such as the data source's authority, update frequency, data completeness, accuracy, and historical records, a recognition index is calculated for each data source, ranging from 0 to 1. Data sources with a threshold of 0.6 or higher are typically considered reliable. Finally, a recognition index matrix is calculated for each knowledge point. A weighted average algorithm is used to synthesize the support level of each data source for the knowledge point. Knowledge points with a recognition index below 0.5 are filtered out, retaining high-quality data for the next step. This step aims to ensure the data quality and reliability of the initial knowledge system, laying the foundation for subsequent knowledge base construction.
[0046] The specific implementation of step S02 involves first constructing an accuracy assessment model, which includes five core equations. The basic accuracy calculation equations employ a multi-factor weighted summation method, incorporating the acceptance index, data source reliability coefficient (typically 0.1–0.9), and time decay factor (typically...). The system takes λ (0.1–0.3) and expert scores (out of 1–10) as inputs and outputs a basic accuracy value. The weighting adjustment equation generates adjustment coefficients through a non-linear combination of citation frequency (normalized to 0–1), application scenario coverage (usually a percentage), the number of related knowledge items, and a risk level coefficient (1–8). The conflict penalty equation calculates a penalty value based on the number of conflicts, severity (0–1), the mean accuracy of conflicting knowledge, and a resolution difficulty coefficient (1–5). Experimental validation of the adjustment equations considers the number of experiments, data consistency coefficient (0.5–1), condition strictness (1–10), method reliability (0.6–0.95), and result standard deviation to calculate adjustment coefficients. The final accuracy synthesis equation weights and combines the aforementioned outputs to generate a final accuracy index. Then, a hierarchical clustering algorithm is used to classify and analyze knowledge points. Redundant knowledge is marked based on semantic similarity and content overlap. Conflicting knowledge is discovered through logical relationship verification of knowledge points. Erroneous knowledge is identified based on the accuracy index, and defective knowledge is discovered through integrity assessment. Finally, a knowledge association matrix is constructed, and cosine similarity and association rule mining algorithms are used to quantify the association strength between knowledge points, with values ranging from 0 to 1. Knowledge point pairs with an association strength greater than 0.7 are typically considered strongly associated. The purpose of this step is to establish a scientific knowledge evaluation system and improve the quality of the knowledge base.
[0047] The specific implementation of step S03 involves first defining a multidimensional knowledge space based on the physicochemical properties, hazard categories, usage conditions, and storage requirements of hazardous chemicals. A grid partitioning algorithm is then used to divide the knowledge space into regular units. Next, the grid centrality of each knowledge point is calculated using the eigenvector centrality algorithm, considering the number of connections, connection strength, and accuracy index of the knowledge point. The centrality value ranges from 0 to 1, with knowledge points having a centrality greater than 0.8 typically considered core knowledge. Then, grid size parameters are set, generally selecting a 3- to 8-dimensional grid structure based on data scale and application requirements. Smaller sizes (e.g., 3×3×3) provide fine-grained representations, while larger sizes (e.g., 8×8×8) provide a macroscopic overview. Finally, a multidimensional knowledge grid space is constructed, using tensor representation to store knowledge information from each dimension, forming a complete gridded knowledge mapping structure, enabling multidimensional organization and rapid retrieval of knowledge. This step aims to establish a structured knowledge organization method, providing a spatial framework for subsequent conflict resolution and risk association.
[0048] The specific implementation of step S04 involves first constructing a knowledge point relationship graph, where nodes represent knowledge points, edges represent relationships between knowledge points, and edge weights represent the intensity of conflict between knowledge points. Next, the Kruskal algorithm is used to construct a minimum spanning tree. This algorithm selects the knowledge connection method with the least conflict, effectively avoiding cyclic conflicts and is suitable for handling knowledge conflicts in complex networks. Then, a knowledge conflict degree matrix is constructed, where matrix element values represent the degree of conflict between knowledge points, ranging from 0 to 1. Knowledge point pairs with a conflict degree greater than 0.6 are typically marked as severely conflicting and processed first. For deficient knowledge, a matrix completion algorithm based on the low-rank assumption is applied. Missing information is inferred from the relationships between existing knowledge points. For low-confidence completion results (confidence threshold usually set to 0.7), the missing label is retained for verification. Finally, the gridded clustering degree is calculated, and the Moran index method from spatial statistics is used to identify knowledge-dense and sparse regions. Regions with a clustering degree greater than 0.8 are identified as knowledge-dense regions, and regions with a clustering degree less than 0.2 are identified as knowledge-blank regions. The purpose of this step is to resolve knowledge conflicts, improve the knowledge structure, and enhance the consistency of the knowledge base.
[0049] The specific implementation of step S05 involves first inputting the physicochemical property data of hazardous chemicals (such as flash point, boiling point, toxicity indicators, etc.) and grid segmentation information (from step S03) as input feature vectors into a pre-trained chemical knowledge grid representation model. This model employs a multi-layer bidirectional transformer architecture, capable of capturing complex nonlinear relationships between chemical properties. Next, a chemical risk association network is constructed, using graph neural network technology to establish possible reactions, synergistic toxicity, and cascade risk relationships between hazardous chemicals. The association strength threshold is typically set to 0.65. Then, a hazardous chemical grid index is calculated. This index comprehensively considers the inherent hazard of the chemical (weight 0.6) and knowledge completeness (weight 0.4) to quantify the hazard level. The index ranges from 0 to 10, where 8–10 represents extremely high hazard, 6–8 represents high hazard, 4–6 represents moderate hazard, 2–4 represents low hazard, and 0–2 represents slight hazard. Finally, a grid dispersion analysis method is applied, using an entropy calculation model to assess the uniformity of knowledge distribution. A dispersion greater than 0.7 indicates uniform knowledge distribution, while a dispersion less than 0.3 indicates extremely uneven knowledge distribution, requiring supplementation of relevant domain knowledge. Finally, a mapping relationship is established between grid nodes, and tensor mapping functions are used to connect knowledge points from different dimensions to form a complete knowledge network. This step aims to create a risk association system for hazardous chemicals and improve the accuracy of risk identification.
[0050] The specific implementation of step S06 involves first designing a knowledge reasoning mechanism that integrates rule-based reasoning based on first-order predicate logic and statistical reasoning methods based on Bayesian networks to achieve both precise and uncertain reasoning for knowledge about hazardous chemicals. Next, for extremely long knowledge content (typically referring to knowledge points with a dimensionality exceeding 1000 or a text length exceeding 5000 characters), a grid-based knowledge compression function is applied for dimensionality reduction. This compression function combines principal component analysis and an autoencoder. Inputs include knowledge vector representation, grid centrality weights (typically 0.1–0.9), knowledge relevance coefficients (typically 0.1–0.8), compression ratio parameters (typically set to 0.3–0.7), and information importance thresholds (typically 0.6). The compression function outputs a compressed knowledge representation vector (dimensionality typically reduced by 50%–80%) and an information loss assessment index (typically required to be below 0.2). Then, a hierarchical knowledge representation model is constructed, using a tree structure to organize knowledge. Upper-level nodes store abstract concepts, and lower-level nodes store specific details, optimizing the storage structure. Finally, knowledge simplification and key information retention are achieved by ranking knowledge based on importance and retaining key information with an importance index greater than 0.7. The purpose of this step is to optimize knowledge representation and storage efficiency, and to ensure the integrity and accessibility of key information.
[0051] The specific implementation of step S07 involves first implementing an adaptive optimization algorithm based on a chemical knowledge grid representation model. This algorithm combines gradient descent and simulated annealing strategies, enabling dynamic adjustment of grid structure parameters according to knowledge updates. The algorithm's convergence condition is typically set to a grid structure change rate of less than 0.01 for five consecutive iterations or reaching the maximum number of iterations (100). Next, a differentiated storage strategy is designed based on knowledge update frequency: knowledge points with high update frequency (e.g., updated more than once per month) are stored in the fast access layer; knowledge points with medium update frequency (e.g., updated once per quarter) are stored in the standard access layer; and knowledge points with low update frequency (e.g., updated less than once per year) are stored in the archive layer. Then, a grid boundary fuzzy processing mechanism is established, employing fuzzy set theory to handle cross-grid knowledge mapping problems. A membership threshold, typically between 0.4 and 0.6, is set to resolve the fuzzy attribution of knowledge at grid boundaries. Finally, intelligent grid structure adjustment is implemented. Based on knowledge base usage feedback and the characteristics of newly added knowledge, grid parameters, including grid dimension, grid size, and grid connectivity, are automatically optimized to form the final gridded hazardous chemical knowledge base. This step aims to improve the adaptability and scalability of the knowledge base, enabling dynamic optimization and continuous improvement.
[0052] The chemical knowledge grid representation model employs a multi-layer bidirectional transformer network architecture, specifically comprising four core layers: a hazardous chemical term embedding layer using a 300-dimensional vector space to represent chemical terms and attributes, employing a domain-adaptive pre-training method to capture chemical semantics; a multi-head grid attention layer containing eight attention heads, each corresponding to a different hazardous characteristic dimension (such as flammability, toxicity, reactivity, etc.), with attention weights determined by grid centrality, grid size, and knowledge association matrix; a two-layer feedforward neural network layer with a hidden layer dimension of 1024, using the GELU activation function to process the attention layer output; and a grid representation output layer generating a 512-dimensional gridded knowledge representation while outputting grid index information. The model employs skip connections and layer normalization mechanisms to enhance training stability, with a total of approximately 45 million parameters.
[0053] The steps for establishing the training dataset for the chemical knowledge grid representation model are as follows: First, data sources are collected, including International Data on Hazardous Chemicals (ICSC cards), hazardous chemical accident case reports (covering major accidents over the past 20 years), chemical structure database information (such as PubChem data), and descriptions of hazardous chemicals from professional journal articles (collecting no fewer than 5000 articles). Next, a triplet knowledge representation is constructed, converting the collected unstructured data into a "subject-relationship-object" format, such as "...". "Possesses - strong corrosiveness", generating at least 500,000 triples. Knowledge graph node representations are then generated, and the TransE and ComplEx algorithms are used to map the triples to a vector space. A knowledge grid mapping matrix is designed to establish the mapping relationship between knowledge graph nodes and the grid space. Positive and negative sample pairs are constructed, with positive samples being related chemical pairs and negative samples being randomly sampled unrelated chemical pairs, in a ratio of 1:3. A gridded knowledge training corpus is formed, containing texts such as descriptions of hazardous chemical properties, safe operating procedures, and emergency response measures. Training task types are designed, including mask prediction, grid relationship prediction, and hazardous attribute classification tasks. Finally, the dataset is divided into training, validation, and test sets in a 7:1.5:1.5 ratio to ensure a balanced distribution of chemical categories in each set.
[0054] A second aspect of the present invention provides a computer-readable storage medium storing program instructions, which, when executed in a computer, are used to perform the above-described method for establishing a gridded hazardous chemicals knowledge base.
[0055] A third aspect of the present invention provides a gridded hazardous chemicals knowledge base establishment system, comprising the aforementioned computer-readable storage medium, wherein the system is any one of a computer, a server, or a microcontroller, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.
[0056] The mathematical model or calculation process involved in this invention will be described in detail below.
[0057] The calculation process of the approval index matrix in step S01 can be specifically represented as follows:
[0058] ;
[0059] In the formula, For knowledge points In the data source The recognition index in China; For data source The reliability coefficient ranges from 0 to 1. For knowledge points In the data source The existence state in the data is 1 if it exists and 0 if it does not exist; For data source The time factor indicates the timeliness of data source updates.
[0060] The formula for calculating the overall recognition index is as follows:
[0061] ;
[0062] In the formula, For knowledge points The overall recognition index; For data source Weighting coefficients; This represents the total number of data sources.
[0063] The parameter acquisition method is as follows:
[0064] The data can be obtained by evaluating its authority, historical accuracy, and credibility. The analytic hierarchy process (AHP) can be used to calculate the result, which is obtained after multiple experts (no fewer than 5) score the data source. Through function Calculation, where This is the time decay coefficient, typically taken as 0.1 to 0.3; The current time; The last update time of the data source; The Delphi method, typically determined by an expert panel, is used to comprehensively assess factors such as the frequency of data source usage, coverage, and data integrity.
[0065] The accuracy evaluation model equations in step S02 are specifically expressed as follows:
[0066] 1. Basic equation for accuracy calculation:
[0067] ;
[0068] In the formula, For knowledge points The basic accuracy index ranges from 0 to 1. For knowledge points The recognition index; For knowledge points The reliability coefficient of the data source ranges from 0.1 to 0.9. For knowledge points The difference between the release time and the current time (in years); This is the time decay factor, with a value ranging from 0.1 to 0.3; For knowledge points The average score from experts, with a maximum score of 10. These are the weight coefficients of each factor, and they satisfy... .
[0069] 2. Weighting adjustment equation:
[0070] ;
[0071] In the formula, For knowledge points The weighting adjustment coefficient; For knowledge points The citation frequency (normalized to 0-1); For knowledge points Application scenario coverage (percentage value); For knowledge points The amount of related knowledge; For knowledge points The risk level coefficient (1-8); These are the weight coefficients of each factor, and they satisfy... .
[0072] 3. Conflict Penalty Equation:
[0073] ;
[0074] In the formula, For knowledge points Conflict penalty value; For knowledge points The number of conflicts; The severity coefficient of the conflict (0-1); To be related to knowledge points Average accuracy of conflicting knowledge points; The difficulty level of conflict resolution is rated as 1-5. These are the weight coefficients of each factor, and they satisfy... .
[0075] 4. Experimental verification of the adjustment equation:
[0076] ;
[0077] In the formula, For knowledge points The experimental verification adjustment coefficient; The number of experimental verifications; The coefficient of consistency of experimental data is 0.5 to 1. The experimental conditions were rated on a scale of 1 to 10. The reliability coefficient of the experimental method is (0.6–0.95). The standard deviation of the experimental results; These are the weight coefficients of each factor, and they satisfy... .
[0078] 5. Final accuracy synthesis equation:
[0079] ;
[0080] In the formula, For knowledge points The final accuracy index; Basic accuracy index; This is the weighting adjustment factor; This is the conflict penalty value; To verify the adjustment coefficient in the experiment; This is a smoothing factor, typically ranging from 0.01 to 0.1.
[0081] The parameter acquisition method is as follows:
[0082] Based on a questionnaire survey of domain experts, it is usually determined that... ; Determined based on the importance of the application scenario, typically ; Determined based on the assessment of the degree of conflict impact, usually ; Determined based on experimental reliability experience, typically .
[0083] The accuracy evaluation model employs a non-linear weighted combination approach, considering multiple influencing factors. A logarithmic relationship is used for citation frequency to reflect the principle of diminishing marginal utility; a square relationship is used for application scenario coverage to emphasize the importance of comprehensive coverage; a square root relationship is used for the quantity of related knowledge to express the diminishing value resulting from the increase of related knowledge; a logistic function is used to handle the number of experiments, reflecting the decrease in gain after a certain number of experimental verifications; conflict penalty is handled with non-linear normalization to avoid extreme penalties; and the final accuracy uses a multiplicative combination to ensure the comprehensive influence of all factors, and a smoothing factor is introduced to prevent calculation anomalies caused by zero values.
[0084] The formula for calculating the knowledge association matrix in step S02 is as follows:
[0085] ;
[0086] In the formula, Representing knowledge points With knowledge points The strength of the association between them is calculated using the following formula:
[0087] ;
[0088] In the formula, Knowledge point vector and Cosine similarity; This refers to the support function in association rule mining. For knowledge points and The number of times they co-occur; The maximum number of co-occurrences among all knowledge point pairs; Let be the weight coefficient, and satisfy... .
[0089] The parameter acquisition method is as follows:
[0090] Knowledge point vector The knowledge point text is converted into a 300- to 768-dimensional vector using a text embedding model (such as Word2Vec or BERT); Calculate knowledge points using the Apriori algorithm. and The probability of them occurring simultaneously; This was obtained by counting the co-occurrence frequency of the two knowledge points in all documents; The importance of each factor is determined based on its significance in practical applications, and is usually... .
[0091] The formula for calculating the grid centrality in step S03 is as follows:
[0092] ;
[0093] In the formula, For knowledge points The grid centrality; For knowledge points and The strength of the association; For knowledge points The final accuracy index; For knowledge points The degree of connectivity (the number of knowledge points directly connected to it); The maximum connectivity among all knowledge points; For knowledge points The betweenness centrality; It represents the maximum betweenness centrality among all knowledge points.
[0094] The parameter acquisition method is as follows:
[0095] Obtain from the knowledge association matrix in step S02; Obtain from the accuracy evaluation model output in step S02; By calculating the knowledge relevance matrix and its relation to knowledge points The number of knowledge points with a correlation strength greater than a threshold (usually 0.3) is obtained; Calculated using the following formula: ,in For knowledge points arrive The number of shortest paths, For the knowledge points of arrive The number of shortest paths.
[0096] The grid-based centrality calculation employs a combination of weighted average, logarithmic normalization, and linear enhancement to reflect the importance of a knowledge point's position within the knowledge network. Weighted average considers the quality of adjacent knowledge points, logarithmic normalization addresses connectivity and reduces the impact of extreme values, and the betweenness centrality enhancement term emphasizes the "bridging" role of nodes within the network.
[0097] The formula for calculating the knowledge conflict degree matrix in step S04 is as follows:
[0098] ;
[0099] In the formula, Representing knowledge points With knowledge points The degree of conflict between them is calculated using the following formula:
[0100] ;
[0101] In the formula, Knowledge point vector and semantic similarity; The strength of logical conflict is determined by predicate logic rules. The degree of numerical conflict represents the degree of inconsistency in numerical knowledge. The maximum numerical conflict level among all knowledge point pairs; Let be the weight coefficient, and satisfy... .
[0102] The parameter acquisition method is as follows:
[0103] It is obtained by calculating the cosine similarity or dot product of knowledge point vectors; The rule engine detects logical contradictions between knowledge points. If there is a complete contradiction, the value is 1; if there is a partial contradiction, the value is 0.1 to 0.9. For numerical knowledge, through formulas Calculation, where and For the corresponding numerical value; Based on the importance of conflict type, it is usually determined that... .
[0104] The matrix completion algorithm used in step S04 is implemented as follows:
[0105] The matrix completion model for defective knowledge can be represented as:
[0106] ;
[0107] In the formula, The complete matrix to be completed; This is an observation matrix containing missing values; For projection operators, corresponding to the set of observed entries. ; For matrix The nuclear norm (the sum of all singular values); This is a regularization parameter used to balance fitting error and matrix complexity; it is typically set to 0.1 to 1.
[0108] The formula for calculating the confidence level of the completed knowledge points is as follows:
[0109] ;
[0110] In the formula, To complete the value Confidence level; These are the predicted values in cross-validation; The standard deviation of the prediction error; To be related to knowledge points The number of related missing knowledge points; This represents the total number of knowledge points in the knowledge base.
[0111] The parameter acquisition method is as follows:
[0112] Known elements in the matrix come directly from the knowledge base, while unknown elements are marked as missing. The optimal value is selected using cross-validation. The model is trained on the training set and predicted on the validation set by randomly dividing the known data into a training set and a validation set. It is obtained by calculating the standard deviation of the difference between the predicted value and the true value on the validation set.
[0113] The matrix completion algorithm is based on the low-rank assumption, which posits the existence of a potential low-dimensional structure between knowledge points. It encourages low-rank solutions by minimizing the nuclear norm while maintaining consistency with the observed data. The confidence score calculation combines prediction error and the proportion of relevant missing knowledge points; the former reflects the model's certainty regarding a specific completed value, while the latter considers the completeness of information surrounding the knowledge point.
[0114] The gridd clustering degree calculation in step S04 uses Moran's I index, as shown in the following formula:
[0115] ;
[0116] In the formula, This is the Moran index, which typically ranges from -1 to 1. This represents the total number of knowledge points. These are elements of the spatial weight matrix, reflecting the knowledge points. and Distance relationships in grid space; For knowledge points The attribute values (such as importance or access frequency); This is the average value of all knowledge point attribute values.
[0117] Formula for calculating local grid clustering:
[0118] ;
[0119] In the formula, For knowledge points The degree of clustering in the local area.
[0120] The parameter acquisition method is as follows:
[0121] Through formula Calculation, where For knowledge points and Euclidean distance in grid space This is the distance threshold, typically set to twice the length of the grid cell diagonal; These could be indicators of the importance of knowledge points (such as grid centrality), accuracy index, or access frequency.
[0122] The Moran index is used to quantify spatial autocorrelation, reflecting the degree of clustering of knowledge distribution patterns in a grid space. Positive values indicate clustering of similar values (knowledge-dense areas), negative values indicate clustering of dissimilar values (such as a mixture of knowledge gaps and dense distributions), and values close to zero indicate random distribution. Local clustering calculations identify the knowledge density characteristics around each grid cell, facilitating the location of specific knowledge-dense areas.
[0123] The formula for calculating the hazardous chemicals grid index in step S05 is as follows:
[0124] ;
[0125] In the formula, Chemicals The grid risk index ranges from 0 to 10. It is the inherent hazard index of chemicals, assessed based on GHS classification standards, with a value range of 0 to 10; This is a knowledge completeness index, representing the degree of coverage of relevant knowledge points; The completeness threshold parameter is typically set to 3 to 5; Let be the weight coefficient, and satisfy... .
[0126] Formula for calculating knowledge completeness index:
[0127] ;
[0128] In the formula, In order to be with chemicals A collection of relevant knowledge points; For knowledge points The final accuracy index; For knowledge points The grid centrality; For set The number of elements in the middle.
[0129] The parameter acquisition method is as follows:
[0130] Based on chemical hazard characteristics indicators (such as flash point, LD50, corrosivity, etc.), each hazard category is converted into a score of 0 to 10 using the GHS classification method, and the highest score is taken as the inherent hazard index. By searching the knowledge base for chemicals All knowledge points with a correlation (correlation strength greater than 0.3) are obtained; Determined based on risk management strategy, typically This reflects the dominant role of inherent hazards in risk assessment.
[0131] The grid risk index calculation employs a combination of linear combination and Sigmoid transformation. The linear combination preserves the direct impact of inherent hazards, while the Sigmoid function handles the contribution of knowledge completeness, making the risk assessment of chemicals with insufficient knowledge more conservative. The knowledge completeness index design considers both the quality (accuracy and centrality) and quantity (logarithmic scaling) of relevant knowledge, and the logarithmic transformation reflects the diminishing marginal utility of knowledge quantity.
[0132] The grid dispersion calculation in step S05 uses the entropy method, and the formula is as follows:
[0133] ;
[0134] In the formula, This represents the grid dispersion, with a value ranging from 0 to 1. This represents the total number of grid cells. For the first The normalized density of knowledge points in each grid cell is calculated using the following formula: ,in For the first Number of knowledge points in each grid cell.
[0135] Formula for calculating dimensional dispersion:
[0136] ;
[0137] In the formula, For dimension Dispersion on; For dimension The number of grid cells on; For dimension Upper Normalized knowledge density of each unit.
[0138] The parameter acquisition method is as follows:
[0139] This was obtained by counting the number of knowledge points in each grid cell. The mesh size is determined based on the meshing parameters set in step S03, for example, a 3×3×3 mesh corresponds to... ; The number of grid divisions in each dimension is the same as the gridding size parameter.
[0140] Grid-based dispersion is based on the concept of information entropy, quantifying the uniformity of knowledge distribution. A higher entropy value indicates a more uniform distribution, while a lower entropy value indicates that knowledge is concentrated in a few grid cells. Normalization processing (divided by) This makes the dispersion of grids of different sizes comparable. Dimensional dispersion calculation provides detailed information on the characteristics of knowledge distribution in each independent dimension, which helps to identify uneven knowledge distribution in certain dimensions.
[0141] The grid knowledge compression function in step S06 can be expressed as:
[0142] ;
[0143] The formula for calculating the compressed vector is as follows:
[0144] ;
[0145] Formula for calculating information loss assessment index:
[0146] ;
[0147] In the formula, For knowledge points The original vector representation; For gridded centrality weights; This is the knowledge relevance coefficient; This is the compression ratio parameter; This is the information importance threshold; This is the compressed knowledge representation vector; As an indicator for assessing information loss; To Perform PCA dimensionality reduction to Dimensional operations, among which ; for An identity matrix of order 1; This represents the average grid centrality of all knowledge points. This is an important feature enhancement matrix, defined by features whose importance exceeds a threshold. Characteristic components; For knowledge points Importance index, expressed by formula calculate.
[0148] The parameter acquisition method is as follows:
[0149] Obtained from step S03, typically a 300- to 768-dimensional vector; Obtained from the grid centrality calculated in step S03; Extracted from the knowledge association matrix formed in step S02, the calculation formula is as follows: ; The compression ratio parameter is manually set, typically between 0.3 and 0.7; The threshold for the importance of retained information, determined based on knowledge importance analysis, is typically 0.6. The matrix was obtained through error analysis of the autoencoder reconstruction of the original vectors. Features with larger reconstruction errors were considered more important, and the top features were selected. One is considered an important feature, with an enhancement coefficient of 0.5 to 1 at the corresponding position, and 0 at the other positions.
[0150] The grid knowledge compression function combines PCA dimensionality reduction and important feature preservation mechanisms. PCA dimensionality reduction provides the basic compression effect, while the feature enhancement matrix adjusts the degree of key information retention based on the importance of knowledge points. Information loss assessment is based on reconstruction error and knowledge importance; important knowledge is allowed a lower information loss rate, while general knowledge can accept a higher compression rate. The compression ratio parameter controls the overall compression intensity, and the threshold parameter determines the boundaries of information that needs to be prioritized for retention.
[0151] The adaptive optimization algorithm in step S07 can be expressed as:
[0152] ;
[0153] In the formula, For the first The mesh structure parameters for the next iteration; For the updated mesh structure parameters; The learning rate is typically set to 0.01–0.05. The gradient of the objective function with respect to the grid structure; This is the evaluation function value for the current mesh structure; For the first The temperature parameter for the next iteration is calculated using the following formula: ,in The initial temperature (usually set to 1), This is the cooling rate (usually set to 0.95–0.99).
[0154] The objective function is defined as:
[0155] ;
[0156] In the formula, Knowledge coverage of the grid structure; For retrieval efficiency metrics; For grid complexity; This is the mesh structure from the previous version; Let be the weight coefficient, and satisfy... .
[0157] The parameter acquisition method is as follows:
[0158] The objective function is calculated by approximating the calculation using the finite difference method and then calculating the change of the objective function after a small perturbation of each grid parameter. This is obtained by calculating the proportion of knowledge points that can be quickly retrieved in the grid structure out of the total number of knowledge points; Calculated by averaging the reciprocal of the average response time of multiple random retrieval requests; The complexity of mesh connectivity is calculated using a weighted sum of factors, including the number of dimensions. Determined based on actual application requirements, typically .
[0159] The adaptive optimization algorithm combines gradient descent and simulated annealing strategies. Gradient descent provides the optimization direction, while simulated annealing increases the possibility of escaping local optima. The objective function is designed to balance knowledge coverage, retrieval efficiency, structural complexity, and structural stability. The last term uses the sigmoid function to smoothly handle the penalty for structural changes, avoiding excessively drastic structural adjustments.
[0160] The mesh boundary fuzzing mechanism in step S07 adopts fuzzy set theory and can be expressed as:
[0161] ;
[0162] In the formula, For knowledge points For grid cells Membership degree; For knowledge points To grid cell center The distance; The fuzzy radius parameter is typically set to 0.5 to 1 times the side length of the grid cell. This is a shape parameter that controls the rate of decrease in membership degree; it is typically set to 2.
[0163] Cross-grid knowledge mapping calculation formula:
[0164] ;
[0165] In the formula, For knowledge points In grid cells The representation in; For knowledge points The original vector representation; To be compatible with grid cells Adjacent grid cell sets; For knowledge points With grid cells The average correlation strength of all knowledge points in the text.
[0166] The parameter acquisition method is as follows:
[0167] It is obtained by calculating the Euclidean distance between the feature vector of the knowledge point and the center vector of the grid cell in the feature space; Determined based on the size of the grid cell; if the side length of the grid cell is... ,but Usually set to ; The value is determined through fuzzy boundary experiments and is usually set to 2. Extracted from the knowledge association matrix, the calculation formula is as follows: ,in For grid cells Number of knowledge points.
[0168] The grid boundary fuzzing mechanism uses an improved Bell-shaped membership function, which provides a smoother boundary transition compared to traditional triangular or trapezoidal membership functions. The cross-grid knowledge mapping formula combines the original representation and the influence of adjacent grids, ensuring a natural transition of knowledge at the boundary through membership weighting, avoiding blind spots in knowledge retrieval caused by hard boundaries.
[0169] Specifically, the principle of this invention is as follows: Based on gridded knowledge representation theory and multi-source data fusion technology, this invention expresses and organizes hazardous chemical knowledge by constructing a multi-dimensional grid structure, thereby reducing reliance on manual operation and enhancing the expression of knowledge relevance. Its core lies in transforming traditional linear or hierarchical knowledge structures into gridded structures, enabling hazardous chemical knowledge to form an interconnected network in a multi-dimensional space. Each grid cell represents a set of knowledge under a certain attribute combination, and the connections between grids represent the relationships between knowledge points.
[0170] In terms of reducing manual operations, this invention achieves automatic data source filtering through an acceptance index matrix. By calculating the degree to which each knowledge point is recognized by multiple authoritative data sources, an objective data quality assessment system is established, replacing the traditional filtering process that relies on manual judgment. The accuracy assessment model automatically calculates the accuracy index of knowledge points through five key equations and can automatically mark redundant, conflicting, erroneous, and defective knowledge, significantly reducing manual intervention.
[0171] In enhancing the expression of knowledge relevance, the innovation of this invention lies in constructing a complete grid characteristic measurement system. By calculating grid centrality, the importance and influence of knowledge points within the entire grid structure are determined; by setting grid size, knowledge representation at different granularities is achieved; and by calculating grid clustering and dispersion, knowledge-dense areas and knowledge distribution characteristics are identified. These indicators collectively constitute the quantitative basis for expressing knowledge relevance, enabling the accurate capture of complex chemical hazard characteristic association patterns.
[0172] The chemical knowledge grid representation model of this invention adopts a multi-layer bidirectional transformer network architecture. It automatically learns the correlations of different types of chemical properties through a multi-head grid attention mechanism. The Kruskal algorithm and matrix completion algorithm automate the handling of knowledge conflicts and defects. The grid knowledge compression function and adaptive optimization algorithm further enhance the system's processing capabilities, enabling the knowledge base to dynamically adjust the grid structure according to knowledge updates. This represents a dual breakthrough in both theoretical and technical aspects, reducing reliance on manual intervention and enhancing knowledge correlation.
[0173] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.
[0174] The specific implementation of step S01 involves first collecting hazardous chemical data from multiple sources, including the National Hazardous Chemicals Database, International Chemical Safety Cards, and academic journal databases, using web crawling technology, database interfaces, and document parsing tools. Next, a hash table data structure is used to organize the multi-source data, and data fusion technology is employed to merge the same knowledge point from different sources into a unified representation, forming... 3D data matrix, where Indicates the number of knowledge points. This indicates the number of data sources. Then, based on factors such as data source authority, update frequency, data integrity, accuracy, and historical records, a credibility index is calculated for each data source, ranging from 0 to 1. Data sources with a threshold of 0.6 or higher are typically considered reliable data sources. The credibility index matrix is calculated using the following formula: In the formula, For knowledge points In the data source The recognition index in China; For data source The reliability coefficient ranges from 0 to 1. For knowledge points In the data source The existence state in the data is 1 if it exists and 0 if it does not exist; For data source The time factor represents the timeliness of data source updates, and the calculation formula is: ,in The time decay coefficient is typically taken as 0.1 to 0.3. Finally, the overall acceptance index for each knowledge point is calculated using a weighted average algorithm: In the formula, For knowledge points The overall recognition index; For data source Weighting coefficients; This represents the total number of data sources. Knowledge points with an acceptance index below 0.5 are filtered out, retaining only high-quality data for the next step. This step aims to ensure the data quality and reliability of the initial knowledge system, laying the foundation for subsequent knowledge base construction.
[0175] The specific implementation of step S02 involves first constructing an accuracy assessment model, which comprises five core equations. The basic accuracy calculation equations employ a multi-factor weighted summation method: In the formula, For knowledge points The basic accuracy index ranges from 0 to 1. For knowledge points The recognition index; For knowledge points The reliability coefficient of the data source ranges from 0.1 to 0.9. For knowledge points The difference between the release time and the current time (in years); This is the time decay factor, with a value ranging from 0.1 to 0.3; For knowledge points The average score from experts, with a maximum score of 10. These are the weight coefficients of each factor, and they satisfy... ,generally The weighting adjustment equation generates adjustment coefficients through nonlinear combination: In the formula, For knowledge points The weighting adjustment coefficient; For knowledge points Citation frequency (normalized to 0-1); For knowledge points Application scenario coverage (percentage value); For knowledge points The amount of related knowledge; For knowledge points The risk level coefficient (1-8); These are the weight coefficients of each factor, and they satisfy... ,generally The conflict penalty equation calculates the penalty value based on multiple factors: In the formula, For knowledge points Conflict penalty value; For knowledge points The number of conflicts; The severity coefficient of the conflict (0-1); To be related to knowledge points Average accuracy of conflicting knowledge points; Difficulty level of conflict resolution (1-5); These are the weight coefficients of each factor, and they satisfy... ,generally Experimental verification of the adjustment equation and calculation of adjustment coefficients based on multiple factors: In the formula, For knowledge points The experimental verification adjustment coefficient; The number of experimental verifications; The coefficient of consistency of experimental data (0.5 to 1); Score the stringency of the experimental conditions (1-10); The reliability coefficient of the experimental method is (0.6–0.95). The standard deviation of the experimental results; These are the weight coefficients of each factor, and they satisfy... ,generally The final accuracy synthesis equation combines the aforementioned outputs in a weighted manner: In the formula, For knowledge points The final accuracy index; Basic accuracy index; This is the weighting adjustment factor; This is the conflict penalty value; To verify the adjustment coefficient in the experiment; The smoothing factor is typically set between 0.01 and 0.1. Next, a hierarchical clustering algorithm is used to classify and analyze the knowledge points. Redundant knowledge is marked based on semantic similarity and content overlap. Conflicting knowledge is discovered through verification of logical relationships between knowledge points. Erroneous knowledge is identified based on accuracy indices, and deficient knowledge is discovered through completeness assessment. Finally, a knowledge association matrix is constructed. ,in Representing knowledge points With knowledge points The strength of the association between them is calculated using the following formula: In the formula, Knowledge point vector and Cosine similarity; This refers to the support function in association rule mining. For knowledge points and The number of times they co-occur; The maximum number of co-occurrences among all knowledge point pairs; Let be the weight coefficient, and satisfy... ,generally The purpose of this step is to establish a scientific knowledge evaluation system and improve the quality of the knowledge base.
[0176] The specific implementation of step S03 involves first defining a multi-dimensional knowledge space based on the physicochemical properties, hazard categories, usage conditions, and storage requirements of hazardous chemicals, and then dividing the knowledge space into regular units using a grid partitioning algorithm. Next, the grid centrality of each knowledge point is calculated, specifically using the eigenvector centrality algorithm. In the formula, For knowledge points The grid centrality; For knowledge points and The strength of the association; For knowledge points The final accuracy index; For knowledge points Connectivity (the number of knowledge points directly connected to it); The maximum connectivity among all knowledge points; For knowledge points Betweenness centrality; The centrality is determined by the maximum betweenness among all knowledge points. The centrality value ranges from 0 to 1, with knowledge points having a centrality greater than 0.8 typically considered core knowledge. Next, the grid size parameters are set, generally choosing a 3- to 8-dimensional grid structure based on data scale and application requirements. Smaller sizes (e.g., 3×3×3) provide fine-grained representations, while larger sizes (e.g., 8×8×8) provide a macroscopic overview. Finally, a multi-dimensional knowledge grid space is constructed, using tensor representation to store knowledge information from each dimension, forming a complete gridded knowledge mapping structure, enabling multi-dimensional organization and rapid retrieval of knowledge. This step aims to establish a structured knowledge organization method, providing a spatial framework for subsequent conflict resolution and risk association.
[0177] The specific implementation of step S04 involves first constructing a knowledge point relationship graph, where nodes represent knowledge points, edges represent relationships between knowledge points, and edge weights represent the conflict intensity between knowledge points. Next, the Kruskal algorithm is used to construct a minimum spanning tree. This algorithm selects the knowledge connection method with the least conflict, effectively avoiding cyclic conflicts and is suitable for handling knowledge conflicts in complex networks. Then, a knowledge conflict degree matrix is constructed. ,in Representing knowledge points With knowledge points The degree of conflict between them is calculated using the following formula: In the formula, Knowledge point vector and semantic similarity; The strength of logical conflict is determined by predicate logic rules. The degree of numerical conflict represents the degree of inconsistency in numerical knowledge. The maximum numerical conflict level among all knowledge point pairs; Let be the weight coefficient, and satisfy... ,generally For defective knowledge, a matrix completion algorithm based on the low-rank assumption is applied for processing, and its model can be expressed as: In the formula, The complete matrix to be completed; This is an observation matrix containing missing values; For projection operators, corresponding to the set of observed entries. ; For matrix The nuclear norm (the sum of all singular values); This is the regularization parameter, used to balance fitting error and matrix complexity, typically taken as 0.1 to 1. The formula for calculating the confidence score of the completed knowledge point is: In the formula, To complete the value Confidence level; These are the predicted values in cross-validation; The standard deviation of the prediction error; To be related to knowledge points The number of related missing knowledge points; This represents the total number of knowledge points in the knowledge base. Knowledge point pairs with a conflict score greater than 0.6 are typically marked as severely conflicting and processed first. Finally, the grid clustering degree is calculated using the Moran index method from spatial statistics. In the formula, This is the Moran index, which typically ranges from -1 to 1. This represents the total number of knowledge points. These are elements of the spatial weight matrix, reflecting the knowledge points. and Distance relationships in grid space; For knowledge points Attribute values (such as importance or access frequency); This is the average of all knowledge point attribute values. The formula for calculating local grid clustering is: In the formula, For knowledge points The clustering degree of the local area. Areas with a clustering degree greater than 0.8 are identified as knowledge-intensive areas, and areas with a clustering degree less than 0.2 are identified as knowledge-deficient areas. The purpose of this step is to resolve knowledge conflicts, improve the knowledge structure, and enhance the consistency of the knowledge base.
[0178] The specific implementation of step S05 involves first inputting the physicochemical property data of hazardous chemicals (such as flash point, boiling point, toxicity indicators, etc.) and grid segmentation information (from step S03) as input feature vectors into a pre-trained chemical knowledge grid representation model. This model employs a multi-layer bidirectional transformer architecture, capable of capturing complex nonlinear relationships between chemical properties. Next, a chemical risk association network is constructed, using graph neural network technology to establish possible reactions, synergistic toxicity, and cascade risk relationships between hazardous chemicals; the association strength threshold is typically set to 0.65. Then, the hazardous chemical grid index is calculated. In the formula, Chemicals The grid risk index ranges from 0 to 10. It is the inherent hazard index of chemicals, assessed based on GHS classification standards, with a value range of 0 to 10; This is a knowledge completeness index, representing the degree of coverage of relevant knowledge points; The completeness threshold parameter is typically set to 3 to 5; Let be the weight coefficient, and satisfy... ,generally The formula for calculating the knowledge completeness index is: In the formula, In order to be with chemicals A collection of relevant knowledge points; For knowledge points The final accuracy index; For knowledge points The grid centrality; For set The number of elements. The index ranges from 0 to 10, where 8 to 10 indicates extremely high risk, 6 to 8 indicates high risk, 4 to 6 indicates moderate risk, 2 to 4 indicates low risk, and 0 to 2 indicates slight risk. Next, a gridded dispersion analysis method is applied, using an entropy calculation model: In the formula, This represents the grid dispersion, with a value ranging from 0 to 1. This represents the total number of grid cells. For the first The normalized density of knowledge points in each grid cell is calculated using the following formula: ,in For the first The number of knowledge points in each grid cell. The formula for calculating dimensional dispersion is: In the formula, For dimension Dispersion on; For dimension The number of grid cells on; For dimension Upper The normalized knowledge density of each unit is calculated. A dispersion greater than 0.7 indicates uniform knowledge distribution, while a dispersion less than 0.3 indicates highly uneven knowledge distribution, requiring supplementation with relevant domain knowledge. Finally, a mapping relationship is established between grid nodes, using tensor mapping functions to connect knowledge points of different dimensions, forming a complete knowledge network. This step aims to form a risk association system for hazardous chemicals and improve the accuracy of risk identification.
[0179] The specific implementation of step S06 involves first designing a knowledge reasoning mechanism that integrates rule-based reasoning based on first-order predicate logic and statistical reasoning methods based on Bayesian networks to achieve both precise and uncertain reasoning for knowledge about hazardous chemicals. Then, for extremely long knowledge content (typically referring to knowledge points with a dimension exceeding 1000 or a text length exceeding 5000 characters), a grid-based knowledge compression function is applied for dimensionality reduction. The formula for calculating the compressed vector is: The formula for calculating the information loss assessment index is as follows: In the formula, For knowledge points The original vector representation; For gridded centrality weights; This is the knowledge relevance coefficient; This is the compression ratio parameter; This is the information importance threshold; This is the compressed knowledge representation vector; As an indicator for assessing information loss; To Perform PCA dimensionality reduction to Dimensional operations, among which ; for An identity matrix of order 1; This represents the average grid centrality of all knowledge points. Enhancement matrix for important features; For knowledge points Importance index, expressed by formula The compression function employs a combination of principal component analysis and an autoencoder. Inputs include knowledge vector representation, grid centrality weights (typically 0.1–0.9), knowledge relevance coefficients (typically 0.1–0.8), compression ratio parameters (typically set to 0.3–0.7), and information importance thresholds (typically 0.6). The compression function outputs a compressed knowledge representation vector (dimensionality typically reduced by 50%–80%) and an information loss assessment index (typically required to be below 0.2). A hierarchical knowledge representation model is then constructed, organizing knowledge using a tree structure. Upper-level nodes store abstract concepts, while lower-level nodes store specific details, optimizing the storage structure. Finally, knowledge simplification and key information retention are achieved by ranking knowledge importance and retaining key information with an importance index greater than 0.7. The purpose of this step is to optimize knowledge representation and storage efficiency, ensuring the integrity and accessibility of key information.
[0180] The specific implementation of step S07 is to first implement an adaptive optimization algorithm based on a chemical knowledge grid representation model: In the formula, For the first The mesh structure parameters for the next iteration; For the updated mesh structure parameters; The learning rate is typically set to 0.01–0.05. The gradient of the objective function with respect to the grid structure; This is the evaluation function value for the current mesh structure; For the first The temperature parameter for the next iteration is calculated using the following formula: ,in This is the initial temperature (usually set to 1). The cooling rate is typically set to 0.95–0.99. The objective function is defined as: In the formula, Knowledge coverage of the grid structure; For retrieval efficiency metrics; For grid complexity; This is the mesh structure from the previous version; Let be the weight coefficient, and satisfy... ,generally This algorithm combines gradient descent and simulated annealing strategies, enabling dynamic adjustment of mesh structure parameters based on knowledge updates. The convergence condition is typically set to a mesh structure change rate of less than 0.01 for five consecutive iterations or reaching the maximum number of iterations (100). Next, a differentiated storage strategy is designed based on knowledge update frequency: knowledge points with high update frequency (e.g., updated more than once a month) are stored in the fast access layer, knowledge points with medium update frequency (e.g., updated once a quarter) are stored in the standard access layer, and knowledge points with low update frequency (e.g., updated less than once a year) are stored in the archive layer. Then, a mesh boundary fuzzy processing mechanism is established, employing fuzzy set theory. In the formula, For knowledge points For grid cells Membership degree; For knowledge points To grid cell center The distance; The fuzzy radius parameter is typically set to 0.5 to 1 times the side length of the grid cell. The shape parameter controls the rate of decrease in membership degree, typically set to 2. The formula for calculating cross-mesh knowledge mapping is: In the formula, For knowledge points In grid cells The representation in; For knowledge points The original vector representation; To be compatible with grid cells Adjacent grid cell sets; For knowledge points With grid cells The average association strength of all knowledge points is calculated. A membership threshold, typically between 0.4 and 0.6, is set to address the issue of ambiguous attribution of knowledge at grid boundaries. Finally, intelligent adjustment of the grid structure is implemented. Based on feedback from knowledge base usage and the characteristics of newly added knowledge, grid parameters, including grid dimension, grid size, and grid connectivity, are automatically optimized to form the final gridded hazardous chemicals knowledge base.
[0181] To better understand and implement this invention, Example 2, a specific application scenario, is provided below: A chemical industrial park houses over 200 types of hazardous chemicals. Researchers in the park decided to establish a grid-based hazardous chemical knowledge base to improve safety management. First, the researchers collected 2650 hazardous chemical knowledge points from eight data sources, including the National Hazardous Chemicals Database, the US NIOSH Chemical Hazards Database, the EU ECHA Chemicals Database, the China Material Safety Data Sheets (MSDS) system, chemical industry journal articles, and accident case reports.
[0182] In the first step, the researchers calculated the acceptance index for each data source, and the results are shown in Table 1:
[0183] Table 1. Evaluation Results of Hazardous Chemicals Data Source Acceptance Index
[0184]
[0185] Figure 2 The chart presents the acceptance index assessment results of eight hazardous chemical data sources in the embodiment. The horizontal axis represents the different data source names, and the vertical axis represents the values of each indicator. Each data source has four indicators: reliability coefficient (CR₍ⱼ₎), time factor (TF₍ⱼ₎), weight coefficient (W₍ⱼ₎), and the final acceptance index. A red dashed line marks the acceptance threshold of 0.5; data sources below this threshold are excluded. The chart shows that the US NIOSH Chemical Hazard Database has the highest acceptance index at 0.85, while safety training materials have the lowest at 0.48 and are excluded from the knowledge base. This chart visually demonstrates the quality assessment results of each data source in the first step of the embodiment. After screening, knowledge sources with an acceptance index below 0.5 were excluded, retaining a total of seven data sources. For each knowledge point, researchers calculated a comprehensive acceptance index and filtered out those with an acceptance index below 0.5, ultimately retaining 2245 high-quality knowledge points.
[0186] In the second step, the researchers constructed an accuracy assessment model. Taking benzene and toluene, two common hazardous chemicals, as examples, they calculated the final accuracy index, and some results are shown in Table 2:
[0187] Table 2. Accuracy Assessment Results of Some Hazardous Chemicals Knowledge Points
[0188]
[0189] Figure 3 The data is presented in radar chart format, showing the properties of benzene (…). ) and toluene ( The accuracy assessment results for six knowledge points of two common hazardous chemicals are shown in the figure. The figure contains five curves, representing the basic accuracy, normalized weight adjustment coefficient, amplified conflict penalty value, amplified experimental verification adjustment coefficient, and final accuracy index. The radar chart format allows for a visual comparison of the performance of different knowledge points across various assessment dimensions. For example, K0135 (flash point of benzene) has the highest final accuracy index of 0.94, while K0216 (LD50 value of toluene) is relatively lower. This chart corresponds to the accuracy assessment model section in the second step of the embodiment. Simultaneously, researchers constructed a knowledge correlation matrix to determine the correlation strength between knowledge points. For example, the correlation strength between the flash point of benzene and its explosion hazard knowledge point is 0.82, and its correlation strength with the storage conditions knowledge point is 0.75.
[0190] Thirdly, the researchers designed a 4D grid structure based on four dimensions: "physicochemical properties," "hazard category," "storage requirements," and "emergency response." Each dimension was divided into five levels, forming a 5×5×5×5 grid space with a total of 625 grid cells. The grid centrality of key knowledge points was calculated, and some results are shown in Table 3.
[0191] Table 3 Results of Grid Centrality Calculation for Key Knowledge Points
[0192]
[0193] Figure 4 This is a bubble scatter plot showing the grid centrality analysis results for five key knowledge points. The horizontal axis represents betweenness centrality, the vertical axis represents the sum of association strengths, the bubble size represents connectivity, and the color intensity represents grid centrality. A red dashed line is also added to the plot to represent a trend line, revealing the linear relationship between betweenness centrality and the sum of association strengths. The plot shows that K0311 (sulfuric acid) has the highest grid centrality (0.92) and the largest betweenness centrality (1562.7), indicating that this knowledge point plays a crucial role in the grid structure. This chart corresponds to the grid centrality calculation section in step three of the embodiment. Figure 5 This is a 3D scatter plot illustrating the distribution of hazardous chemical knowledge within a four-dimensional grid structure. The three axes represent the physical and chemical properties dimension, the hazard category dimension, and the storage requirements dimension, each divided into five levels. The color and size of the points represent the values in the fourth dimension—emergency response. The plot also marks the locations of five high-risk chemicals: benzene (… ), acetylene ( Sodium cyanide (NaCN), sulfuric acid ( ) and toluene ( This 3D visualization clearly shows the distribution characteristics of hazardous chemicals in a 4D grid space, especially the tendency of high-risk chemicals to concentrate in specific areas of the grid space. This diagram corresponds to the 4D grid structure designed in step three of the embodiment.
[0194] In the fourth step, the researchers applied Kruskal's algorithm to construct a minimum spanning tree of knowledge, addressing 82 sets of conflicting knowledge points. For example, the hazardous properties of nitric acid conflicted across different data sources; the most reliable description was determined by calculating the conflict degree matrix. Simultaneously, a matrix completion algorithm was used to process 157 deficient knowledge points, such as partial physical property data of certain novel composite hazardous chemicals. The researchers calculated the grid-based clustering degree, identifying a knowledge density of 0.86 for the flammable and explosive chemicals region, indicating high clustering; while the knowledge density for the novel nanomaterials region was only 0.23, indicating significant knowledge gaps.
[0195] In the fifth step, researchers input the hazardous chemical attribute data and grid segmentation information into a pre-trained chemical knowledge grid representation model to construct a chemical risk association network. Using 36 high-risk chemicals within the park as the core, a hazardous chemical grid index was calculated, and some results are shown in Table 4.
[0196] Table 4 Calculation results of grid index for some hazardous chemicals
[0197]
[0198] Using gridded dispersion analysis, the overall knowledge dispersion was calculated to be 0.65, which is at a moderately uniform level. The dispersion of the physicochemical properties dimension was 0.82, while the dispersion of the emergency response dimension was only 0.48, indicating that emergency knowledge is unevenly distributed. Figure 6 This is a bubble scatter plot showing the grid risk index analysis results for eight major hazardous chemicals. The horizontal axis represents the inherent hazard index, the vertical axis represents the knowledge completeness, the bubble size represents the grid risk index, and the color represents the risk level (red indicates extremely high risk, orange indicates high risk). The plot divides the entire area into four quadrants, marked with different colored backgrounds: yellow represents low risk-high knowledge, red represents high risk-high knowledge, blue represents low risk-low knowledge, and purple represents high risk-low knowledge (the most dangerous). From the plot, it can be seen that benzene (… Sodium chlorate has the highest grid risk index at 8.82, classifying it as extremely high-risk; while sodium chlorate ( The grid risk index for [the chemical] is low, at 7.46, classifying it as a high-risk level. This chart corresponds to the hazardous chemical grid index calculation result in step five of the embodiment.
[0199] In the sixth step, the researchers designed a knowledge reasoning mechanism, employing a grid-based knowledge compression function to reduce the dimensionality of extremely long knowledge entries. Taking the text describing the physicochemical safety properties of ammonium nitrate (original length 8520 characters) as an example, after processing with the compression function, a knowledge representation vector with a 72% reduction in dimensionality was obtained, with an information loss assessment index of 0.17, better than the set threshold of 0.2. A hierarchical knowledge representation model was used to optimize the storage structure. The top layer contains abstract concepts such as "flammable and explosive," "toxic," and "corrosive," refining each layer to specific parameters of particular substances, achieving knowledge simplification while retaining key information.
[0200] In the seventh step, the researchers implemented an adaptive optimization algorithm based on a chemical knowledge grid representation model to dynamically adjust the grid structure. Initially, the algorithm's learning rate was set to 0.03, the initial temperature to 1.0, and the cooling rate to 0.97. After 83 iterations, the grid structure change rate decreased to 0.008, reaching the convergence condition. Based on the knowledge update frequency, the researchers designed a three-layer storage strategy: a fast access layer (updated monthly, 265 knowledge points), a standard access layer (updated quarterly, 1582 knowledge points), and an archive layer (updated annually, 398 knowledge points). A grid boundary fuzziness processing mechanism was established, setting the membership threshold to 0.5, thus solving the cross-grid knowledge mapping problem. The final gridded hazardous chemicals knowledge base contains 625 grid cells, 2245 knowledge points, and a total storage capacity of approximately 850MB.
[0201] Since the implementation of this grid-based hazardous chemicals knowledge base, the park's safety management efficiency has significantly improved. Chemical risk assessment time has been reduced from an average of 4.5 hours to 0.8 hours, and the identification accuracy rate has increased from 85% to 96.5%. Emergency response plan generation time has been reduced from an average of 28 minutes to 5 minutes, and it can provide precise guidance for complex multi-chemical mixture accidents.
[0202] Traditional hazardous chemical knowledge management primarily employs relational databases or simple knowledge graphs, which suffer from inconsistent data quality, difficulty in handling knowledge conflicts, limited expression of relationships, and low knowledge retrieval efficiency. Traditional methods, relying solely on simple hierarchical classification and keyword indexing, struggle to handle complex relationships between chemicals and cannot adapt to the demands of dynamic knowledge updates. The gridded hazardous chemical knowledge base establishment method of this invention offers the following advancements compared to traditional approaches: First, it establishes a systematic knowledge quality assessment system, ensuring knowledge quality through dual screening using both recognition and accuracy indices; second, it employs a multi-dimensional grid structure to organize knowledge, overcoming the limitations of traditional planar knowledge organization and enabling rapid retrieval and reasoning from multiple perspectives; third, it innovatively introduces the Kruskal algorithm and matrix completion algorithm, effectively resolving knowledge conflicts and deficiencies; fourth, it comprehensively characterizes the importance and distribution characteristics of knowledge through the calculation of three indicators: grid centrality, clustering, and dispersion; and fifth, it designs an adaptive optimization algorithm and boundary fuzzy processing mechanism, achieving dynamic optimization and continuous improvement of the knowledge base, significantly enhancing the level of chemical risk management.
[0203] It should be noted that the variables involved in this invention are explained in detail in Tables 5, 6, and 7 below.
[0204] Table 5. Variable Explanation Table (Part 1)
[0205]
[0206] Table 6. Variable Explanation Table (Part Two)
[0207]
[0208] Table 7. Variable Explanation Table (Part 3)
[0209]
[0210] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for establishing a grid-based hazardous chemicals knowledge base, characterized in that, include: Collect data on hazardous chemicals to establish an initial knowledge system, use multi-source data fusion technology to form a data matrix, and filter data sources based on the acceptance index; Construct an accuracy evaluation model, label redundant knowledge, conflicting knowledge, erroneous knowledge, and defective knowledge, and form a knowledge correlation matrix; Establish a gridded knowledge mapping structure and calculate the gridded centrality to determine the weights of knowledge points; A knowledge point relationship graph is constructed, where nodes represent knowledge points, edges represent relationships between knowledge points, and edge weights represent the conflict intensity between knowledge points. The Kruskal algorithm is used to construct a minimum spanning tree of knowledge on this relationship graph. By selecting the knowledge connection method with the lowest conflict intensity, cyclic conflicts are effectively avoided to handle knowledge conflicts. A matrix completion algorithm is applied to process the missing knowledge. After completion, the confidence of each knowledge point is evaluated by two factors: the comprehensive prediction error and the proportion of missing knowledge points around the knowledge point to the total number of knowledge points in the knowledge base. The former reflects the certainty of the model for a specific completed value, while the latter considers the completeness of the information around the knowledge point. Construct a chemical risk association network and calculate a hazardous chemical grid index to quantify the hazard level; Design a knowledge reasoning mechanism and use a grid knowledge compression function for dimensionality reduction; An adaptive optimization algorithm combining gradient descent and simulated annealing is implemented to dynamically adjust the mesh structure, where gradient descent provides the optimization direction and simulated annealing increases the possibility of escaping local optima; Based on the knowledge update frequency, a three-tier differentiated storage strategy is designed, storing knowledge points with high update frequency in the fast access layer, knowledge points with medium update frequency in the standard access layer, and knowledge points with low update frequency in the archive layer. A fuzzy processing mechanism for grid boundaries is established. Fuzzy set theory is used to handle cross-grid knowledge mapping problems. An improved Bell-type membership function is used to calculate the membership degree of each knowledge point to each adjacent grid cell. The original representation of the knowledge point and the influence of the knowledge points in the adjacent grid cells are weighted by the membership degree, so that the knowledge points achieve a natural transition at the grid boundary rather than hard boundary truncation, forming a gridded hazardous chemical knowledge base. The cross-grid knowledge mapping calculation formula is as follows: ; In the formula, For knowledge points In grid cells The representation in; For knowledge points The original vector representation; For knowledge points For grid cells The membership degree is calculated using fuzzy set theory; To be compatible with grid cells Adjacent grid cell sets; For knowledge points With grid cells The average correlation strength of all knowledge points in the text.
2. The method for establishing a gridded hazardous chemicals knowledge base according to claim 1, characterized in that, The accuracy assessment model includes a basic accuracy calculation equation, a weight adjustment equation, a conflict penalty equation, an experimental verification adjustment equation, and a final accuracy synthesis equation. The basic accuracy calculation equation is used to calculate the initial accuracy value of the knowledge points. The final accuracy synthesis equation is used to calculate the final accuracy index by integrating various factors.
3. The method for establishing a gridded hazardous chemicals knowledge base according to claim 2, characterized in that, In establishing a gridded knowledge mapping structure, knowledge grid units are divided according to the properties of hazardous chemicals, and grid size is set to divide the knowledge granularity, forming a multi-dimensional knowledge grid space; grid size refers to the size of the knowledge grid division granularity.
4. The method for establishing a gridded hazardous chemicals knowledge base according to claim 3, characterized in that, Grid centrality is an indicator that describes the importance of a knowledge point in the entire grid structure. Knowledge points with high centrality have more connections and stronger influence. Knowledge points with a centrality greater than 0.8 are considered core knowledge. The calculation of grid centrality comprehensively considers the weighted average component of the quality of adjacent knowledge points, the structural component that uses logarithmic function normalization to process connectivity to reduce the impact of extreme values, and the positional component that introduces betweenness centrality to reflect the bridging role of knowledge points in the network.
5. The method for establishing a gridded hazardous chemicals knowledge base according to claim 4, characterized in that, The knowledge conflict degree matrix is a two-dimensional array that records contradictory or inconsistent information in the knowledge base. It is used to identify knowledge that needs to be addressed in a focused manner. The edge weights represent the intensity of conflict between knowledge points. In the application of gridded dispersion analysis of knowledge distribution characteristics, gridded dispersion is an index describing the uniformity of knowledge points distributed in the grid space.
6. The method for establishing a gridded hazardous chemicals knowledge base according to claim 5, characterized in that, The Hazardous Chemicals Grid Index is a composite index that comprehensively considers the hazardous characteristics of chemicals and the completeness of knowledge about them. It is used to quantitatively assess the risk level of chemicals and establishes mapping relationships between grid nodes to improve the accuracy of risk identification.
7. The method for establishing a gridded hazardous chemicals knowledge base according to claim 6, characterized in that, The grid knowledge compression function is used to reduce the dimensionality of extremely long knowledge in the hazardous chemicals knowledge base and retain key information. The inputs include knowledge vector representation, grid centrality weight, knowledge correlation coefficient, compression ratio parameter and information retention importance threshold. The output is the compressed knowledge representation vector and information loss assessment index.
8. The method for establishing a gridded hazardous chemicals knowledge base according to claim 7, characterized in that, The specific structure of the chemical knowledge grid representation model is a multi-layer bidirectional transformer network architecture, which includes a hazardous chemical word embedding layer, a multi-head grid attention layer, a feedforward neural network layer, and a gridded representation output layer. The parameters of the gridded multi-head attention mechanism are jointly determined by the gridded centrality, the gridded size, and the knowledge association matrix.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, which, when executed in a computer, are used to perform the method for establishing a gridded hazardous chemicals knowledge base according to any one of claims 1-8.
10. A grid-based hazardous chemicals knowledge base establishment system, characterized in that, The system includes the computer-readable storage medium of claim 9, wherein the system is any one of a computer, a server, or a microcontroller, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.