Grid dangerous chemical knowledge base establishment method, medium and system

By building a grid-based hazardous chemicals knowledge base and utilizing multi-source data fusion and automated processing technology, we have solved the problems of traditional hazardous chemicals knowledge bases relying on manual operations and lacking correlation between knowledge, and achieved more efficient risk identification and prevention and control capabilities.

CN120633795AActive Publication Date: 2025-09-12INSPECTION & QUARANTINE TECH CENT SHANDONG ENTRY EXIT INSPECTION & QUARANTINE BUREAU
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510753821.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-12
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The construction of traditional hazardous chemicals knowledge base relies on a large amount of manual operations, and the correlation between knowledge is insufficiently expressed, resulting in insufficient accuracy and comprehensiveness in risk assessment and prevention and control.

Method used

Multi-source data fusion technology is used to construct a grid-based hazardous chemicals knowledge base, data sources are screened through accuracy assessment models, a grid-based knowledge mapping structure is established, Kruskal algorithm is used to handle conflicts, matrix completion algorithm is applied to handle defective knowledge, knowledge reasoning mechanism is designed for dimensionality reduction, and adaptive optimization algorithm is implemented to dynamically adjust the grid structure.

Benefits of technology

It achieves the automation and accuracy of the correlation expression between hazardous chemical knowledge, reduces manual operations, improves risk identification and prevention and control capabilities, and provides more intelligent and systematic knowledge support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633795A_ABST
    Figure CN120633795A_ABST
Patent Text Reader

Abstract

The invention provides a gridding dangerous chemical knowledge base establishment method, medium and system, and belongs to the technical field of knowledge base establishment methods, and the method comprises the steps: firstly collecting data, and screening high-quality data through an acceptance degree index; then constructing an accuracy evaluation model to automatically mark various types of knowledge and forming a knowledge correlation degree matrix to quantify correlation strength; establishing a gridding knowledge mapping structure, dividing knowledge grid units, calculating gridding centrality and determining knowledge point weights; automatically constructing a minimum spanning tree to process knowledge conflicts, and processing defect knowledge by applying a matrix completion algorithm; a chemical risk association network is constructed, and dangerous chemical grid indexes are calculated to quantify danger levels; then designing a knowledge reasoning mechanism and processing super-long knowledge by adopting a grid knowledge compression function; and finally, implementing an adaptive optimization algorithm to dynamically adjust the grid structure, and forming a grid dangerous chemical knowledge base which reduces manual dependence and enhances relevance expression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge base establishment methods, and specifically relates to a grid-based hazardous chemicals knowledge base establishment method, medium, and system. Background Art

[0002] Hazardous chemicals knowledge bases are crucial infrastructure for chemical safety management and risk prevention, providing knowledge support for safe production, accident prevention, and emergency response. Traditionally, hazardous chemicals knowledge bases have relied on manual collection and organization. Professionals review relevant literature, manuals, experimental data, and accident reports to extract information on the chemicals' physical and chemical properties, hazardous properties, storage and transportation requirements, and emergency response methods. This information is then organized and categorized according to specific structures to form a structured knowledge system. This expert-based knowledge base construction approach has proven effective in practical applications.

[0003] However, the construction of traditional hazardous chemical knowledge bases has significant shortcomings. First, manual data collection and organization consumes a significant amount of human resources. With the increasing number of chemical types and the continuous updating of safety standards, manual operations make it difficult to timely update and maintain the knowledge base. Second, existing knowledge bases generally use simple classification or hierarchical structures to organize knowledge, which makes it difficult to express the complex interactions between chemicals and the correlation between their hazardous properties. For example, related knowledge such as the synergistic hazardous effects produced by the mixing of certain chemicals and the changes in hazardous properties under specific environmental conditions is difficult to effectively express and mine.

[0004] The current technical issues with hazardous chemical knowledge base construction, such as high reliance on manual operations and insufficient expression of inter-knowledge correlations, severely restrict the accuracy and comprehensiveness of hazardous chemical risk assessment and prevention and control. A new approach that reduces manual reliance and enhances the ability to express knowledge correlations is urgently needed to overcome this technical bottleneck. Specifically, existing technologies for hazardous chemical knowledge base construction often rely heavily on manual operations and suffer from insufficient expression of inter-knowledge correlations. Summary of the Invention

[0005] In view of this, the present invention provides a grid-based hazardous chemical knowledge base establishment method, medium and system, which can solve the technical problems in the prior art that the construction of hazardous chemical knowledge base often relies on a large amount of manual operations and has insufficient expression of the correlation between knowledge.

[0006] The present invention is implemented as follows: In a first aspect, the present invention provides a method for establishing a grid-based hazardous chemical knowledge base, comprising: collecting hazardous chemical data to establish an initial knowledge system, using multi-source data fusion technology to form a data matrix, and screening data sources according to a recognition index; constructing an accuracy assessment model, marking redundant knowledge, conflicting knowledge, erroneous knowledge, and defective knowledge, and forming a knowledge correlation matrix; establishing a grid-based knowledge mapping structure, calculating grid-based centrality to determine knowledge point weights; using the Kruskal algorithm to construct a knowledge minimum spanning tree to handle knowledge conflicts, and applying a matrix completion algorithm to handle defective knowledge; constructing a chemical risk association network, calculating a hazardous chemical grid index to quantify the hazard level; designing a knowledge reasoning mechanism, and using a grid knowledge compression function for dimensionality reduction processing; implementing an adaptive optimization algorithm to dynamically adjust the grid structure to form a grid-based hazardous chemical knowledge base.

[0007] Among them, the accuracy evaluation model includes the accuracy basic calculation equation, the weight adjustment equation, the conflict penalty equation, the experimental verification adjustment equation, and the final accuracy synthesis equation; the accuracy basic calculation equation is used to calculate the initial accuracy value of the knowledge point; the final accuracy synthesis equation is used to comprehensively calculate the final accuracy index based on various factors.

[0008] Among them, in establishing a grid knowledge mapping structure, the knowledge grid units are divided according to the properties of hazardous chemicals, and the grid size is set to divide the knowledge granularity to form a multi-dimensional knowledge grid space; the grid size refers to the size of the knowledge grid division granularity, the smaller size provides a more detailed knowledge representation, and the larger size provides a more macro knowledge overview.

[0009] Among them, grid centrality is an indicator that describes the importance of knowledge points in the entire grid structure. Knowledge points with high centrality usually have more connections and stronger influence. In the calculation of grid aggregation, grid aggregation is an indicator that measures the density of knowledge points in a certain area, which is used to identify key knowledge areas and blank areas.

[0010] Among them, the knowledge conflict degree matrix is ​​a two-dimensional array that records contradictory or inconsistent information in the knowledge base, which is used to identify the knowledge that needs to be resolved; the edge weight is used to represent the intensity of the conflict between knowledge points; in the application of grid dispersion to analyze the knowledge distribution characteristics, grid dispersion is an indicator that describes the uniformity of the distribution of knowledge points in the grid space.

[0011] Among them, the hazardous chemicals grid index is a composite indicator that comprehensively considers the hazardous characteristics of chemicals and the completeness of knowledge, and is used to quantitatively assess the risk level of chemicals; establishes a mapping relationship between grid nodes to improve the accuracy of risk identification.

[0012] Among them, the grid knowledge compression function is used to reduce the dimensionality of extremely long knowledge in the hazardous chemicals knowledge base and retain key information. The input includes knowledge vector representation, grid centrality weight, knowledge correlation coefficient, compression ratio parameter and information retention importance threshold. The output is the compressed knowledge representation vector and information loss evaluation index.

[0013] Among them, the specific structure of the chemical knowledge grid representation model is a multi-layer bidirectional transformer network architecture, which includes a hazardous chemical word embedding layer, a multi-head grid attention layer, a feedforward neural network layer, and a grid representation output layer. The parameters of the grid multi-head attention mechanism are jointly determined by the grid centrality, grid size, and knowledge correlation matrix.

[0014] A second aspect of the present invention provides a computer-readable storage medium having program instructions stored therein. When the program instructions are run in a computer, the program instructions are used to execute the above-mentioned method for establishing a grid-based hazardous chemical knowledge base.

[0015] The third aspect of the present invention provides a grid-based hazardous chemical knowledge base establishment system, which includes the above-mentioned computer-readable storage medium. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set in the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.

[0016] By introducing multi-source data fusion technology and automated data screening mechanisms, this invention achieves efficient collection and integration of hazardous chemical data, significantly reducing reliance on manual operations. By constructing a multidimensional knowledge grid structure and a grid representation model, this invention creatively addresses the problem of insufficient representation of the relationships between hazardous chemical knowledge. Using accuracy assessment models, knowledge relevance matrices, and other technical approaches, this invention achieves automated hierarchical analysis of knowledge points and quantifies the strength of relationships, enabling precise representation of complex relationships between knowledge points. Through innovative metrics such as grid centrality and grid aggregation, this invention captures the distribution characteristics and influence of knowledge points in grid space, providing a systematic perspective for a comprehensive understanding of the hazardous properties of chemicals.

[0017] By constructing a grid-based knowledge representation system and an automated processing mechanism, this method effectively solves the core technical problems of relying on a large amount of manual operations and insufficient expression of the correlation between knowledge in the construction of traditional hazardous chemical knowledge bases. It provides more intelligent, systematic and comprehensive knowledge support for chemical safety management, and greatly enhances the risk identification and prevention and control capabilities of hazardous chemicals. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flow chart of the method of the present invention.

[0019] Figure 2 This is a comparison chart of the recognition index of each data source in Example 2.

[0020] Figure 3 This is a comparison chart of the accuracy assessment of hazardous chemicals knowledge points in Example 2.

[0021] Figure 4 This is a gridded centrality analysis diagram of key knowledge points in Example 2.

[0022] Figure 5 This is a distribution map of hazardous chemicals knowledge in the four-dimensional grid structure in Example 2.

[0023] Figure 6 This is the hazardous chemicals grid index and risk assessment diagram in Example 2. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0025] like Figure 1 FIG. 1 is a flowchart of a method for establishing a grid-based hazardous chemicals knowledge base according to the first aspect of the present invention. The method comprises the following steps: S01. Collect hazardous chemicals data to establish an initial knowledge system. Use multi-source data fusion technology to form a data matrix. Filter data sources based on the recognition index. Calculate the recognition index matrix for each knowledge point and set a threshold to filter low-quality data. S02. Construct an accuracy assessment model, calculate the accuracy index of each knowledge point, perform hierarchical analysis on the knowledge points and mark redundant knowledge, conflicting knowledge, erroneous knowledge, and defective knowledge, form a knowledge correlation matrix, and quantify the correlation strength; S03. Establish a grid knowledge mapping structure, divide the knowledge grid units according to the properties of hazardous chemicals, calculate the grid centrality to determine the weight of the knowledge points, set the grid size to divide the knowledge granularity, and form a multi-dimensional knowledge grid space; S04. Use Kruskal's algorithm to construct a minimum spanning tree to handle knowledge conflicts, construct a knowledge conflict degree matrix to measure the degree of conflict, apply a matrix completion algorithm to handle defective knowledge, calculate grid aggregation to identify knowledge-intensive areas, and use edge weights to represent the intensity of conflicts between knowledge points. S05. Input hazardous chemical attribute data and grid block information into a pre-trained chemical knowledge grid representation model to construct a chemical risk association network, calculate the hazardous chemical grid index to quantify the hazard level, apply grid dispersion to analyze knowledge distribution characteristics, establish mapping relationships between grid nodes, and improve risk identification accuracy; S06. Design a knowledge reasoning mechanism that integrates logical rules and statistical methods. Use grid knowledge compression functions to reduce the dimension of extremely long knowledge. Build a hierarchical knowledge representation model to optimize the storage structure, achieving knowledge simplification and key information retention. S07. Implement an adaptive optimization algorithm based on the chemical knowledge grid representation model to dynamically adjust the grid structure, design differentiated storage strategies based on the knowledge update frequency, establish a grid boundary fuzzy processing mechanism to solve the cross-grid knowledge mapping problem, realize intelligent adjustment of the grid structure, and obtain the final gridded hazardous chemicals knowledge base.

[0026] The recognition index is a quantitative indicator that measures the degree to which a piece of hazardous chemical knowledge is recognized by multiple authoritative data sources. Its value ranges from 0 to 1, with higher values ​​indicating greater recognition. The accuracy index is a quantitative measure of the precision of knowledge points, derived through experimental verification and theoretical model calculations, and is used to screen high-quality knowledge points. The knowledge relevance matrix is ​​a two-dimensional array that represents the degree of interconnectedness between hazardous chemical knowledge points. Each element represents the strength of the association between two knowledge points. The knowledge conflict matrix is ​​a two-dimensional array that records contradictory or inconsistent information within the knowledge base and is used to identify conflicting knowledge points that require resolution. Grid centrality is an indicator that describes the importance of a knowledge point within the overall grid structure. Knowledge points with high centrality typically have more connections and greater influence. Grid size refers to the granularity of the knowledge grid. Smaller sizes provide a more refined knowledge representation, while larger sizes provide a more macroscopic knowledge overview. Grid clustering measures the density of knowledge points within a region and is used to identify key knowledge areas and gaps. Grid dispersion describes the uniformity of the distribution of knowledge points within the grid space. High dispersion indicates uniform knowledge distribution, while low dispersion indicates uneven knowledge distribution. The Hazardous Chemicals Grid Index is a composite indicator that comprehensively considers the hazardous properties of chemicals and the completeness of knowledge, and is used to quantitatively assess the risk level of chemicals.

[0027] Among them, the grid knowledge compression function is used to reduce the dimensionality of extremely long knowledge in the hazardous chemicals knowledge base and retain key information. The input includes the knowledge vector representation obtained from S03, the grid centrality weight calculated from S03, the knowledge correlation coefficient extracted from the knowledge correlation matrix formed in S02, the manually set compression ratio parameter, and the retained information importance threshold determined based on knowledge importance analysis. The output is the compressed knowledge representation vector used for the hierarchical knowledge representation model construction in S06 and the information loss evaluation index used for the knowledge base performance evaluation in S09.

[0028] Among them, the specific structure of the chemical knowledge grid representation model is a multi-layer bidirectional transformer network architecture, which includes a hazardous chemical word embedding layer, a multi-head grid attention layer, a feedforward neural network layer, and a grid representation output layer. The parameters of the grid multi-head attention mechanism are determined by the grid centrality calculated in S03, the grid size set in S03, and the knowledge association matrix formed in S02. The number of grid attention heads matches the grid division dimension, and each attention head captures the association pattern of different types of chemical hazardous characteristics.

[0029] Among them, the steps for establishing the training data set in the chemical knowledge grid representation model training process specifically include collecting international hazardous chemical safety data sheets, hazardous chemical accident case reports, chemical structure database information, and hazardous chemical descriptions in professional journal literature, constructing triple knowledge representations, generating knowledge graph node representations, designing knowledge grid mapping matrices, constructing positive and negative sample pairs, forming a gridded knowledge training corpus, designing training task types, and dividing the training set, validation set, and test set.

[0030] Among them, the steps of training the chemical knowledge grid representation model specifically include initializing model parameters, designing knowledge mask recovery tasks to train grid representation capabilities, building grid node prediction tasks to enhance the model's understanding of grid structure, designing hazardous chemical property association prediction tasks to improve risk identification capabilities, using a contrastive learning framework to optimize the grid representation space, combining knowledge distillation technology to integrate expert knowledge, using a distributed training system to accelerate the training process, evaluating model performance through a validation set, using an early stopping strategy to prevent overfitting, saving training weights and grid mapping parameters, and building a model deployment package for subsequent applications.

[0031] The accuracy evaluation model includes an accuracy basic calculation equation, a weight adjustment equation, a conflict penalty equation, an experimental verification adjustment equation, and a final accuracy synthesis equation; The accuracy basic calculation equation is used to calculate the initial accuracy value of the knowledge point. The input includes the knowledge point recognition index matrix obtained from S01, the knowledge point data source reliability coefficient, the knowledge point release time decay factor, and the knowledge point expert score mean. The output is the knowledge point basic accuracy index. The weight adjustment equation is used to adjust the accuracy assessment weight according to the importance of the knowledge point. The input includes the knowledge point reference frequency, the knowledge point application scenario coverage, the amount of knowledge associated with the knowledge point, and the knowledge point danger level coefficient. The output is the knowledge point weight adjustment coefficient. The conflict penalty equation is used to reduce the accuracy score of knowledge points with conflicts. The input includes the number of conflicts between knowledge points, the conflict severity coefficient, the average accuracy of the conflicting knowledge points, and the conflict resolution difficulty coefficient. The output is the knowledge point conflict penalty value. The experimental verification adjustment equation is used to correct the accuracy according to the experimental data. The input includes the number of experimental verifications, the consistency coefficient of the experimental data, the experimental condition strictness score, the experimental method reliability coefficient, and the experimental result standard deviation. The output is the experimental verification adjustment coefficient. The final accuracy synthesis equation is used to comprehensively calculate the final accuracy index by combining various factors. The input includes the basic accuracy index of the knowledge point, the knowledge point weight adjustment coefficient, the knowledge point conflict penalty value, and the experimental verification adjustment coefficient. The output is the final accuracy index of the knowledge point, which is used for the knowledge point hierarchical analysis and labeling in S02.

[0032] The specific implementation of the above steps is described in detail below.

[0033] The specific implementation of step S01 is to first collect hazardous chemical data from multiple sources such as the National Hazardous Chemicals Database, International Chemical Safety Cards, and academic journal databases. Then, a hash table data structure is used to organize multi-source data, and data fusion technology is used to merge the same knowledge points from different sources into a unified representation to form a dimensional data matrix, where Indicates the number of knowledge points, represents the number of data sources. The acceptance index for each data source is then calculated based on factors such as its authority, update frequency, data completeness, and historical accuracy. The value ranges from 0 to 1, with a threshold of 0.6 or above typically considered reliable. Finally, a recognition index matrix is ​​calculated for each knowledge point. A weighted average algorithm is used to synthesize the support for each knowledge point from each data source. Knowledge points with an acceptance index below 0.5 are filtered out, retaining high-quality data for the next step. This step aims to ensure the data quality and reliability of the initial knowledge system, laying the foundation for subsequent knowledge base construction.

[0034] The specific implementation of step S02 is to first construct an accuracy evaluation model, which contains five core equations. The basic accuracy calculation equation adopts a multi-factor weighted summation method, which takes the recognition index, data source reliability coefficient (usually 0.1 to 0.9), time decay factor (usually , λ is 0.1-0.3) and expert ratings (on a scale of 1-10) as inputs, and outputs a baseline accuracy value. The weight adjustment equation generates an adjustment coefficient through a nonlinear combination of citation frequency (normalized to 0-1), application scenario coverage (typically a percentage), the number of associated knowledge points, and the risk level coefficient (1-8). The conflict penalty equation calculates a penalty value based on the number of conflicts, severity (0-1), mean accuracy of conflicting knowledge, and resolution difficulty coefficient (1-5). The experimental validation adjustment equation calculates the adjustment coefficient based on the number of experiments, data consistency coefficient (0.5-1), condition stringency (1-10), method reliability (0.6-0.95), and result standard deviation. The final accuracy synthesis equation weights and combines the above outputs to generate the final accuracy index. A hierarchical clustering algorithm is then used to classify and analyze knowledge points. Redundant knowledge is marked based on semantic similarity and content overlap. Conflicting knowledge is identified through logical relationship verification of knowledge points. Incorrect knowledge is identified based on the accuracy index, and defective knowledge is identified through completeness assessment. Finally, we construct a knowledge association matrix. We use cosine similarity and an association rule mining algorithm to quantify the strength of associations between knowledge points. The value ranges from 0 to 1, and knowledge point pairs with an association strength greater than 0.7 are generally considered strongly associated. The goal of this step is to establish a scientific knowledge evaluation system and improve the quality of the knowledge base.

[0035] The specific implementation of step S03 involves first defining a multidimensional knowledge space based on the physical and chemical properties, hazard categories, usage conditions, and storage requirements of hazardous chemicals. This knowledge space is then divided into regular units using a gridding algorithm. Next, the gridded centrality of each knowledge point is calculated using the eigenvector centrality algorithm, which takes into account the number of connections, connection strength, and accuracy index of the knowledge point. Centrality values ​​range from 0 to 1, and knowledge points with a centrality greater than 0.8 are typically considered core knowledge. The gridding size parameter is then set, typically selecting a 3- to 8-dimensional grid structure based on the data scale and application requirements. Smaller sizes (e.g., 3×3×3) provide detailed representation, while larger sizes (e.g., 8×8×8) provide a macro overview. Finally, a multidimensional knowledge grid space is constructed, using tensor representation to store knowledge information across all dimensions. This complete gridded knowledge mapping structure enables multidimensional knowledge organization and rapid retrieval. This step aims to establish a structured knowledge organization approach, providing a spatial framework for subsequent conflict resolution and risk association.

[0036] The specific implementation of step S04 involves first constructing a relationship graph between knowledge points, where nodes represent knowledge points, edges represent relationships between knowledge points, and edge weights represent the intensity of conflict between knowledge points. A minimum spanning tree is then constructed using the Kruskal algorithm. This algorithm selects knowledge connections with minimal conflict, effectively avoiding cyclic conflicts and is suitable for handling knowledge conflicts in complex networks. A knowledge conflict degree matrix is ​​then constructed. Matrix element values ​​represent the degree of conflict between knowledge points, ranging from 0 to 1. Knowledge point pairs with a conflict degree greater than 0.6 are typically marked as severely conflicting and prioritized. For defective knowledge, a matrix completion algorithm based on the low-rank assumption is applied. Missing information is inferred from the relationships between existing knowledge points. For low-confidence completion results (the confidence threshold is typically set at 0.7), the missing label is retained for verification. Finally, the gridded clustering degree is calculated, and the Moran index method from spatial statistics is used to identify knowledge-dense and knowledge-sparse areas. Areas with a clustering degree greater than 0.8 are identified as knowledge-dense areas, while areas with a clustering degree less than 0.2 are identified as knowledge-gap areas. This step aims to resolve knowledge conflicts, improve the knowledge structure, and enhance the consistency of the knowledge base.

[0037] Step S05 is implemented by first taking the physical and chemical property data of hazardous chemicals (such as flash point, boiling point, and toxicity index) and grid segmentation information (from step S03) as input feature vectors and feeding them into a pre-trained chemical knowledge grid representation model. This model utilizes a multi-layer bidirectional transformer architecture, capable of capturing complex nonlinear relationships between chemical properties. Next, a chemical risk association network is constructed, using graph neural network technology to establish possible reactions, synergistic toxicity, and cascading risk relationships between hazardous chemicals. The association strength threshold is typically set at 0.65. Next, a hazardous chemical grid index is calculated, which comprehensively considers the inherent hazard of the chemical (weighted by 0.6) and the completeness of knowledge (weighted by 0.4) to quantify the hazard level. The index ranges from 0 to 10, with 8-10 representing extremely high hazard, 6-8 representing high hazard, 4-6 representing moderate hazard, 2-4 representing low hazard, and 0-2 representing minimal hazard. A grid-based dispersion analysis method is then applied, employing an entropy calculation model to assess the uniformity of knowledge distribution. A dispersion greater than 0.7 indicates uniform knowledge distribution, while a dispersion less than 0.3 indicates extremely uneven knowledge distribution, indicating the need for supplementary domain knowledge. Finally, a mapping relationship is established between grid nodes, and tensor mapping functions are used to connect knowledge points of different dimensions to form a complete knowledge network. This step aims to form a hazardous chemical risk association system and improve risk identification accuracy.

[0038] The specific implementation of step S06 involves first designing a knowledge reasoning mechanism that integrates rule-based reasoning based on first-order predicate logic and statistical reasoning methods based on Bayesian networks to achieve both precise and uncertain reasoning about hazardous chemical knowledge. Next, a grid knowledge compression function is applied to reduce the dimensionality of extremely long knowledge content (typically knowledge points with more than 1000 dimensions or text length exceeding 5000 characters). This compression function combines principal component analysis with an autoencoder. Its inputs include the knowledge vector representation, grid centrality weights (typically 0.1-0.9), knowledge relevance coefficients (typically 0.1-0.8), a compression ratio parameter (typically set to 0.3-0.7), and an information importance threshold (typically 0.6). The compression function outputs a compressed knowledge representation vector (typically with a 50%-80% reduction in dimensionality) and an information loss assessment index (typically set to less than 0.2). A hierarchical knowledge representation model is then constructed, organizing knowledge in a tree-like structure, with upper-level nodes storing abstract concepts and lower-level nodes storing specific details, optimizing the storage structure. Finally, knowledge simplification and key information retention are achieved, knowledge importance is sorted, and key information with an importance index greater than 0.7 is retained. The purpose of this step is to optimize knowledge representation and storage efficiency and ensure the integrity and accessibility of key information.

[0039] The specific implementation of step S07 involves first implementing an adaptive optimization algorithm based on a chemical knowledge grid representation model. This algorithm combines gradient descent and simulated annealing strategies to dynamically adjust grid structure parameters based on knowledge updates. The algorithm typically converges when the grid structure change rate is less than 0.01 for five consecutive iterations or when a maximum number of 100 iterations is reached. Next, a differentiated storage strategy is designed based on the frequency of knowledge updates: knowledge points with high update frequency (e.g., updated more than once a month) are stored in a fast access layer, knowledge points with medium update frequency (e.g., updated once a quarter) are stored in a standard access layer, and knowledge points with low update frequency (e.g., updated less than once a year) are stored in an archive layer. A grid boundary fuzzy processing mechanism is then established, employing fuzzy set theory to address cross-grid knowledge mapping. A membership threshold is typically set between 0.4 and 0.6 to address the ambiguity of knowledge attribution at grid boundaries. Finally, intelligent grid structure adjustment is implemented, automatically optimizing grid parameters, including grid dimension, grid size, and grid connectivity, based on knowledge base usage feedback and the characteristics of newly added knowledge, to form the final gridded hazardous chemical knowledge base. This step aims to improve the adaptability and scalability of the knowledge base and achieve dynamic optimization and continuous improvement of the knowledge base.

[0040] The chemical knowledge grid representation model utilizes a multi-layer bidirectional transformer network architecture, consisting of four core layers: a hazardous chemical word embedding layer that uses a 300-dimensional vector space to represent chemical terms and properties, employing domain-adaptive pre-training to capture chemical semantics; a multi-head grid attention layer consisting of eight attention heads, each corresponding to a different hazard characteristic dimension (such as flammability, toxicity, and reactivity); attention weights are determined by grid centrality, grid size, and a knowledge relevance matrix; a two-layer feedforward neural network layer with a hidden layer dimension of 1024, using the GELU activation function to process the attention layer output; and a grid representation output layer that generates a 512-dimensional grid knowledge representation and outputs grid index information. The model employs skip connections and layer normalization to enhance training stability, and has a total of approximately 45 million parameters.

[0041] The steps to establish the training data set for the chemical knowledge grid representation model are as follows: First, collect data sources, including international hazardous chemical safety data sheets (such as ICSC cards), hazardous chemical accident case reports (covering major accidents in the past 20 years), chemical structure database information (such as PubChem data), and hazardous chemical descriptions in professional journals (collect no less than 5,000 documents). Then construct a triple knowledge representation and convert the collected unstructured data into a "subject-relation-object" format, such as " - has - highly corrosive", a total of no less than 500,000 triples are generated. Then, a knowledge graph node representation is generated, and the TransE and ComplEx algorithms are used to map the triples to the vector space. A knowledge grid mapping matrix is ​​designed to establish a mapping relationship between the knowledge graph nodes and the grid space. Positive and negative sample pairs are constructed, with positive samples being related chemical pairs and negative samples being randomly sampled unrelated chemical pairs, with a ratio of 1:3. A gridded knowledge training corpus is formed, containing text such as descriptions of hazardous chemical properties, safe operating procedures, and emergency response measures. Training task types are designed, including mask prediction, grid relationship prediction, and hazardous attribute classification tasks. Finally, the data set is divided into training set, validation set, and test set in a ratio of 7:1.5:1.5 to ensure a balanced distribution of chemical categories in each set.

[0042] A second aspect of the present invention provides a computer-readable storage medium having program instructions stored therein. When the program instructions are run in a computer, the program instructions are used to execute the above-mentioned method for establishing a grid-based hazardous chemical knowledge base.

[0043] The third aspect of the present invention provides a grid-based hazardous chemical knowledge base establishment system, which includes the above-mentioned computer-readable storage medium. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set in the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.

[0044] The mathematical model or calculation process involved in the present invention is described in detail below.

[0045] The calculation process of the recognition index matrix in step S01 can be specifically expressed as follows: ; Where, For knowledge points In the data source The recognition index in For data sources The reliability coefficient ranges from 0 to 1; For knowledge points In the data source The existence status of , if it exists, it is 1, if it does not exist, it is 0; For data sources The time factor indicates the timeliness of data source updates.

[0046] The calculation formula for the comprehensive recognition index is as follows: ; Where, For knowledge points Comprehensive recognition index; For data sources The weight coefficient of The total number of data sources.

[0047] The parameter acquisition method is: It is obtained by evaluating the authority, historical accuracy and credibility of the data source. It can be calculated using the analytic hierarchy process and is calculated by multiple experts (no less than 5) scoring the data source. Through the function Calculate, where is the time attenuation coefficient, generally ranging from 0.1 to 0.3; is the current time; The last update time of the data source; It is determined by an expert group based on a comprehensive assessment of the frequency of use, coverage and data integrity of the data source, usually using the Delphi method.

[0048] The accuracy evaluation model equations in step S02 are specifically expressed as follows: 1. Accuracy basic calculation equation: ; Where, For knowledge points The basic accuracy index ranges from 0 to 1; For knowledge points Recognition index; For knowledge points The data source reliability coefficient ranges from 0.1 to 0.9; For knowledge points The difference between the release time and the current time (in years); is the time attenuation factor, ranging from 0.1 to 0.3; For knowledge points The average expert rating of , with a full score of 10 points; are the weight coefficients of each factor, and satisfy .

[0049] 2. Weight adjustment equation: ; Where, For knowledge points The weight adjustment coefficient of For knowledge points Citation frequency (normalized to 0–1); For knowledge points Application scenario coverage (percentage value); For knowledge points The amount of associated knowledge; For knowledge points Danger level coefficient (1 to 8); are the weight coefficients of each factor, and satisfy .

[0050] 3. Conflict penalty equation: ; Where, For knowledge points The conflict penalty value; For knowledge points the number of conflicts; is the conflict severity coefficient (0-1); For knowledge points The average accuracy of conflicting knowledge points; is the difficulty level of conflict resolution (1 to 5); are the weight coefficients of each factor, and satisfy .

[0051] 4. Experimental verification of the adjustment equation: ; Where, For knowledge points The experimental verification adjustment coefficient of is the number of experimental verifications; is the consistency coefficient of experimental data (0.5-1); Score the stringency of experimental conditions (1–10); is the reliability coefficient of the experimental method (0.6 to 0.95); is the standard deviation of the experimental results; are the weight coefficients of each factor, and satisfy .

[0052] 5. Final accuracy synthesis equation: ; Where, For knowledge points The final accuracy index of is the basic accuracy index; is the weight adjustment coefficient; is the conflict penalty value; Adjustment coefficients for experimental verification; is the smoothing factor, generally ranging from 0.01 to 0.1.

[0053] The parameter acquisition method is: Based on a questionnaire survey of domain experts, usually ; Determined according to the importance of application scenarios, usually ; Based on the assessment of the impact of the conflict, usually ; Determined based on experimental reliability experience, usually .

[0054] The accuracy assessment model uses a nonlinear weighted combination approach, taking into account multiple influencing factors. A logarithmic relationship is used for citation frequency to reflect the principle of diminishing marginal utility; a square relationship is used for application scenario coverage to emphasize the importance of comprehensive coverage; a square root relationship is used for the amount of associated knowledge to reflect the decreasing value of associated knowledge growth; a logical function is used to handle the number of experiments, reflecting the diminishing gain after a certain number of experimental verifications; conflict penalties are processed using nonlinear normalization to avoid extreme penalties; and a multiplicative combination is used for final accuracy to ensure the combined influence of various factors, and a smoothing factor is introduced to prevent calculation anomalies caused by zero values.

[0055] The calculation formula of the knowledge relevance matrix in step S02 is as follows: ; Where, Representing knowledge points And knowledge points The correlation strength between them is calculated as follows:

[0056] ; Where, is the knowledge point vector and The cosine similarity of It is the support function in association rule mining; For knowledge points and The number of co-occurrences of is the maximum number of co-occurrences among all knowledge point pairs; is the weight coefficient and satisfies .

[0057] The parameter acquisition method is: Knowledge point vector Convert the knowledge point text into a 300-768 dimensional vector using a text embedding model (such as Word2Vec or BERT); Calculate knowledge points using the Apriori algorithm and Probability of simultaneous occurrence; Obtained by counting the co-occurrence frequency of two knowledge points in all documents; It is determined according to the importance of each factor in practical applications. .

[0058] The grid centrality calculation formula in step S03 is as follows: ; Where, For knowledge points The grid centrality of For knowledge points and The strength of association; For knowledge points The final accuracy index of For knowledge points The connectivity of a knowledge point (the number of knowledge points directly connected to it); is the maximum connectivity among all knowledge points; For knowledge points betweenness centrality; is the maximum betweenness centrality among all knowledge points.

[0059] The parameter acquisition method is: Obtained from the knowledge relevance matrix in step S02; Obtained from the accuracy assessment model output of step S02; By calculating the knowledge relevance matrix The number of knowledge points with association strength greater than a threshold (usually 0.3) is obtained; Calculated by the following formula: ,in For knowledge points arrive The number of shortest paths, For passing knowledge points of arrive The number of shortest paths.

[0060] Grid-based centrality calculation uses a combination of weighted averaging, logarithmic normalization, and linear enhancement to reflect the importance of a knowledge point's position in the knowledge network. Weighted averaging considers the quality of adjacent knowledge points, logarithmic normalization processes connectivity to reduce the impact of extreme values, and betweenness centrality enhancement emphasizes the "bridge" role of a node in the network.

[0061] The calculation formula of the knowledge conflict degree matrix in step S04 is as follows: ; Where, Representing knowledge points And knowledge points The degree of conflict between them is calculated as follows: ; Where, is the knowledge point vector and Semantic similarity of The strength of logical conflicts is detected by predicate logic rules; is the degree of numerical conflict, which indicates the degree of inconsistency of numerical knowledge; is the maximum numerical conflict degree among all knowledge point pairs; is the weight coefficient and satisfies .

[0062] The parameter acquisition method is: Obtained by calculating the cosine similarity or dot product of knowledge point vectors; The rule engine detects logical contradictions between knowledge points. If there is a complete contradiction, the value is 1; if there is a partial contradiction, the value is 0.1-0.9; For numerical knowledge, through the formula Calculate, where and is the corresponding numerical value; Determined based on the importance of the conflict type, usually .

[0063] The matrix completion algorithm applied in step S04 is specifically implemented as follows: The matrix completion model of defect knowledge can be expressed as: ; Where, is the complete matrix to be completed; is the observation matrix containing missing values; is the projection operator, corresponding to the set of observed entries ; is a matrix The nuclear norm of (the sum of all singular values); It is a regularization parameter used to balance the fitting error and matrix complexity, and is usually set to 0.1 to 1.

[0064] The confidence calculation formula for the completed knowledge point is: ; Where, To complete the value confidence level; is the predicted value in cross validation; is the standard deviation of the prediction error; For knowledge points The number of relevant missing knowledge points; is the total number of knowledge points in the knowledge base.

[0065] The parameter acquisition method is: The known elements in the matrix come directly from the knowledge base, and the unknown elements are marked as missing; Select the optimal value through cross-validation method; The known data is randomly divided into a training set and a validation set, the model is trained on the training set and the validation set is predicted; It is obtained by calculating the standard deviation of the difference between the predicted value and the true value on the validation set.

[0066] The matrix completion algorithm is based on the low-rank assumption, assuming the existence of underlying low-dimensional structure between knowledge points. It encourages low-rank solutions by minimizing the nuclear norm while maintaining consistency with observed data. The confidence calculation combines the prediction error and the proportion of relevant missing knowledge points. The former reflects the model's certainty about a specific completion value, while the latter considers the completeness of the information surrounding the knowledge point.

[0067] The gridding aggregation degree in step S04 is calculated using Moran's index (I), and the formula is as follows: ; Where, is the Moran index, and its value range is usually -1 to 1; is the total number of knowledge points; is the spatial weight matrix element, reflecting the knowledge point and distance relations in grid space; For knowledge points Attribute values ​​(such as importance or access frequency); is the average value of all knowledge point attribute values.

[0068] The calculation formula of local grid aggregation is: ; Where, For knowledge points The degree of aggregation in the local area.

[0069] The parameter acquisition method is: By formula Calculate, where For knowledge points and Euclidean distance in grid space, is the distance threshold, usually set to twice the diagonal length of the grid cell; It can be an importance index of a knowledge point (such as grid centrality), an accuracy index, or an access frequency.

[0070] The Moran index quantifies spatial autocorrelation, reflecting the degree of clustering of knowledge distribution patterns within a grid space. Positive values ​​indicate clustering of similar values ​​(knowledge-dense areas), negative values ​​indicate clustering of dissimilar values ​​(e.g., knowledge gaps interspersed with dense distribution), and values ​​close to zero indicate random distribution. Local clustering calculations identify knowledge density characteristics around each grid cell, facilitating the identification of specific knowledge-dense areas.

[0071] The calculation formula for the hazardous chemical grid index in step S05 is as follows: ; Where, For chemicals The grid risk index ranges from 0 to 10; The inherent hazard index of chemicals is assessed based on the GHS classification standard, with a value range of 0 to 10; is the knowledge completeness index, which indicates the coverage of relevant knowledge points; is the completeness threshold parameter, usually ranging from 3 to 5; is the weight coefficient and satisfies .

[0072] Knowledge completeness index calculation formula: ; Where, For chemicals A collection of related knowledge points; For knowledge points The final accuracy index of For knowledge points The grid centrality of For collection The number of elements in .

[0073] The parameter acquisition method is: Based on the chemical's hazardous characteristics (such as flash point, LD50, corrosiveness, etc.), the GHS classification method is used to convert each hazard category into a score of 0 to 10, and the highest score is taken as the inherent hazard index; By searching the knowledge base for chemicals All knowledge points with correlation (correlation strength greater than 0.3) are obtained; Determined according to risk management strategy, usually , reflecting the dominant role of inherent hazards in risk assessment.

[0074] The grid risk index is calculated using a combination of linear combinations and sigmoid transformations. The linear combination preserves the direct impact of inherent hazard, while the sigmoid function accounts for the contribution of knowledge completeness, making risk assessments more conservative for chemicals with limited knowledge. The knowledge completeness index takes into account both the quality (accuracy and centrality) and quantity (logarithmic scaling) of relevant knowledge. The logarithmic transformation reflects the diminishing marginal returns of knowledge quantity.

[0075] The gridding dispersion calculation in step S05 adopts the entropy method, and the formula is as follows: ; Where, is the grid dispersion, ranging from 0 to 1; is the total number of grid cells; For the The normalized density of knowledge points in a grid cell is calculated as follows: ,in For the The number of knowledge points in a grid cell.

[0076] Dimensional dispersion calculation formula: ; Where, Dimension dispersion on ; Dimension The number of grid cells on ; Dimension Previous The normalized knowledge density of each unit.

[0077] The parameter acquisition method is: Obtained by counting the number of knowledge points in each grid cell; Determined according to the grid size parameters set in step S03, for example, a 3×3×3 grid corresponds to ; The number of mesh divisions in each dimension, same as the Meshing Size parameter.

[0078] Grid dispersion is based on the concept of information entropy, which quantifies the uniformity of knowledge distribution. A higher entropy value indicates a more uniform distribution, while a lower entropy value indicates that knowledge is concentrated in a few grid cells. ) makes the dispersion of grids of different sizes comparable. Dimensional dispersion calculations provide detailed information about the distribution characteristics of knowledge in each independent dimension, helping to identify problems of uneven knowledge distribution in certain dimensions.

[0079] The grid knowledge compression function in step S06 can be expressed as: ; The calculation formula of the compressed vector is: ; Information loss assessment index calculation formula: ; Where, For knowledge points The original vector representation of ; is the grid centrality weight; is the knowledge correlation coefficient; is the compression ratio parameter; is the information importance threshold; is the compressed knowledge representation vector; is an information loss assessment indicator; For Perform PCA dimensionality reduction to dimensional operations, where ; for The unit matrix of order; is the average value of the gridded centrality of all knowledge points; Enhance the matrix for important features, by the importance exceeding the threshold The characteristic composition of For knowledge points The importance index of calculate.

[0080] The parameter acquisition method is: Obtained from step S03, usually a 300-768 dimensional vector; Obtained from the gridding centrality calculated in step S03; Extracted from the knowledge relevance matrix formed in step S02, the calculation formula is: ; It is a manually set compression ratio parameter, usually set to 0.3 to 0.7; The importance threshold of retained information determined based on knowledge importance analysis is usually 0.6; The matrix is ​​obtained by analyzing the reconstruction error of the original vector from the autoencoder. The features with large reconstruction error are considered more important. As important features, the corresponding positions are set to enhancement coefficients of 0.5 to 1, and the rest of the positions are set to 0.

[0081] The grid knowledge compression function combines PCA dimensionality reduction with a key feature retention mechanism. PCA dimensionality reduction provides the basic compression effect, while the feature enhancement matrix adjusts the degree of retention of key information based on the importance of the knowledge points. Information loss is assessed based on reconstruction error and knowledge importance, with a lower information loss rate for important knowledge and a higher compression rate for less important knowledge. The compression ratio parameter controls the overall compression strength, while the threshold parameter determines the threshold of information to be retained.

[0082] The adaptive optimization algorithm in step S07 can be expressed as: ; Where, For the The grid structure parameters of the iteration; is the updated grid structure parameter; is the learning rate, usually set to 0.01~0.05; is the gradient of the objective function with respect to the grid structure; is the evaluation function value of the current grid structure; For the The temperature parameter of the iteration is calculated as follows: ,in is the initial temperature (usually set to 1), is the cooling rate (usually set to 0.95-0.99).

[0083] The objective function is defined as: ; Where, is the knowledge coverage of the grid structure; is the retrieval efficiency index; is the mesh complexity; This is the grid structure of the previous version; is the weight coefficient and satisfies .

[0084] The parameter acquisition method is: Finite difference method is used for approximate calculation, and the objective function change is calculated after a small perturbation of each grid parameter; It is obtained by calculating the proportion of knowledge points that can be quickly retrieved in the grid structure to the total knowledge points; Calculated by simulating the inverse of the average response time of multiple random retrieval requests; Computed as a weighted sum of mesh connectivity complexity and number of dimensions; Determined based on actual application requirements, usually .

[0085] The adaptive optimization algorithm combines gradient descent and simulated annealing strategies. Gradient descent provides optimization direction, while simulated annealing increases the likelihood of escaping local optima. The objective function design balances knowledge coverage, retrieval efficiency, structural complexity, and structural stability. The latter uses a sigmoid function to smooth the penalty for structural changes and avoid overly drastic structural adjustments.

[0086] The grid boundary fuzzy processing mechanism in step S07 adopts fuzzy set theory and can be expressed as: ; Where, For knowledge points For grid cells The degree of membership; For knowledge points to grid cells center distance; is the fuzzy radius parameter, which is usually set to 0.5 to 1 times the side length of the grid unit; is a shape parameter that controls the rate of membership decrease and is usually set to 2.

[0087] Cross-grid knowledge mapping calculation formula: ; Where, For knowledge points In the grid cell In the representation; For knowledge points The original vector representation of ; For grid cells A collection of adjacent grid cells; For knowledge points With grid cells The average association strength of all knowledge points in .

[0088] The parameter acquisition method is: It is obtained by calculating the Euclidean distance between the feature vector of the knowledge point and the center vector of the grid unit in the feature space; Determined by the grid unit size, if the grid unit side length is ,but Usually set to ; Determined by fuzzy boundary experiment optimization, usually 2; Extracted from the knowledge relevance matrix, the calculation formula is: ,in For grid cells The number of knowledge points.

[0089] The mesh boundary fuzzification mechanism uses an improved Bell-type membership function, which provides smoother boundary transitions compared to traditional triangular or trapezoidal membership functions. The cross-mesh knowledge mapping formula combines the original representation and the influence of adjacent meshes. Through membership weighting, it ensures a natural transition of knowledge at boundaries, avoiding knowledge retrieval blind spots caused by hard boundaries.

[0090] Specifically, the present invention is based on grid-based knowledge representation theory and multi-source data fusion technology. It expresses and organizes hazardous chemical knowledge by constructing a multidimensional grid structure, thereby reducing reliance on manual operations and enhancing the expression of knowledge relevance. Its core principle is to transform traditional linear or hierarchical knowledge structures into a grid-based structure, allowing hazardous chemical knowledge to form an interconnected network in a multidimensional space. Each grid cell represents a knowledge set under a specific attribute combination, and the connections between grid cells represent the relationships between knowledge points.

[0091] To reduce manual effort, this invention automatically screens data sources using a recognition index matrix. By calculating the degree to which each knowledge point is recognized by multiple authoritative data sources, an objective data quality assessment system is established, replacing the traditional screening process that relies on manual judgment. The accuracy assessment model automatically calculates the accuracy index of knowledge points through a set of five key equations and automatically labels redundant, conflicting, erroneous, and defective knowledge, significantly reducing manual intervention.

[0092] The innovation of this invention in enhancing the expression of knowledge relevance lies in the construction of a comprehensive grid property measurement system. By calculating grid centrality, the importance and influence of knowledge points within the overall grid structure are determined; by setting the grid size, knowledge representation at varying granularities is achieved; and by calculating grid aggregation and dispersion, knowledge-intensive regions and knowledge distribution characteristics are identified. Together, these metrics form the quantitative foundation for expressing knowledge relevance, enabling the precise capture of complex patterns in the association of chemical hazard properties.

[0093] The chemical knowledge grid representation model of this invention utilizes a multi-layer bidirectional transformer network architecture. It automatically learns the associations between different chemical properties through a multi-head grid attention mechanism. Kruskal's algorithm and matrix completion algorithm automatically handle knowledge conflicts and deficiencies. The grid knowledge compression function and adaptive optimization algorithm further enhance the system's processing capabilities, enabling the knowledge base to dynamically adjust the grid structure based on knowledge updates. This achieves a dual breakthrough in both theoretical and technical aspects: reducing manual reliance and enhancing knowledge relevance.

[0094] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows.

[0095] The specific implementation of step S01 is to first collect hazardous chemical data from multiple sources such as the National Hazardous Chemicals Database, International Chemical Safety Cards, and academic journal databases through web crawler technology, database interfaces, and document parsing tools. Then, a hash table data structure is used to organize multi-source data, and data fusion technology is used to merge the same knowledge points from different sources into a unified representation to form a dimensional data matrix, where Indicates the number of knowledge points, Represents the number of data sources. The recognition index of each data source is then calculated based on factors such as the data source's authority, update frequency, data completeness, and accuracy history. The value ranges from 0 to 1, and data sources with a threshold of 0.6 or above are generally considered reliable. The following formula is used to calculate the recognition index matrix: , where For knowledge points In the data source The recognition index in For data sources The reliability coefficient ranges from 0 to 1; For knowledge points In the data source The existence status of , if it exists, it is 1, if it does not exist, it is 0; For data sources The time factor represents the timeliness of data source update, and the calculation formula is ,in is the time attenuation coefficient, which is generally set between 0.1 and 0.3. Finally, the comprehensive recognition index of each knowledge point is calculated using the weighted average algorithm: , where For knowledge points Comprehensive recognition index; For data sources The weight coefficient of = is the total number of data sources. Knowledge points with a recognition index below 0.5 are filtered out, retaining high-quality data for the next step. This step aims to ensure the data quality and reliability of the initial knowledge system, laying the foundation for subsequent knowledge base construction.

[0096] The specific implementation of step S02 is to first construct an accuracy evaluation model, which includes five core equations. The basic accuracy calculation equation adopts a multi-factor weighted summation method: , where For knowledge points The basic accuracy index ranges from 0 to 1; For knowledge points Recognition index; For knowledge points The data source reliability coefficient ranges from 0.1 to 0.9; For knowledge points The difference between the release time and the current time (in years); is the time attenuation factor, ranging from 0.1 to 0.3; For knowledge points The average expert rating of , with a full score of 10 points; are the weight coefficients of each factor, and satisfy ,generally The weight adjustment equation generates the adjustment coefficients through nonlinear combination: , where For knowledge points The weight adjustment coefficient of For knowledge points Citation frequency (normalized to 0–1); For knowledge points Application scenario coverage (percentage value); For knowledge points The amount of associated knowledge; For knowledge points Danger level coefficient (1 to 8); are the weight coefficients of each factor, and satisfy ,generally The conflict penalty equation calculates the penalty value based on multiple factors: , where For knowledge points The conflict penalty value; For knowledge points the number of conflicts; is the conflict severity coefficient (0-1); For knowledge points The average accuracy of conflicting knowledge points; is the difficulty level of conflict resolution (1 to 5); are the weight coefficients of each factor, and satisfy ,generally The experimental verification adjustment equation comprehensively calculates the adjustment coefficient of multiple factors: , where For knowledge points The experimental verification adjustment coefficient of is the number of experimental verifications; is the consistency coefficient of experimental data (0.5-1); Score the stringency of the experimental conditions (1–10); is the reliability coefficient of the experimental method (0.6-0.95); is the standard deviation of the experimental results; are the weight coefficients of each factor, and satisfy ,generally The final accuracy synthesis equation performs a weighted combination of the above outputs: , where For knowledge points The final accuracy index of is the basic accuracy index; is the weight adjustment coefficient; is the conflict penalty value; Adjustment coefficients for experimental verification; is a smoothing factor, generally ranging from 0.01 to 0.1. Then, a hierarchical clustering algorithm is used to classify and analyze knowledge points. Redundant knowledge is marked based on semantic similarity and content overlap. Conflicting knowledge is discovered through logical relationship verification of knowledge points. Wrong knowledge is identified based on the accuracy index. Defective knowledge is discovered through integrity assessment. Finally, a knowledge relevance matrix is ​​constructed. ,in Representing knowledge points And knowledge points The correlation strength between them is calculated as follows: , where is the knowledge point vector and The cosine similarity of It is the support function in association rule mining; For knowledge points and The number of co-occurrences of is the maximum number of co-occurrences among all knowledge point pairs; is the weight coefficient and satisfies ,generally The purpose of this step is to establish a scientific knowledge evaluation system and improve the quality of the knowledge base.

[0097] The specific implementation of step S03 is to first define a multidimensional knowledge space based on the physical and chemical properties, hazard categories, usage conditions, storage requirements, and other attributes of hazardous chemicals, and then use a grid partitioning algorithm to divide the knowledge space into regular units. Then, the grid centrality of each knowledge point is calculated, specifically using the eigenvector centrality algorithm: , where For knowledge points The grid centrality of For knowledge points and The strength of association; For knowledge points The final accuracy index of For knowledge points The connectivity of the (the number of knowledge points directly connected to it) is the maximum connectivity among all knowledge points; For knowledge points Betweenness centrality; The maximum betweenness centrality among all knowledge points. Centrality values ​​range from 0 to 1, and knowledge points with a centrality greater than 0.8 are typically considered core knowledge. Next, grid size parameters are set. Generally, a 3- to 8-dimensional grid structure is selected based on the data scale and application requirements. Smaller sizes (such as 3×3×3) provide detailed representation, while larger sizes (such as 8×8×8) provide a macro overview. Finally, a multidimensional knowledge grid space is constructed, using tensor representation to store knowledge information across various dimensions. This complete gridded knowledge mapping structure enables multidimensional knowledge organization and rapid retrieval. This step aims to establish a structured knowledge organization approach, providing a spatial framework for subsequent conflict resolution and risk association.

[0098] The specific implementation of step S04 is to first construct a relationship graph between knowledge points, where nodes represent knowledge points, edges represent relationships between knowledge points, and edge weights represent the intensity of conflicts between knowledge points. Then, the minimum spanning tree is constructed using the Kruskal algorithm. This algorithm selects the knowledge connection method with the least conflict, effectively avoids cyclic conflicts, and is suitable for handling knowledge conflicts in complex networks. Then, a knowledge conflict degree matrix is ​​constructed. ,in Representing knowledge points And knowledge points The degree of conflict between them is calculated as follows: , where is the knowledge point vector and Semantic similarity of The strength of logical conflicts is detected by predicate logic rules; is the degree of numerical conflict, which indicates the degree of inconsistency of numerical knowledge; is the maximum numerical conflict degree among all knowledge point pairs; is the weight coefficient and satisfies ,generally For defect knowledge, a matrix completion algorithm based on low-rank assumption is applied, and its model can be expressed as: , where is the complete matrix to be completed; is the observation matrix containing missing values; is the projection operator, corresponding to the set of observed entries ; is a matrix the nuclear norm of (the sum of all singular values); is a regularization parameter used to balance fitting error and matrix complexity, usually ranging from 0.1 to 1. The confidence calculation formula for the completed knowledge point is: , where To complete the value confidence level; is the predicted value in cross validation; is the standard deviation of the prediction error; For knowledge points The number of relevant missing knowledge points; is the total number of knowledge points in the knowledge base. Knowledge point pairs with a conflict degree greater than 0.6 are usually marked as severely conflicting and are processed first. Finally, the gridded clustering degree is calculated using the Moran index method in spatial statistics: , where is the Moran index, and its value range is usually -1 to 1; is the total number of knowledge points; is the spatial weight matrix element, reflecting the knowledge point and distance relations in grid space; For knowledge points Attribute values ​​(such as importance or access frequency); is the average value of all knowledge point attribute values. The calculation formula for local grid aggregation is: , where For knowledge points The degree of clustering in the local area. Areas with a clustering degree greater than 0.8 are identified as knowledge-intensive areas, while areas with a clustering degree less than 0.2 are identified as knowledge-gap areas. The purpose of this step is to resolve knowledge conflicts, improve the knowledge structure, and enhance the consistency of the knowledge base.

[0099] The specific implementation of step S05 is to first input the physical and chemical property data of hazardous chemicals (such as flash point, boiling point, toxicity index, etc.) and grid block information (from step S03) as input feature vectors and input them into a pre-trained chemical knowledge grid representation model. This model uses a multi-layer bidirectional transformer architecture to capture the complex nonlinear relationships between chemical properties. Next, a chemical risk association network is constructed, using graph neural network technology to establish possible reactions, synergistic toxicity, and cascading risk relationships between hazardous chemicals. The association strength threshold is typically set at 0.65. The hazardous chemical grid index is then calculated: , where For chemicals The grid risk index ranges from 0 to 10; The inherent hazard index of chemicals is assessed based on the GHS classification standard, with a value range of 0 to 10; is the knowledge completeness index, which indicates the coverage of relevant knowledge points; is the completeness threshold parameter, usually ranging from 3 to 5; is the weight coefficient and satisfies ,generally The calculation formula of knowledge completeness index is: , where For chemicals A collection of related knowledge points; For knowledge points The final accuracy index of For knowledge points The grid centrality of For collection The index ranges from 0 to 10, where 8 to 10 indicates extremely high risk, 6 to 8 indicates high risk, 4 to 6 indicates moderate risk, 2 to 4 indicates low risk, and 0 to 2 indicates slight risk. The gridded dispersion analysis method is then applied, using the entropy calculation model: , where is the grid dispersion, ranging from 0 to 1; is the total number of grid cells; For the The normalized density of knowledge points in a grid cell is calculated as follows: ,in For the The number of knowledge points in a grid cell. The formula for calculating dimensional dispersion is: , where Dimension dispersion on ; Dimension The number of grid cells on ; Dimension Previous The normalized knowledge density of each unit is calculated. A dispersion greater than 0.7 indicates uniform knowledge distribution, while a dispersion less than 0.3 indicates extremely uneven knowledge distribution, indicating the need for supplementary domain knowledge. Finally, a mapping relationship is established between grid nodes, using tensor mapping functions to connect knowledge points of different dimensions to form a complete knowledge network. This step aims to establish a hazardous chemical risk association system and improve risk identification accuracy.

[0100] The specific implementation of step S06 involves first designing a knowledge reasoning mechanism that integrates rule-based reasoning based on first-order predicate logic and statistical reasoning based on Bayesian networks to achieve both precise and uncertain reasoning about hazardous chemical knowledge. Next, for extremely long knowledge content (typically knowledge points with more than 1,000 dimensions or text length exceeding 5,000 characters), a grid knowledge compression function is applied to perform dimensionality reduction: , where the calculation formula for the compressed vector is: , the information loss evaluation index calculation formula is: , where For knowledge points The original vector representation of ; is the grid centrality weight; is the knowledge correlation coefficient; is the compression ratio parameter; is the information importance threshold; is the compressed knowledge representation vector; is an information loss assessment indicator; For Perform PCA dimensionality reduction to dimensional operations, where ; for The unit matrix of order; is the average value of the gridded centrality of all knowledge points; Enhance the matrix for important features; For knowledge points The importance index of Calculation. This compression function combines principal component analysis and autoencoders. Its inputs include a knowledge vector representation, a gridded centrality weight (typically 0.1–0.9), a knowledge relevance coefficient (typically 0.1–0.8), a compression ratio parameter (typically set to 0.3–0.7), and an information importance threshold (typically 0.6). The compression function outputs a compressed knowledge representation vector (typically with a 50%–80% reduction in dimensionality) and an information loss assessment index (typically required to be less than 0.2). A hierarchical knowledge representation model is then constructed, organizing knowledge in a tree-like structure. Upper-level nodes store abstract concepts, while lower-level nodes store specific details, optimizing the storage structure. Finally, knowledge simplification and key information retention are achieved. Knowledge importance is ranked, and key information with an importance index greater than 0.7 is retained. This step aims to optimize knowledge representation and storage efficiency while ensuring the integrity and accessibility of key information.

[0101] The specific implementation of step S07 is to first implement an adaptive optimization algorithm based on the chemical knowledge grid representation model: , where For the The grid structure parameters of the iteration; is the updated grid structure parameter; is the learning rate, usually set to 0.01~0.05; is the gradient of the objective function with respect to the grid structure; is the evaluation function value of the current grid structure; For the The temperature parameter of the iteration is calculated as follows: ,in is the initial temperature (usually set to 1), is the cooling rate (usually set to 0.95-0.99). The objective function is defined as: , where is the knowledge coverage of the grid structure; is the retrieval efficiency index; is the mesh complexity; This is the grid structure of the previous version; is the weight coefficient and satisfies ,generally . The algorithm combines gradient descent and simulated annealing strategies, and can dynamically adjust the grid structure parameters according to the knowledge update situation. The algorithm convergence condition is usually set to the grid structure change rate of less than 0.01 for 5 consecutive iterations or reaching the maximum number of iterations of 100 times. Then, a differentiated storage strategy is designed according to the knowledge update frequency. Knowledge points with high update frequency (such as more than 1 update per month) are stored in the fast access layer, knowledge points with medium update frequency (such as 1 update per quarter) are stored in the standard access layer, and knowledge points with low update frequency (such as less than 1 update per year) are stored in the archive layer. Then, a grid boundary fuzzy processing mechanism is established, using fuzzy set theory: , where For knowledge points For grid cells The degree of membership; For knowledge points to grid cells center distance; is the fuzzy radius parameter, which is usually set to 0.5 to 1 times the side length of the grid unit; is a shape parameter that controls the membership decrease rate and is usually set to 2. The calculation formula for cross-grid knowledge mapping is: , where For knowledge points In the grid cell In the representation; For knowledge points The original vector representation of ; For grid cells A collection of adjacent grid cells; For knowledge points With grid cells The average association strength of all knowledge points in the grid is calculated. The membership threshold is typically set between 0.4 and 0.6 to address the ambiguity of knowledge attribution at the grid boundaries. Finally, the grid structure is intelligently adjusted. Based on knowledge base usage feedback and newly added knowledge features, grid parameters, including grid dimension, grid size, and grid connectivity, are automatically optimized to form the final gridded hazardous chemicals knowledge base.

[0102] To better understand and implement the present invention, Example 2, a specific application scenario, is provided below: Over 200 hazardous chemicals are concentrated within a chemical park. Park researchers decided to establish a grid-based hazardous chemical knowledge base to improve safety management. First, the researchers collected 2,650 hazardous chemical knowledge points from eight data sources: the National Hazardous Chemicals Database, the U.S. NIOSH Chemical Hazard Database, the EU ECHA Chemical Database, the China Material Safety Data Sheet (MSDS) system, chemical industry journals, and accident case reports.

[0103] In the first step, the researchers calculated the recognition index of each data source. The results are shown in Table 1: Table 1 Evaluation results of the hazardous chemicals data source recognition index

[0104] Figure 2 The results of the recognition index evaluation of the 8 hazardous chemical data sources in the embodiment are shown. The horizontal axis represents the names of different data sources, and the vertical axis represents the numerical values ​​of various indicators. Each data source has 4 indicators: reliability coefficient (CR₍ⱼ₎), time factor (TF₍ⱼ₎), weight coefficient (W₍ⱼ₎) and the final recognition index. The figure also marks the recognition threshold of 0.5 with a red dotted line. Data sources below this threshold will be excluded. As can be seen from the figure, the US NIOSH Chemical Hazard Database has the highest recognition index of 0.85, while the safety training materials have the lowest recognition index of 0.48 and are excluded from the knowledge base. This chart intuitively shows the quality assessment results of each data source in the first step of the embodiment. After screening, knowledge sources with a recognition index lower than 0.5 were excluded, and a total of 7 data sources were retained. For each knowledge point, the researchers calculated the comprehensive recognition index and filtered out knowledge points with a recognition index lower than 0.5, ultimately retaining 2,245 high-quality knowledge points.

[0105] In the second step, the researchers constructed an accuracy assessment model. Using knowledge points on two common hazardous chemicals, benzene and toluene, as examples, they calculated their final accuracy index. Some of the results are shown in Table 2: Table 2 Accuracy evaluation results of some hazardous chemicals knowledge points

[0106] Figure 3 The radar chart shows the ) and toluene ( ) Accuracy evaluation results of 6 knowledge points of two common hazardous chemicals. The figure contains 5 curves, representing the basic accuracy, the normalized weight adjustment coefficient, the amplified conflict penalty value, the amplified experimental verification adjustment coefficient and the final accuracy index. Through the form of a radar chart, the performance of different knowledge points in each evaluation dimension can be intuitively compared. For example, K0135 (the flash point of benzene) has the highest final accuracy index of 0.94, while K0216 (the LD50 value of toluene) is relatively low. This chart corresponds to the accuracy evaluation model part of the second step in the embodiment. At the same time, the researchers constructed a knowledge correlation matrix to determine the correlation strength between knowledge points. For example, the correlation strength between the flash point knowledge point of benzene and its explosion hazard knowledge point is 0.82, and the correlation strength with the storage conditions knowledge point is 0.75.

[0107] In the third step, the researchers designed a four-dimensional grid structure based on the four dimensions of "physical and chemical properties," "hazard classification," "storage requirements," and "emergency disposal." Each dimension was divided into five levels, forming a 5×5×5×5 grid space with a total of 625 grid cells. The grid centrality of key knowledge points was calculated, and some of the results are shown in Table 3. Table 3 Calculation results of grid centrality of key knowledge points

[0108] Figure 4 It is a bubble scatter plot that shows the grid centrality analysis results of 5 key knowledge points. The horizontal axis represents betweenness centrality, the vertical axis represents the sum of association strengths, the bubble size represents connectivity, and the color depth represents grid centrality. A red dotted line is also added to the figure to represent the trend line, revealing the linear relationship between betweenness centrality and the sum of association strengths. As can be seen from the figure, K0311 (sulfuric acid) has the highest grid centrality (0.92) and the largest betweenness centrality (1562.7), indicating that this knowledge point has an important position in the grid structure. This chart corresponds to the grid centrality calculation part of the third step in the embodiment. Figure 5 It is a three-dimensional scatter plot that shows the distribution of hazardous chemical knowledge in a four-dimensional grid structure. The three coordinate axes in the figure represent the physical and chemical properties dimension, the hazard category dimension, and the storage requirement dimension, each of which is divided into five levels. The color and size of the points represent the value of the fourth dimension - the emergency response dimension. The figure also marks the locations of five high-risk chemicals: benzene ( ), acetylene ( ), sodium cyanide (NaCN), sulfuric acid ( ) and toluene ( This 3D visualization clearly shows the distribution of hazardous chemicals within the 4D grid space, particularly the tendency of high-risk chemicals to be concentrated in specific areas of the grid space. This diagram corresponds to the 4D grid structure designed in step 3 of the example.

[0109] In the fourth step, the researchers applied the Kruskal algorithm to construct a minimum spanning tree of knowledge and processed 82 sets of conflicting knowledge points. For example, the hazardous properties of nitric acid conflicted in different data sources, and the most reliable description was determined by calculating the conflict degree matrix. At the same time, the matrix completion algorithm was used to process 157 defective knowledge points, such as some physical property data of certain new composite hazardous chemicals. The researchers calculated the grid aggregation degree and identified that the knowledge density of the flammable and explosive chemicals area was 0.86, which is highly aggregated; while the knowledge density of the new nanomaterials area was only 0.23, indicating a significant knowledge gap.

[0110] In the fifth step, the researchers input the hazardous chemical attribute data and grid block information into a pre-trained chemical knowledge grid representation model to construct a chemical risk association network. The hazardous chemical grid index was calculated based on the 36 high-risk chemicals in the park. Some of the results are shown in Table 4: Table 4 Calculation results of grid index for some hazardous chemicals

[0111] Applying grid dispersion analysis, the overall knowledge dispersion was calculated to be 0.65, which is at a medium uniform level; the dispersion of the physical and chemical properties dimension was 0.82, while the dispersion of the emergency response dimension was only 0.48, indicating that emergency knowledge was unevenly distributed. Figure 6 It is a bubble scatter plot that shows the grid risk index analysis results of 8 major hazardous chemicals. The horizontal axis represents the inherent hazard index, the vertical axis represents the knowledge completeness, the bubble size represents the grid risk index, and the color represents the risk level (red represents extremely high risk, orange represents high risk). The figure divides the entire area into four quadrants and uses different color backgrounds to mark them: the yellow area represents the low risk-high knowledge area, the red area represents the high risk-high knowledge area, the blue area represents the low risk-low knowledge area, and the purple area represents the high risk-low knowledge area (the most dangerous). As can be seen from the figure, benzene ( ) has the highest grid risk index of 8.82, which is an extremely high risk level; while sodium chlorate ( ) has a low grid risk index of 7.46, which is a high risk level. This chart corresponds to the calculation results of the hazardous chemical grid index in step 5 of the embodiment.

[0112] In the sixth step, the researchers designed a knowledge inference mechanism, employing a grid knowledge compression function to reduce the dimensionality of extremely long knowledge. For example, using a text describing the physical and chemical safety properties of ammonium nitrate (originally 8520 characters long) as an example, the compression function yielded a knowledge representation vector with a 72% reduction in dimensionality. The information loss assessment index was 0.17, exceeding the threshold of 0.2. A hierarchical knowledge representation model was employed to optimize the storage structure. The top layer encompasses abstract concepts such as "flammable and explosive," "toxic," and "corrosive," which are then refined layer by layer to the specific parameters of specific substances, achieving knowledge simplification while retaining key information.

[0113] In the seventh step, the researchers implemented an adaptive optimization algorithm based on the chemical knowledge grid representation model to dynamically adjust the grid structure. Initially, the algorithm learning rate was set to 0.03, the initial temperature was 1.0, and the cooling rate was 0.97. After 83 iterations of optimization, the grid structure change rate dropped to 0.008, reaching the convergence condition. Based on the frequency of knowledge updates, the researchers designed a three-tier storage strategy: a quick access layer (updated monthly, 265 knowledge points), a standard access layer (updated quarterly, 1582 knowledge points), and an archive layer (updated annually, 398 knowledge points). A grid boundary fuzzy processing mechanism was established, and the membership threshold was set to 0.5 to solve the problem of cross-grid knowledge mapping. The resulting grid-based hazardous chemical knowledge base contains 625 grid cells, 2245 knowledge points, and a total storage capacity of approximately 850MB.

[0114] Since the implementation of this grid-based hazardous chemical knowledge base, the park's safety management efficiency has significantly improved. Chemical risk assessment time has been reduced from an average of 4.5 hours to 0.8 hours, and identification accuracy has increased from 85% to 96.5%. Emergency response plan generation time has been reduced from an average of 28 minutes to 5 minutes, and precise guidance can be provided for complex multi-chemical incidents.

[0115] Traditional hazardous chemical knowledge management mainly uses relational databases or simple knowledge graph methods, which face problems such as uneven data quality, difficult to handle knowledge conflicts, limited expression of associations, and low knowledge retrieval efficiency. Traditional methods rely solely on simple hierarchical classification and keyword indexing, making it difficult to handle the complex associations between chemicals and unable to adapt to the needs of dynamic knowledge updates. Compared with traditional means, the grid-based hazardous chemical knowledge base establishment method of the present invention has brought the following improvements: First, a systematic knowledge quality assessment system is established, and knowledge quality is ensured through dual screening of recognition index and accuracy index; second, the use of a multi-dimensional grid structure to organize knowledge breaks through the limitations of traditional planar knowledge organization and can quickly retrieve and reason from multiple angles; third, the innovative introduction of Kruskal algorithm and matrix completion algorithm effectively solves the problems of knowledge conflicts and defects; fourth, through the calculation of three indicators of grid centrality, aggregation and dispersion, the importance and distribution characteristics of knowledge are comprehensively characterized; fifth, an adaptive optimization algorithm and boundary fuzzy processing mechanism are designed to achieve dynamic optimization and continuous improvement of the knowledge base, greatly improving the level of chemical risk management.

[0116] It should be noted that the variables involved in the present invention are explained in detail as shown in Tables 5, 6 and 7 below.

[0117] Table 5 Variable Explanation Table (Part I)

[0118] Table 6 Variable Explanation Table (Part II)

[0119] Table 7 Variable Explanation Table (Part 3)

[0120] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.

Claims

1. A method for establishing a grid-based hazardous chemicals knowledge base, characterized in that: include: Collect hazardous chemicals data to establish an initial knowledge system, use multi-source data fusion technology to form a data matrix, and screen data sources based on the recognition index; Construct an accuracy assessment model, mark redundant knowledge, conflicting knowledge, erroneous knowledge, and defective knowledge, and form a knowledge relevance matrix; Establish a grid knowledge mapping structure, calculate the grid centrality and determine the weight of the knowledge points; Kruskal algorithm is used to construct the knowledge minimum spanning tree to deal with knowledge conflicts, and matrix completion algorithm is used to deal with defective knowledge. Construct a chemical risk association network and calculate the hazardous chemical grid index to quantify the hazard level; Design a knowledge reasoning mechanism and use grid knowledge compression function to reduce dimensionality; Implement adaptive optimization algorithms to dynamically adjust the grid structure and form a grid-based hazardous chemicals knowledge base.

2. The method for establishing a grid-based hazardous chemicals knowledge base according to claim 1, characterized in that: The accuracy evaluation model includes a basic accuracy calculation equation, a weight adjustment equation, a conflict penalty equation, an experimental verification adjustment equation, and a final accuracy synthesis equation; the basic accuracy calculation equation is used to calculate the initial accuracy value of the knowledge point; and the final accuracy synthesis equation is used to comprehensively calculate the final accuracy index based on various factors.

3. The method for establishing a grid-based hazardous chemicals knowledge base according to claim 2, characterized in that: In establishing a grid knowledge mapping structure, knowledge grid units are divided according to the properties of hazardous chemicals, and the grid size is set to divide the knowledge granularity to form a multi-dimensional knowledge grid space; the grid size refers to the size of the knowledge grid division granularity.

4. The method for establishing a grid-based hazardous chemicals knowledge base according to claim 3, characterized in that: Grid centrality is an indicator that describes the importance of a knowledge point in the entire grid structure.

5. The method for establishing a grid-based hazardous chemicals knowledge base according to claim 4, characterized in that: The knowledge conflict degree matrix is ​​a two-dimensional array that records contradictory or inconsistent information in the knowledge base and is used to identify conflicting knowledge that needs to be resolved. The edge weights are used to represent the intensity of conflicts between knowledge points. In the analysis of knowledge distribution characteristics using grid dispersion, grid dispersion is an indicator that describes the uniformity of the distribution of knowledge points in the grid space.

6. The method for establishing a grid-based hazardous chemicals knowledge base according to claim 5, characterized in that: The hazardous chemicals grid index is a composite indicator that comprehensively considers the hazardous characteristics of chemicals and the completeness of knowledge. It is used to quantitatively assess the risk level of chemicals; it establishes a mapping relationship between grid nodes to improve the accuracy of risk identification.

7. The method for establishing a grid-based hazardous chemicals knowledge base according to claim 6, characterized in that: The grid knowledge compression function is used to reduce the dimensionality of extremely long knowledge in the hazardous chemicals knowledge base and retain key information. The input includes knowledge vector representation, grid centrality weight, knowledge correlation coefficient, compression ratio parameter, and information importance threshold. The output is the compressed knowledge representation vector and information loss evaluation index.

8. The method for establishing a grid-based hazardous chemicals knowledge base according to claim 7, characterized in that: The specific structure of the chemical knowledge grid representation model is a multi-layer bidirectional transformer network architecture, which includes a hazardous chemical word embedding layer, a multi-head grid attention layer, a feedforward neural network layer, and a grid representation output layer. The parameters of the grid multi-head attention mechanism are jointly determined by the grid centrality, grid size, and knowledge relevance matrix.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the method for establishing a grid-based hazardous chemical knowledge base according to any one of claims 1 to 8.

10. A grid-based hazardous chemicals knowledge base establishment system, characterized in that: The computer-readable storage medium according to claim 9 is included, the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

Citation Information

Patent Citations

  • Knowledge cube model algorithm based on graph theory

    CN103309979A

  • Knowledge question-answering system and method based on chemical knowledge base

    CN111522934A

  • Image emotion recognition method based on fuzzy knowledge neural network

    CN112836718A

  • Image fusion classification method based on minimum spanning tree and evidence theory

    CN118411548A

  • Method for constructing chemical-plastic industry chain knowledge graph by using graph convolutional network

    CN119250172A