Chemical semantic recognition method and system

Through the chemical natural language processing method based on regular expressions and chemical physical analysis model, the problem of low semantic recognition accuracy in the chemical industry is solved, and efficient and accurate recognition and feedback of the Chinese intelligent question-and-answer system of chemical industry is realized.

CN120387456APending Publication Date: 2025-07-29CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410110264.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing semantic recognition model has low recognition accuracy in the chemical industry and cannot accurately identify user consultation problems, resulting in feedback content errors or delayed push.

Method used

The chemical natural language data preprocessing method based on regular expression is used to construct a chemical physical natural language semantic analysis model, and the chemical text is recognized by the hidden Markov model and the conditional random field model, and the chemical vocabulary library and the physical property auxiliary analysis model are combined to generate feedback content.

Benefits of technology

It improves the accuracy and efficiency of natural language processing in the chemical industry, and realizes the accurate identification and feedback of the Chinese intelligent question and answer system in the chemical industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387456A_ABST
    Figure CN120387456A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a chemical semantic recognition method, and belongs to the technical field of refining devices. The method comprises the steps of collecting to-be-recognized text information, and preprocessing the text information based on a regular expression; performing word segmentation processing on the preprocessed text information based on a pre-constructed chemical word segmentation library; taking the text information subjected to word segmentation processing as input parameters, and sequentially executing training of a plurality of chemical physical property auxiliary analysis models to obtain a chemical physical property semantic recognition result; and generating feedback content based on the chemical physical property semantic recognition result. According to the scheme, through application of the regular expression, the physical property-based hidden Markov model and the conditional random field model, the machine recognition accuracy and efficiency of the chemical text are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic recognition, and particularly to a chemical semantic recognition method and a chemical semantic recognition system. Background Art

[0002] Since chemical enterprises involve many risk devices during operation, their operational stability is very important. To ensure the stable operation of chemical equipment, relevant operators need to have sufficient operating knowledge. If there are uncertain operation contents, the operators need to be able to quickly obtain relevant knowledge. Based on this, operators have a great demand for chemical knowledge consultation. To meet the response speed, a corresponding intelligent question-and-answer system is essential. With the rapid development of the Internet and artificial intelligence technologies, the Chinese word segmentation technology in natural language processing has become increasingly mature, and semantic understanding has begun to be applied industrially. However, for the chemical industry, the existing semantic recognition models do not have the corresponding industry characteristic recognition ability, resulting in insufficient generalization ability and low recognition accuracy of the existing semantic recognition models. This leads to the inability to accurately identify the user's consultation questions, resulting in incorrect or delayed push of feedback content. Aiming at the problem of low recognition accuracy of the existing semantic recognition solutions for the chemical industry, a new semantic recognition solution for the chemical industry needs to be proposed. Summary of the Invention

[0003] The purpose of the embodiments of the present invention is to provide a chemical semantic recognition method and system to at least solve the problem of low recognition accuracy of the existing semantic recognition solutions for the chemical industry.

[0004] To achieve the above purpose, the first aspect of the present invention provides a chemical semantic recognition method, the method includes: collecting text information to be recognized, and preprocessing the text information based on regular expressions; performing word segmentation processing on the preprocessed text information based on a pre-constructed chemical word segmentation library; using the text information after word segmentation processing as input parameters, and sequentially performing training of multiple chemical physical property auxiliary analysis models to obtain chemical physical property semantic recognition results; generating feedback content based on the chemical physical property semantic recognition results.

[0005] Optionally, the preprocessing the text information based on regular expressions includes: extracting information on the list of hazardous chemicals based on an open-source chemical knowledge graph; performing regularization processing on the typical chemical materials in the text information based on the information on the list of hazardous chemicals.

[0006] Optionally, the performing word segmentation processing on the preprocessed text information based on a pre-constructed chemical word segmentation library includes: performing word segmentation processing on the preprocessed text information based on an initial semantic recognition model; performing dictionary matching on the text information after word segmentation processing based on the pre-constructed chemical word segmentation library, and performing word segmentation adjustment.

[0007] Optionally, the matching rule for dictionary matching of the text information after word segmentation is based on the forward maximum matching method, the reverse maximum matching method, or the bidirectional matching method.

[0008] Optionally, the chemical property auxiliary analysis model includes: a chemical property semantic analysis auxiliary model, a property auxiliary analysis model based on hidden Markov, and a property auxiliary analysis model based on conditional random field.

[0009] Optionally, performing the training of the chemical property semantic analysis auxiliary model includes: classifying the chemical properties of chemicals and grading them based on the degree of hazard; matching chemicals based on regular expressions and matching physical properties based on the matched chemicals; performing physical property semantic auxiliary analysis based on the matched physical properties.

[0010] Optionally, performing the training of the property auxiliary analysis model based on hidden Markov includes: using the output of the chemical property semantic analysis auxiliary model as the input of the property auxiliary analysis model based on hidden Markov, and calculating the counting probability of the text information after word segmentation under each property classification; calculating the property probability value of the statement based on the counting probability of the text information after word segmentation under each property classification, and the calculation rule is:

[0011] P w =max(a 1i ,a 2i ,…,a ni )

[0012] where P w is the property probability value; a ni is the counting probability under the nth property classification.

[0013] Optionally, performing the training of the property auxiliary analysis model based on conditional random field includes: using the output of the property auxiliary analysis model based on hidden Markov as the input of the property auxiliary analysis model based on conditional random field, and generating an undirected graph model, expressed as:

[0014] G=(V,E)

[0015] where G is the generated undirected graph; V is the set of chemicals; E is the set of relationships between each chemical and other word segments; outputting the semantic word segmentation result based on the undirected graph model as the chemical property semantic recognition result.

[0016] Optionally, generating feedback content based on the chemical property semantic recognition result includes: determining the word segmentation result based on the chemical property semantic recognition result; obtaining each analysis semantic and the associated semantic between each word segment and other word segments based on the word segmentation result as semantic elements; generating complete text information based on the semantic elements as feedback content.

[0017] In a second aspect of the present invention, a chemical semantic recognition system is provided. The system includes: a collection unit for collecting text information to be recognized and preprocessing the text information based on regular expressions; a word segmentation unit for performing word segmentation on the preprocessed text information based on a pre-constructed chemical industry word segmentation library; a training unit for using the text information after word segmentation as input parameters to sequentially execute the training of multiple chemical property auxiliary analysis models to obtain a chemical property semantic recognition result; and a feedback unit for generating feedback content based on the chemical property semantic recognition result.

[0018] In a third aspect of the present invention, a computer-readable storage medium is provided. Instructions are stored on the computer-readable storage medium, and when running on a computer, the computer is caused to execute the above-mentioned chemical semantic recognition method.

[0019] In a fourth aspect of the present invention, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned chemical semantic recognition method is implemented.

[0020] Through the above technical solutions, the present invention proposes a method for preprocessing chemical natural language data based on regular expressions, and completes the preprocessing of chemical professional natural language through specific regular expressions. A natural language semantic analysis model based on chemical properties is constructed. By applying chemical property data, it assists in the semantic analysis of chemical natural language, so as to facilitate the development of a chemical professional Chinese intelligent question-answering system. It solves the problems of obscurity and low accuracy in natural language processing in the chemical industry in China, and provides the machine recognition accuracy and efficiency of chemical texts through the application of regular expressions, hidden Markov models based on physical properties, and conditional random field models.

[0021] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific embodiment part. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:

[0023] Figure 1 is a flowchart of the steps of a chemical semantic recognition method provided by an embodiment of the present invention;

[0024] Figure 2 is a schematic diagram of the training process of a chemical property auxiliary analysis model provided by an embodiment of the present invention;

[0025] Figure 3It is the system structure diagram of the chemical semantic recognition system provided by an embodiment of the present invention. Specific embodiments

[0026] The following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.

[0027] Since chemical enterprises involve many risk devices during operation, their operational stability is very important. To ensure the stable operation of chemical equipment, relevant operators need to have sufficient operation knowledge. If there are uncertain operation contents, the operators need to be able to quickly obtain relevant knowledge. Based on this, operators have a great demand for chemical knowledge consultation. To meet the response speed, a corresponding intelligent question-answering system is essential. With the rapid development of the Internet and artificial intelligence technologies, the Chinese word segmentation technology in natural language processing has become increasingly mature, and semantic understanding has begun to be applied industrially. However, for the chemical industry, the existing semantic recognition models do not have the corresponding industry-specific recognition capabilities, resulting in insufficient generalization ability and low recognition accuracy of the existing semantic recognition models. As a result, it is impossible to accurately recognize the user's consultation questions, leading to incorrect or delayed push of feedback content. Aiming at the problem of low recognition accuracy of the existing semantic recognition solutions for the chemical industry, the present invention proposes a chemical semantic recognition method and a chemical natural language data preprocessing method based on regular expressions, and completes the preprocessing of chemical professional natural language through specific regular expressions. A natural language semantic analysis model based on the physical properties of chemical substances is constructed. By applying the physical property data of chemical substances, it assists in the semantic analysis of chemical natural language, so as to develop a chemical professional Chinese intelligent question-answering system. It solves the problems of obscurity and low accuracy in natural language processing in China's chemical industry, and provides the machine recognition accuracy and efficiency of chemical texts through the application of regular expressions, hidden Markov models based on physical properties, and conditional random field models.

[0028] Figure 1 It is the flowchart of the chemical semantic recognition method provided by an embodiment of the present invention. As Figure 1 shown, an embodiment of the present invention provides a chemical semantic recognition method, and the method includes:

[0029] Step S10: Collect the text information to be recognized, and preprocess the text information based on regular expressions.

[0030] Specifically, the preprocessing of the text information based on regular expressions includes: extracting the list information of hazardous chemicals based on the open-source chemical knowledge graph; performing regularization processing on the typical chemical materials in the text information based on the list information of hazardous chemicals.

[0031] Example 1:

[0032] Regular expressions are just patterns we use to retrieve letters and numbers in text. We perform preliminary retrieval and semantic replacement of chemicals in the way of applying regular expressions. Metacharacters are the basic elements of regular expressions. Here, metacharacters do not have their usual meanings but are interpreted with a certain special meaning. Some metacharacters have special meanings when written inside square brackets. The metacharacters are as follows:

[0033] Metacharacter description: Matches any character except the newline character. [] Character class: Matches any character contained in the square brackets. [^] Negated character class: Matches any character not contained in the square brackets. After completing the first step of character parsing, perform parsing of hazardous chemical corpus. Based on 118 chemical elements such as common hydrogen H, oxygen O, argon Ar, fluorine F, nitrogen N, helium He, neon Ne, phosphorus P, chlorine Cl, krypton Kr, boron B, sulfur S, sodium Na, arsenic As, carbon C, radon Rn, xenon Xe, astatine At... etc., perform regular expression parsing. For example: Sodium hydride is parsed as [^hydrogen*]. At the same time, parse the chemical formulas of related elements.

[0034] Based on the hazardous chemical MSDS database and the national list of dangerous goods database, perform preliminary preprocessing of data and processing of regular expressions for typical chemical materials, such as: [peroxide*], etc.

[0035] Step S20: Based on the pre-constructed chemical word segmentation library, perform word segmentation processing on the preprocessed text information.

[0036] Specifically, perform word segmentation processing on the preprocessed text information based on the initial semantic recognition model; based on the pre-constructed chemical word segmentation library, perform dictionary matching on the text information after word segmentation processing and perform word segmentation adjustment.

[0037] In the embodiment of the present invention, the present invention scheme proposes a Chinese word segmentation library. Chinese automatic word segmentation refers to automatically splitting Chinese text into words by a computer, that is, making there be spaces between words in a Chinese sentence like in English to mark. Chinese automatic word segmentation is considered to be a most basic link in Chinese natural language processing. Then perform text extraction, that is, a technology for automatically extracting specific messages from a large amount of text data for accessing the database. Simply put, it can be understood as extracting important information from the given text, such as chemicals, chemical equations, chemical properties, chemical engineering related equipment, process information, time, place, person, event, cause, result, number, date, currency, proper nouns, etc. Generally speaking, it is necessary to understand key technologies such as who, what substance, what chemical properties, what process, through what reaction equipment, etc.

[0038] Example 2:

[0039] After constructing a feature dataset, the feature matrix may become overly large, leading to a series of problems such as high computational effort and long training times. Therefore, reducing the dimensionality of the feature matrix is essential. Common feature dimensionality reduction methods in machine learning include models with an L1 penalty, principal component analysis (PCA), and linear discriminant analysis (LDA). PCA and LDA share many similarities, both mapping the original samples into a lower-dimensional sample space. PCA is an unsupervised dimensionality reduction method, while LDA is a supervised dimensionality reduction method. In natural language processing, topic models are commonly used. Topic models combine dimensionality reduction with semantic representation, such as statistical topic models like LSI, LDA, PLSA, and HDP. These models seek to represent text in a low-dimensional space (on different topics), reducing dimensionality while preserving the original text's semantic information as much as possible. Topic models are particularly effective for classifying medium-length text.

[0040] Word segmentation methods based on string matching. The basic idea is to segment the Chinese text to be segmented, especially chemicals and dangerous goods, according to certain rules based on dictionary matching. The text is then matched against words in the dictionary. If a match is successful, the word is segmented according to the dictionary. If a match fails, the word is adjusted or reselected, and this cycle repeats. Representative methods include forward maximum matching, reverse maximum matching, and bidirectional matching. Principal component analysis is applied to each segmentation step to continuously identify Chinese words not in the corpus, and relevant words are added in a timely manner to automatically update the chemical engineering word segmentation database.

[0041] Step S30: using the text information after word segmentation as input, sequentially executing multiple chemical property auxiliary analysis model trainings to obtain chemical property semantic recognition results.

[0042] Specifically, the chemical property auxiliary analysis model includes: a chemical property semantic analysis auxiliary model, a physical property auxiliary analysis model based on hidden Markov, and a physical property auxiliary analysis model based on conditional random field. Figure 2 , including the following steps:

[0043] Step S301: Execute chemical property semantic analysis auxiliary model training.

[0044] Specifically, the physical properties of chemicals are classified and graded based on the degree of hazard; chemicals are matched based on regular expressions, and physical properties are matched based on the matched chemicals; and physical property semantic-assisted analysis is performed based on the matched physical properties.

[0045] In the embodiments of the present invention, based on the physical properties of chemicals, the physical properties of chemicals are classified into health hazards, flammability, reactivity, and special hazards. They are divided into five levels of 0, 1, 2, 3, and 4 according to the degree of hazard. A chemical classification database is established according to the above classification, and the physical properties of the chemicals matched by the regular expression are matched. Semantic analysis assistance is performed through the matched physical property data.

[0046] Step S302: Perform the training of the physical property assisted analysis model based on Hidden Markov.

[0047] Specifically, take the output of the chemical physical property semantic analysis assistance model as the input of the physical property assisted analysis model based on Hidden Markov, and calculate the counting probabilities of the text information after word segmentation under each physical property classification respectively; calculate the physical property probability value of the sentence based on the counting probabilities of the text information after word segmentation under each physical property classification. The calculation rule is:

[0048] P w = max(a 1i , a 2i , …, a ni )

[0049] Among them, P w is the physical property probability value; a ni is the counting probability under the nth physical property classification.

[0050] Example 3:

[0051] According to the chemical corpus to be analyzed, assume that only a single sequence in the training data is trained (referred to as 0) in part. In a real speech recognition system, hundreds or thousands of sequences need to be trained. Since each state can only generate one observation symbol, the sum of the probabilities of the observations b should be 1.0. The probability a of a specific transition between states i and j can be calculated by the number of transitions. We denote it as C(i - j), and then it can be normalized by all the number of transitions starting from the state:

[0052]

[0053] Through the above formula, repeatedly estimate the obtained counts. Start with an estimated value of the transition probability and the observation probability, and then repeatedly use these estimated probabilities to derive better and better probabilities. For an observation, calculate its forward probability to obtain the estimated probability. Then, distribute this probability value over all different paths that contribute to the forward probability, incorporate the physical properties of the chemicals (health hazard h, flammability f, reactivity c, special hazard s) for semantic analysis and estimation. Then, according to the counting probabilities, calculate the probability values of the health hazard hi, flammability fi, reactivity ci, and special hazard si respectively. Then calculate the physical property probability value Pw of the sentence.

[0054] Pw = max(hi, fi, ci, si)

[0055] Correct the semantics of the language sequence through Pw. For example, the statement "Inhaling hydrogen sulfide can cause people to die instantly" is semantically analyzed and corrected based on the health hazard hi of hydrogen sulfide.

[0056] Step S303: Perform training on the physical property assisted analysis model based on the conditional random field.

[0057] Specifically, take the output of the physical property assisted analysis model based on the hidden Markov model as the input of the physical property assisted analysis model based on the conditional random field to generate an undirected graph model, expressed as:

[0058] G = (V, E)

[0059] Among them, G is the generated undirected graph; V is the set of chemicals; E is the set of relationships between each chemical and other word segments; output the semantic word segmentation result based on the undirected graph model as the chemical property semantic recognition result.

[0060] Example 4:

[0061] Apply the conditional random field to Chinese word segmentation, chemical extraction, and physical property analysis. The conditional random field is often used in natural language processing tasks such as sequence labeling and data segmentation, directly modeling the joint distribution, such as the mixture Gaussian model, the hidden Markov model, the Markov random field, etc.

[0062] Let G = (V, E) be an undirected graph, V is the set of nodes (usually each chemical is used as a node, or a substance device related to the chemical can also be used), and E is the set of undirected edges (mainly describing the relationship between the chemical and other word segments, and this relationship is semantically corrected using the physical properties (health hazard h, flammability f, reactivity c, special hazard s) of the chemical). Y = {Yv ∈ V}, that is, each node in V corresponds to a random variable Yv, and its value range is the set of possible labels {Y}. If the observation sequence X is the condition, then each random variable satisfies the following Markov property:

[0063] P(Yv|X, Yw, w ≠ v) = P(Yv|X, Yw, w ∼ v)

[0064] Among them, w - v means that the two nodes are adjacent nodes in the graph G, and (X, Y) is a conditional random variable.

[0065] Step S40: Generate feedback content based on the chemical property semantic recognition result.

[0066] Specifically, based on the chemical property semantic recognition result, determine the word segmentation result; obtain each analysis semantic and the associated semantic between each word segment and other word segments based on the word segmentation result as semantic elements; generate complete text information based on the semantic elements as the feedback content.

[0067] In a possible application scenario, a chemical natural language question-answering solution is constructed based on the chemical semantic recognition solution proposed in the present invention, including:

[0068] 1) Natural language preprocessing: After the system reads in a Chinese sentence, it initially enters the steps of natural language processing. The main preprocessing steps include: Regular expression matching: Apply a chemical corpus preprocessing method based on regular expressions to carry out regular expression data matching, implement regular expression processing of chemicals and chemical-related sentences, and form a preliminary chemical semantic analysis. Word segmentation based on string matching: Perform string matching according to the corpus in the chemical database and the chemical corpus to complete the preliminary sentence preprocessing. Chinese word segmentation: For the sentence after regular expression processing, perform Chinese word segmentation, and at the same time, for the corpus in the chemical corpus, perform Chinese word segmentation correction.

[0069] 2) Semantic analysis: Apply methods such as chemical property semantic analysis assistance, property assistance analysis based on hidden Markov, and property assistance analysis based on conditional random fields to perform semantic analysis, and correct the semantic analysis through chemical properties (health hazard h, flammability f, reactivity c, special hazard s).

[0070] 3) Content generation: Content generation is to use a computer to automatically analyze the semantics in the original literature to form a coherent short passage that comprehensively and accurately reflects the central content. And generate a linear sequence of explanatory sentences according to the semantics, and correct the linear sequence of chemical words and sentences to form the corresponding answer content.

[0071] Furthermore, in a possible application scenario, a chemical natural language question-answering system is constructed based on the chemical semantic recognition solution proposed in the present invention, including:

[0072] 1) Speech recognition device: The speech recognition device mainly completes speech-to-text recognition. Its goal is to automatically convert the speech content of humans into corresponding text by the computer. Different from speaker recognition and speaker verification, the latter attempts to identify or verify the speaker who emits the speech rather than the lexical content contained therein.

[0073] 2) Chemical knowledge question-answering system: Apply the method of chemical natural language processing to the content of the speech recognition device, perform chemical natural language processing on the recognized content, form a question from the processed result, and generate question-and-answer content according to the question.

[0074] 3) Speech synthesis / text reading: According to the answer content generated by the question-answering system, it is converted into speech through the speech conversion module of the natural language question-answering system and text reading is performed.

[0075] Figure 3 It is the system structure diagram of the system for mining influencing factors of the safe production status of a refining and chemical device provided by an embodiment of the present invention. As Figure 3 shown, an embodiment of the present invention provides a system for mining influencing factors of the safe production status of a refining and chemical device, and the system includes:

[0076] An acquisition unit, configured to acquire the text information to be recognized and preprocess the text information based on a regular expression.

[0077] Specifically, the preprocessing of the text information based on the regular expression includes: extracting the list information of hazardous chemicals based on an open-source chemical knowledge graph; performing regularization processing on the typical chemical materials in the text information based on the list information of hazardous chemicals.

[0078] A regular expression is just a pattern we use to retrieve letters and numbers in text. The way of applying regular expressions is used for the preliminary retrieval and semantic replacement of chemicals. Metacharacters are the basic elements of regular expressions. Metacharacters here do not mean what they usually express, but are interpreted with a special meaning. Some metacharacters have special meanings when written inside square brackets. The metacharacters are as follows:

[0079] Metacharacter description, matches any character except the newline character. [] Character class, matches any character contained in the square brackets. [^] Negative character class. Matches any character not contained in the square brackets. After completing the first step of character parsing, perform parsing of the hazardous chemical corpus. Based on 118 chemical elements such as common hydrogen H, oxygen O, argon Ar, fluorine F, nitrogen N, helium He, neon Ne, phosphorus P, chlorine CI, krypton Kr, boron B, sulfur S, sodium Na, arsenic As, carbon C, radon Rn, xenon Xe, astatine At... etc., perform regular expression parsing. For example: sodium hydride is parsed as [^hydrogen*]. At the same time, parse the chemical formulas of related elements.

[0080] Based on the hazardous chemical MSDS database and the national hazardous goods list database, perform preliminary preprocessing of the data and regular expression processing of typical chemical materials, such as: [peroxide*], etc.

[0081] A word segmentation unit, configured to perform word segmentation processing on the preprocessed text information based on a pre-constructed chemical industry word segmentation library.

[0082] Specifically, perform word segmentation processing on the preprocessed text information based on an initial semantic recognition model; perform dictionary matching on the text information after word segmentation processing based on the pre-constructed chemical industry word segmentation library and perform word segmentation adjustment.

[0083] In an embodiment of the present invention, the present invention proposes a Chinese word segmentation library. Automatic Chinese word segmentation refers to the use of a computer to automatically segment Chinese text into words, that is, to mark the words in a Chinese sentence with spaces like in English. Automatic Chinese word segmentation is considered to be one of the most basic links in Chinese natural language processing. Then comes text extraction, which is the technology of automatically extracting specific messages from a large amount of text data for accessing a database. It can be simply understood as extracting important information from a given text, such as chemicals, chemical equations, chemical properties, chemical-related equipment, process information, time, place, people, events, causes, results, numbers, dates, currencies, proper nouns, etc. In layman's terms, it is to understand key technologies such as who is using what substance, what chemical properties, what process, and what reaction equipment.

[0084] After constructing a feature dataset, the feature matrix may become overly large, leading to a series of problems such as high computational effort and long training times. Therefore, reducing the dimensionality of the feature matrix is essential. Common feature dimensionality reduction methods in machine learning include models with an L1 penalty, principal component analysis (PCA), and linear discriminant analysis (LDA). PCA and LDA share many similarities, both mapping the original samples into a lower-dimensional sample space. PCA is an unsupervised dimensionality reduction method, while LDA is a supervised dimensionality reduction method. In natural language processing, topic models are commonly used. Topic models combine dimensionality reduction with semantic representation, such as statistical topic models like LSI, LDA, PLSA, and HDP. These models seek to represent text in a low-dimensional space (on different topics), reducing dimensionality while preserving the original text's semantic information as much as possible. Topic models are particularly effective for classifying medium-length text.

[0085] Word segmentation methods based on string matching. The basic idea is to segment the Chinese text to be segmented, especially chemicals and dangerous goods, according to certain rules based on dictionary matching. The text is then matched against words in the dictionary. If a match is successful, the word is segmented according to the dictionary. If a match fails, the word is adjusted or reselected, and this cycle repeats. Representative methods include forward maximum matching, reverse maximum matching, and bidirectional matching. Principal component analysis is applied to each segmentation step to continuously identify Chinese words not in the corpus, and relevant words are added in a timely manner to automatically update the chemical engineering word segmentation database.

[0086] The training unit is used to take the text information after word segmentation processing as input, execute multiple chemical property auxiliary analysis model training in sequence, and obtain the chemical property semantic recognition results.

[0087] Specifically, the chemical property auxiliary analysis model includes: a chemical property semantic analysis auxiliary model, a property auxiliary analysis model based on Hidden Markov Model, and a property auxiliary analysis model based on Conditional Random Field. As Figure 2 , including the following steps:

[0088] Step S301: Execute the training of the chemical property semantic analysis auxiliary model.

[0089] Specifically, classify the physical properties of chemicals and grade them based on the degree of hazard; match chemicals based on regular expressions and match physical properties based on the matched chemicals; perform physical property semantic auxiliary analysis based on the matched physical properties.

[0090] In the embodiment of the present invention, based on the physical properties of chemicals, the physical properties of chemicals are divided into health hazards, flammability, reactivity, and special hazards. They are divided into five grades of 0, 1, 2, 3, and 4 according to the degree of hazard. Establish a chemical classification database according to the above classification, perform physical property matching for the chemicals matched according to the regular expressions, and perform semantic analysis assistance through the matched physical property data.

[0091] Step S302: Execute the training of the property auxiliary analysis model based on Hidden Markov Model.

[0092] Specifically, take the output of the chemical property semantic analysis auxiliary model as the input of the property auxiliary analysis model based on Hidden Markov Model, and calculate the counting probability of the text information after word segmentation under each property classification respectively; calculate the property probability value of the sentence based on the counting probability of the text information after word segmentation under each property classification, and the calculation rule is:

[0093] P w =max(a 1i ,a 2i ,…,a ni )

[0094] Wherein, P w is the property probability value; a ni is the counting probability under the nth property classification.

[0095] According to the chemical corpus to be analyzed, assume that only a single sequence in the training data is trained (referred to as 0) in part. In a real speech recognition system, hundreds or thousands of sequences need to be trained. Since each state can only generate one observation symbol, the sum of the probabilities of the observations b should be 1.0. The probability a of a specific transition between state i and state j can be calculated by the number of transitions, which we denote as C(i-j), and then it can be normalized by all the transition numbers starting from the state:

[0096]

[0097] Using the above formula, the obtained count is repeatedly estimated. Starting from an estimated value of the transition probability and the observation probability, these estimated probabilities are then repeatedly used to derive better and better probabilities. For an observation, calculate its forward probability to obtain the estimated probability. Then, distribute this probability amount over all the different paths that contribute to the forward probability, incorporating the physical properties of the chemical (health hazard h, flammability f, reactivity c, special hazard s) for semantic analysis and estimation. Then, based on the counting probability, calculate the probability values of the health hazard hi, flammability fi, reactivity ci, and special hazard si respectively. Then, calculate the physical property probability value Pw of the statement.

[0098] Pw = max(hi, fi, ci, si)

[0099] Correct the semantics of the language sequence through Pw. For example, the statement "Inhalation of hydrogen sulfide can cause instant death in humans" is corrected through semantic analysis of the health hazard hi of hydrogen sulfide.

[0100] Step S303: Perform training on the physical property assisted analysis model based on the conditional random field.

[0101] Specifically, take the output of the physical property assisted analysis model based on the hidden Markov model as the input of the physical property assisted analysis model based on the conditional random field, and generate an undirected graph model, denoted as:

[0102] G = (V, E)

[0103] where G is the generated undirected graph; V is the set of chemicals; E is the set of relationships between each chemical and other word segments; output the semantic word segmentation result based on the undirected graph model as the chemical physical property semantic recognition result.

[0104] Apply the conditional random field to Chinese word segmentation, chemical extraction, and physical property analysis. The conditional random field is often used in natural language processing tasks such as sequence labeling and data segmentation, directly modeling the joint distribution, such as the mixture Gaussian model, hidden Markov model, Markov random field, etc.

[0105] Let G = (V, E) be an undirected graph, where V is the set of nodes (usually each chemical serves as a node or a substance device related to the chemical can also be used), and E is the set of undirected edges (mainly describing the relationship between the chemical and other word segments, and this relationship is semantically corrected using the physical properties of the chemical (health hazard h, flammability f, reactivity c, special hazard s)). Y = {Yv ∈ V}, that is, each node in V corresponds to a random variable Yv, and its value range is the set of possible labels {Y}. If the observation sequence X is the condition, then each random variable satisfies the following Markov property:

[0106] P(Yv|X,Yw,w≠v) = P(Yv|X,Yw,w~v)

[0107] Where w-v means that two nodes are adjacent nodes in graph G, and (X,Y) is a conditional random variable.

[0108] A feedback unit for generating feedback content based on the chemical property semantic recognition result.

[0109] Specifically, based on the chemical property semantic recognition result, determine the word segmentation result; obtain each analysis semantics and the associated semantics between each word segmentation and other word segmentations based on the word segmentation result as semantic elements; generate complete text information based on the semantic elements as feedback content.

[0110] In a possible application scenario, a chemical natural language question-answering solution is constructed based on the chemical semantic recognition solution proposed by the present invention, including:

[0111] 1) Natural language preprocessing: After the system reads in a Chinese sentence, it initially enters the steps of natural language processing. The main preprocessing steps include: Regular expression matching: Apply a chemical corpus preprocessing method based on regular expressions to carry out regular expression data matching, implement regular expression processing of chemical substances and sentences related to chemical substances, and form a preliminary chemical semantic analysis. Word segmentation based on string matching: Perform string matching according to the corpus in the chemical substance database and the chemical corpus to complete the preliminary sentence preprocessing. Chinese word segmentation: For the sentence after regular expression processing, perform Chinese word segmentation, and at the same time, for the corpus in the chemical corpus, perform Chinese word segmentation correction.

[0112] 2) Semantic analysis: Apply methods such as chemical property semantic analysis assistance, property assistance analysis based on hidden Markov, and property assistance analysis based on conditional random fields to perform semantic analysis, and perform semantic analysis correction through chemical substance properties (health hazard h, flammability f, reactivity c, special hazard s).

[0113] 3) Content generation: Content generation is to use a computer to automatically analyze the semantics in the original literature to form a coherent short passage that comprehensively and accurately reflects the central content. And generate a linear sequence of explanatory sentences according to the semantics, and perform linear sequence correction on chemical words and sentences to form the corresponding answer content.

[0114] Furthermore, in a possible application scenario, a chemical natural language question-answering system is constructed based on the chemical semantic recognition solution proposed by the present invention, including:

[0115] 1) Speech recognition device: The speech recognition device mainly completes speech-to-text recognition, with the goal of enabling a computer to automatically convert human speech content into corresponding text. Different from speaker recognition and speaker verification, the latter attempts to identify or verify the speaker who emits the speech rather than the lexical content contained therein.

[0116] 2) Chemical engineering knowledge Q&A system: Using the method of chemical engineering natural language processing, it processes the content of the speech recognition device and performs chemical engineering natural language processing on the recognized content, forms questions from the processed results, and generates Q&A content based on the questions.

[0117] 3) Speech synthesis / text reading: According to the answer content generated by the Q&A system, it is converted into speech through the speech conversion module of the natural language Q&A system and text reading is performed.

[0118] The embodiments of the present invention also provide a computer-readable storage medium, on which instructions are stored, and when running on a computer, the computer is made to execute the above-mentioned chemical engineering semantic recognition method.

[0119] Those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a single-chip microcomputer, a chip or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks and other various media that can store program codes.

[0120] The optional embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention. In addition, it should be noted that in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not separately describe various possible combination methods.

[0121] In addition, any combination can be made between various different embodiments of the present invention as long as it does not violate the idea of the embodiments of the present invention, and it should also be regarded as the content disclosed by the embodiments of the present invention.

Claims

1. A chemical semantic recognition method, characterized in that, The method includes: Collecting text information to be recognized and preprocessing the text information based on regular expressions; Performing word segmentation on the preprocessed text information based on a pre-constructed chemical engineering word segmentation library; Using the text information after word segmentation as input parameters, and sequentially performing training of multiple chemical engineering physical property auxiliary analysis models to obtain chemical engineering physical property semantic recognition results; Generating feedback content based on the chemical engineering physical property semantic recognition results.

2. The method according to claim 1, wherein The preprocessing of the text information based on regular expressions includes: Extracting information on the list of hazardous chemicals based on an open-source chemical engineering knowledge graph; Performing regularization processing on typical chemical materials in the text information based on the information on the list of hazardous chemicals.

3. The method according to claim 1, wherein The performing word segmentation on the preprocessed text information based on a pre-constructed chemical engineering word segmentation library includes: Performing word segmentation on the preprocessed text information based on an initial semantic recognition model; Performing dictionary matching on the text information after word segmentation based on the pre-constructed chemical engineering word segmentation library and adjusting the word segmentation.

4. The method according to claim 3, characterized in that, The matching rule for the dictionary matching of the text information after word segmentation is based on the forward maximum matching method, the reverse maximum matching method, or the bidirectional matching method.

5. The method according to claim 1, wherein The chemical engineering physical property auxiliary analysis model includes: A chemical engineering physical property semantic analysis auxiliary model, a physical property auxiliary analysis model based on hidden Markov, and a physical property auxiliary analysis model based on conditional random field.

6. The method according to claim 5, wherein Performing training of the chemical engineering physical property semantic analysis auxiliary model includes: Classifying chemical engineering physical properties and grading them based on the degree of hazard; Performing chemical matching based on regular expressions and performing physical property matching based on the matched chemicals; Performing physical property semantic auxiliary analysis based on the matched physical properties.

7. The method according to claim 6, characterized in that, Performing training of the physical property auxiliary analysis model based on hidden Markov includes: Using the output of the chemical engineering physical property semantic analysis auxiliary model as the input of the physical property auxiliary analysis model based on hidden Markov, and respectively calculating the counting probabilities of the text information after word segmentation under each physical property classification; Calculating the physical property probability value of the statement based on the counting probabilities of the text information after word segmentation under each physical property classification, and the calculation rule is: P w = max(a 1i , a 2i , …, a ni ) Among them, P w is the physical property probability value; a ni is the counting probability under the n-th physical property classification.

8. The method according to claim 7, wherein Performing training of the physical property auxiliary analysis model based on conditional random field includes: Using the output of the physical property auxiliary analysis model based on hidden Markov as the input of the physical property auxiliary analysis model based on conditional random field to generate an undirected graph model, which is expressed as: G=(V,E) where G is the generated undirected graph; V is the set of chemicals; E is the set of relationships between each chemical and other word segments; Outputting the semantic word segmentation result based on the undirected graph model as the chemical engineering physical property semantic recognition result.

9. The method according to claim 1, wherein The generating feedback content based on the chemical engineering physical property semantic recognition results includes: Determining the word segmentation result based on the chemical engineering physical property semantic recognition results; Obtaining each analysis semantics and the associated semantics between each word segment and other word segments based on the word segmentation result as semantic elements; Generating complete text information based on the semantic elements as feedback content.

10. A chemical semantic recognition system, characterized in that, The system includes: A collection unit for collecting text information to be recognized and preprocessing the text information based on regular expressions; A word segmentation unit for performing word segmentation on the preprocessed text information based on a pre-constructed chemical engineering word segmentation library; A training unit, configured to use the text information after word segmentation as an input parameter, and sequentially execute the training of multiple chemical property auxiliary analysis models to obtain a chemical property semantic recognition result; A feedback unit, configured to generate feedback content based on the chemical property semantic recognition result.

11. A computer-readable storage medium, characterized in that, Instructions are stored on the computer-readable storage medium, and when running on a computer, the computer is caused to execute the chemical semantic recognition method described in any one of claims 1-9.

12. An electronic device, the electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the chemical semantic recognition method described in any one of claims 1-9 is implemented.