Construction method and device of theme keyword dictionary, equipment and storage medium

By constructing a topic-keyword dictionary and using LDA clustering technology, the efficiency and accuracy of engineering project type recognition in the pipeline system are solved, and the conversion from unstructured sentences to structured information is realized, unknown patterns are identified, and project quality is improved.

CN120372018APending Publication Date: 2025-07-25PIPECHINA SOUTH CHINA CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510460864.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art cannot quickly and accurately identify the type of engineering project in the pipeline system, resulting in low efficiency and poor accuracy. Especially when known patterns and unknown patterns are confused, traditional LDA methods cannot recognize sentence-level failure patterns.

Method used

By constructing a topic-keyword dictionary, using the potential Dirichlet allocation LDA clustering technology, clustering the training corpus text, updating sentences from known topics, generating the updated topic-keyword dictionary, and identifying sentence-level failure modes.

Benefits of technology

It improves the accuracy and efficiency of project type identification, and can extract structured information from unstructured sentences, identify unknown patterns, and improve project quality and work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372018A_ABST
    Figure CN120372018A_ABST
Patent Text Reader

Abstract

The invention discloses a theme keyword dictionary construction method and device, equipment and a storage medium. The method comprises the following steps: acquiring a training corpus text; based on a to-be-trained topic-keyword dictionary, performing latent Dirichlet allocation LDA clustering on the training corpus text to obtain a clustering result, the clustering result at least comprising a topic corresponding to each sentence in the training corpus text; and updating a to-be-trained topic-keyword dictionary on the basis of sentences with topics being known topics in the training corpus text to obtain an updated topic-keyword dictionary. According to the method, the topic-keyword dictionary is constructed, so that information in corpora can be mined through the topic-keyword dictionary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of data mining, and in particular, to a method, apparatus, device, and storage medium for constructing a thematic keyword dictionary. Background Art

[0002] In the management of engineering projects involving natural language processing, it is usually necessary to manually learn the knowledge in the business field and evaluate the engineering projects to determine whether the engineering projects are projects with known patterns or projects with unknown patterns. For example, traditional methods for identifying pipeline failure modes mainly rely on manual experience and historical data analysis. However, with the increasing complexity of pipeline systems, pipeline failure problems are becoming increasingly serious. Processing information in pipeline systems manually has problems of low efficiency and poor accuracy.

[0003] Therefore, how to quickly mine information in the corresponding field and study the type of engineering projects based on the mined information is an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides a method, apparatus, device, and storage medium for constructing a thematic keyword dictionary. By constructing a theme-keyword dictionary, engineering projects can be studied through the theme-keyword dictionary, so as to solve the problem in the prior art that information in the corresponding field cannot be mined and the type of engineering projects cannot be studied based on the mined information.

[0005] According to one aspect of the present invention, a method for constructing a thematic keyword dictionary is provided. The method includes:

[0006] Obtaining training corpus texts;

[0007] Performing Latent Dirichlet Allocation (LDA) clustering on the training corpus texts based on a thematic keyword dictionary to be trained, and obtaining a clustering result, where the clustering result at least includes the theme corresponding to each sentence in the training corpus texts;

[0008] Updating the thematic keyword dictionary to be trained based on the sentences in the training corpus texts whose theme is a known theme, and obtaining an updated thematic keyword dictionary.

[0009] According to another aspect of the present invention, a device for constructing a thematic keyword dictionary is provided. The device includes:

[0010] An obtaining module, configured to obtain training corpus texts;

[0011] A clustering module, configured to perform Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on the to-be-trained topic-keyword dictionary, so as to obtain a clustering result, where the clustering result at least includes the topic corresponding to each sentence in the training corpus text;

[0012] An updating module, configured to update the to-be-trained topic-keyword dictionary based on the sentences in the training corpus text whose topics are known topics, so as to obtain an updated topic-keyword dictionary.

[0013] According to another aspect of the present invention, there is provided an electronic device, where the electronic device includes: at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the method for constructing a topic-keyword dictionary according to any embodiment of the present invention.

[0016] According to another aspect of the present invention, there is provided a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method for constructing a topic-keyword dictionary according to any embodiment of the present invention is implemented.

[0017] A method, apparatus, device, and storage medium for constructing a topic-keyword dictionary according to an embodiment of the present invention. The method includes: obtaining a training corpus text; performing Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on the to-be-trained topic-keyword dictionary, so as to obtain a clustering result, where the clustering result at least includes the topic corresponding to each sentence in the training corpus text; and updating the to-be-trained topic-keyword dictionary based on the sentences in the training corpus text whose topics are known topics, so as to obtain an updated topic-keyword dictionary. This method can study engineering projects through a topic-keyword dictionary, and solves the problem in the prior art that information in the corresponding field cannot be mined and the type of engineering projects cannot be studied through the mined information.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understandable through the following description. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart of a method for constructing a subject-keyword dictionary provided in Embodiment 1 of the present invention;

[0021] Figure 2 It is a schematic diagram of an original corpus data provided in an embodiment of the present invention;

[0022] Figure 3 It is a schematic diagram of the original corpus data after splitting provided in an embodiment of the present invention;

[0023] Figure 4 It is a schematic diagram of a subject-keyword dictionary provided in an embodiment of the present invention;

[0024] Figure 5 It is a schematic diagram of a weighted subject-keyword dictionary provided in an embodiment of the present invention;

[0025] Figure 6 It is a schematic diagram of the modified initial training corpus provided in an embodiment of the present invention;

[0026] Figure 7 It is a schematic flowchart of constructing a subject-keyword dictionary based on supervised LDA provided in an embodiment of the present invention;

[0027] Figure 8 It is a schematic diagram of a clustering result provided in an embodiment of the present invention;

[0028] Figure 9 It is a schematic diagram of an updated subject-keyword dictionary provided in an embodiment of the present invention;

[0029] Figure 10 It is a schematic diagram of a subject-keyword dictionary after adding a new subject provided in an embodiment of the present invention;

[0030] Figure 11 It is a schematic structural diagram of a device for constructing a subject-keyword dictionary provided in Embodiment 2 of the present invention;

[0031] Figure 12 It is a schematic structural diagram of an electronic device for the method of constructing a subject-keyword dictionary according to an embodiment of the present invention. Detailed implementation manners

[0032] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. It should be understood that the steps recorded in the method embodiments of the present invention can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0033] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, any variations of the terms "including" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0035] It should be noted that the modification of "one" and "a plurality" mentioned in the present invention is illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and do not limit the scope of these messages or information.

[0037] As the complexity of the pipe network system continues to increase, the problem of pipe network failure becomes increasingly serious, bringing many challenges to the normal operation of pipelines in the network. Traditional methods for identifying pipe network failure modes mainly rely on manual experience and historical data analysis, suffering from low efficiency and poor accuracy. In recent years, data-driven methods have gradually gained attention. In particular, the development of machine learning and deep learning technologies has provided new ideas for identifying pipe network failure modes. In engineering projects involving natural language processing, the specific quantity of structured knowledge contained in the overall corpus is generally unknown, which can affect the project progress and engineering quality.

[0038] Latent Dirichlet Allocation (LDA) is an unsupervised topic model that can extract latent topics from a large amount of text. However, the topics of LDA are generally sets of words, which is different from the processing method in natural language mining that mainly focuses on sentences and uses multi-tuples composed of functional elements or framework elements. Traditional LDA methods cannot be used to identify sentence-level failure modes in documents. Especially when known and unknown modes are mixed together, traditional LDA has no recognition ability.

[0039] Based on this, in this embodiment, supervised LDA is used to cluster sentences, and a constructed topic-keyword dictionary can exclude known sentence patterns, thereby discovering the quantity of unknown patterns. This dictionary can understand the overall structure of the corpus and estimate the workload of natural language processing, improving engineering quality, as follows:

[0040] Embodiment 1

[0041] Figure 1 FIG. is a schematic flowchart of a method for constructing a topic-keyword dictionary provided in Embodiment 1 of the present invention. This method is applicable to constructing a topic-keyword dictionary in the corresponding field to study the situation of engineering projects through the topic-keyword dictionary. This method can be executed by a device for constructing a topic-keyword dictionary, where the device can be implemented by software and / or hardware and is generally integrated on an electronic device. In this embodiment, the electronic device includes but is not limited to: devices such as computers.

[0042] As Figure 1 shown, a method for constructing a topic-keyword dictionary provided in Embodiment 1 of the present invention includes the following steps:

[0043] S110. Obtain training corpus text.

[0044] Among them, the training corpus text can be the preprocessed corpus text. The corpus text can be documents such as literature and reports. The field to which the corpus text belongs can be determined based on the field where the theme-keyword dictionary needs to be applied. For example, if the theme-keyword dictionary needs to be applied to the pipeline network risk field, the corpus text can be text of the pipeline network type. If the theme-keyword dictionary needs to be applied to the oil and petrochemical exploration and development field, the corpus text can be text of the oil and petrochemical exploration and development type.

[0045] In this embodiment, the training corpus text can be obtained.

[0046] In one embodiment, the obtaining of the training corpus text includes: obtaining the original corpus data, performing splitting, word segmentation, and stop word removal on the original corpus data to obtain the initial training corpus; determining the known themes, and constructing the theme-keyword dictionary to be trained based on the known themes; transforming the sentences in the initial training corpus that have the same themes as those in the theme-keyword dictionary to be trained to obtain the transformed sentences; and updating the initial training corpus based on the transformed sentences to obtain the training corpus text.

[0047] Among them, the original corpus data can be the unprocessed corpus text. The theme-keyword dictionary can include the themes and the semantic elements corresponding to the themes.

[0048] In this embodiment, after obtaining the original corpus data, the original corpus data can be processed by means such as splitting, word segmentation, and stop word removal to obtain the initial training corpus. The known themes are determined through the initial training corpus, and the theme-keyword dictionary to be trained is constructed based on the known themes. The sentences in the initial training corpus that have the same themes as those in the theme-keyword dictionary to be trained are transformed to obtain the transformed sentences, and the initial training corpus is updated based on the transformed sentences to obtain the training corpus text.

[0049] Exemplarily, Figure 2 is a schematic diagram of an original corpus data provided by an embodiment of the present invention. Figure 3 is a schematic diagram of the original corpus data after splitting provided by an embodiment of the present invention. As Figure 2 shown, in the field of pipeline network risk analysis, this document is a literature with a complete structure. In addition to the causal relationship, there are many other unknown relationships. As Figure 3 shown, after splitting the original corpus data to the sentence level, it can be seen that there are many sentences that are meaningless for risk business analysis, and these need to be removed after the literature is split.

[0050] When constructing the topic dictionary, by determining the known topics, the business knowledge is recorded in the dictionary as a general knowledge that can be used for different tasks. When determining the known topics, only the confirmed sentences are screened as seed annotations. Figure 4 The figure is a schematic diagram of a topic-keyword dictionary provided by an embodiment of the present invention. As Figure 4 shown, the topic-keyword dictionary can also annotate sentences and the documents described by the sentences. The name of a known topic is defined as a risk relationship, and the most typical causal relationship in risk analysis is defined. The semantic elements of the causal relationship include 3 main components: cause, lead to, and result. Sentences in engineering are generally relatively long and generally describe one cause and multiple results. For example, Figure 4 in " Figure 4 the long operation years" leads to 3 results. The 4th sentence in

[0051] describes a causal chain: driving operation --> sudden increase in pipeline pressure --> local bulging --> pipeline burst. These complex relationships will all be expressed as a list of multiple causal triples. The cause and result are not single words. Their basic structure is a triple of (object, parameter, parameter state). For example, "sudden increase in pipeline pressure" is a triple of (pipeline, pressure, sudden increase). Some triples may not be complete. For example, "geographical location" has only 1 word and is not a triple, but combined with its previous text "special geographical location of the pipeline", it can form a triple of (pipeline, geographical location, special), that is, a single word is often a shortened form of a triple.

[0052] In this embodiment, the training initial corpus can be transformed based on the topic-keyword dictionary. Specifically, the sentences in the training initial corpus that are the same as the topics in the to-be-trained topic-keyword dictionary can be found. For this sentence, based on the semantic elements corresponding to the topic in the to-be-trained topic-keyword dictionary, new sentences are generated. The new sentences and the old sentences are used as the transformed sentences corresponding to the sentences with the same topic.

[0053] Exemplarily, for a sentence whose theme is known, the elements in the dictionary are combined into a new sentence, and then added to the existing sentence. For other sentences, they are kept in their original state. The sentences with known themes and the unknown sentences that follow together constitute the training corpus for LDA. Figure 5 A schematic diagram of a weighted theme-keyword dictionary provided by an embodiment of the present invention is shown as Figure 5 shown. New sentence groups can be generated through semantic elements corresponding to the theme. The sentence formed by the theme-keyword is a phrase, and the text after weighting the dictionary formed by combining it with the original mother sentence is a paragraph. Each sentence in this paragraph has a different form, but their themes are the same, all being the typical risk theme of causal relationship. Figure 6 A schematic diagram of a modified initial training corpus provided by an embodiment of the present invention is shown as Figure 6 shown. The 5th line therein is the sentence modified through the theme corresponding to the causal relationship, and the entire text thus becomes 2 sentences, while the other sentences remain unchanged.

[0054] S120. Perform Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on the to-be-trained theme-keyword dictionary to obtain a clustering result, where the clustering result at least includes the theme corresponding to each sentence in the training corpus text.

[0055] Among them, the clustering result may include the theme corresponding to each sentence in the training corpus text.

[0056] In this embodiment, based on the to-be-trained theme-keyword dictionary, the training corpus text can be clustered through the LDA clustering method to obtain a clustering result.

[0057] In one embodiment, the performing Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on the to-be-trained theme-keyword dictionary to obtain a clustering result includes: constructing a distribution matrix based on the training corpus text, where the distribution matrix includes a sentence-theme matrix and a theme-word matrix; and performing LDA clustering on the training corpus text through the to-be-trained theme-keyword dictionary based on a preset number of theme categories and the distribution matrix to obtain a clustering result.

[0058] Among them, the distribution matrix can be a matrix including a sentence-theme matrix and a theme-word matrix. The sentence-theme matrix describes the degree of association between each sentence and each theme. The theme-word matrix describes the degree of association between each theme and each word. The preset number of theme categories can be the number of categories of themes to be clustered, and the preset number of theme categories can be set to be larger than the number of known themes.

[0059] In this embodiment, a distribution matrix can be constructed based on the training corpus text. Based on the preset number of topic categories and the distribution matrix, the training corpus text is subjected to LDA clustering through the topic-keyword dictionary to be trained, and a clustering result is obtained.

[0060] Exemplarily, Figure 7 FIG. is a schematic flowchart of constructing a topic-keyword dictionary based on supervised LDA provided by an embodiment of the present invention. As Figure 7 shown, according to the LDA algorithm, first, the text is preprocessed such as word segmentation and stop word removal to construct an ndw matrix, and two implicit distribution matrices, ndz and ndw, are initialized; through Gibbs sampling, it is trained 50 times; the preset number of topic categories is selected to be 3 more than the known number of topics, and the total number is generally selected to be less than 10. Figure 8 FIG. is a schematic diagram of a clustering result provided by an embodiment of the present invention. As Figure 8 shown, the clustering result may further include keywords corresponding to the topics of each cluster. The total number of the trained results is 4 types of topics (z1 - z4). Among them, there is one more sentence in the original dictionary that is consistent with the causal relationship z1. The structure of the studied sentence is "the number of pipe body fracture accidents is increasing continuously, bringing huge losses to the oilfield". The sorted topic-keyword 3-tuple is (the number of pipe body fracture accidents is increasing continuously, bringing, huge losses to the oilfield). Here, the grammatical structure of "bring...to..." in Chinese can be specifically processed. The other z2 - z4 clustering structures roughly describe several categories such as the phrase structure of the title, detection activities, and physical property descriptions. The method of this embodiment can obtain a clustering result that is roughly consistent with the actual situation only by word distribution without semantic support.

[0061] The traditional LDA algorithm calculates the latent topics of documents and the keyword dictionary corresponding to the topics, and cannot control the clustering result. In this embodiment, by constructing known topic-keyword multi-tuples, the training is supervised, and the topics of sentences are clustered through LDA, so that the clustering result can be controlled. LDA iteratively updates the model parameters through three main loops. It is initialized to randomly assign topics to each document. The topic assignment loop updates the word distribution of each topic and the topic distribution of each document according to the current topic assignment situation and the parameter update loop until LDA gradually converges and can identify the latent topic structure in the text.

[0062] S130. Update the topic-keyword dictionary to be trained based on the sentences with known topics in the training corpus text to obtain an updated topic-keyword dictionary.

[0063] In this embodiment, the topic-keyword dictionary to be trained can be updated by the sentences with known topics in the training corpus text to obtain an updated topic-keyword dictionary.

[0064] In one embodiment, updating the to-be-trained theme-keyword dictionary based on sentences with known themes in the training corpus text to obtain an updated theme-keyword dictionary, including: for sentences with known themes in the training corpus text, determining whether the matching degree between the sentence and the keywords of the corresponding known theme reaches a preset threshold; if so, constructing a corresponding keyword multi-tuple for the sentence; adding the keyword multi-tuple corresponding to the sentence to the to-be-trained theme-keyword dictionary, and adding the sentence as a sample to the labeled samples in the to-be-trained theme-keyword dictionary to obtain an updated theme-keyword dictionary.

[0065] Among them, the preset threshold can be defined according to the actual situation. The keyword multi-tuple can be a multi-tuple composed of semantic elements corresponding to the theme. The labeled sample can be an example sentence in the theme-keyword dictionary.

[0066] In this embodiment, for sentences with known themes in the training corpus text, it can be determined whether the matching degree between the sentence and the keywords of the corresponding known theme reaches a preset threshold. For example, determining how many words in the sentence are included in the keywords of the known theme; if the sentence matches the theme, constructing a corresponding keyword multi-tuple for the sentence, adding the keyword multi-tuple corresponding to the sentence to the to-be-trained theme-keyword dictionary, and adding the sentence as a sample to the labeled samples in the to-be-trained theme-keyword dictionary to obtain an updated theme-keyword dictionary.

[0067] Exemplarily, after confirmation, according to the clustering theme with a known theme and its sentences, searching the to-be-trained theme-keyword dictionary to determine new sentences and their corresponding keyword multi-tuples, repeating until no new sentences appear in the known theme set. For example, according to Figure 8 the first 16 keywords in the keyword list of theme z1 clustered by LDA, it is found that words such as "fracture accident, increase, oilfield, loss" in the sentences of category z1 in the transformed corpus are all in the keyword dictionary. Therefore, the sentence "With the increasing energy demand and the implementation of the national strategy to ensure energy security, the intensity of oil and gas exploration has been continuously increased, and the pipe body fracture accidents have been increasing continuously, bringing huge losses to the oilfield" can be added as a sample to the labeled samples. Figure 9 FIG. is a schematic diagram of an updated theme-keyword dictionary provided by an embodiment of the present invention. As Figure 9 shown, a triple (pipe body fracture accident, bringing, huge losses to the oilfield) is added to the last line, thus forming a new corpus and a new theme-keyword dictionary.

[0068] A method for constructing a topic-keyword dictionary provided in Embodiment 1 of the present invention includes: obtaining a training corpus text; performing Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on a topic-keyword dictionary to be trained to obtain a clustering result, where the clustering result at least includes the topic corresponding to each sentence in the training corpus text; and updating the topic-keyword dictionary to be trained based on the sentences in the training corpus text whose topics are known topics to obtain an updated topic-keyword dictionary. This method can study engineering projects through the topic-keyword dictionary, solving the problem in the prior art that information in the corresponding field cannot be mined and the type of engineering projects cannot be studied based on the mined information.

[0069] Based on the above embodiment, a variant embodiment of the above embodiment is proposed. Here, it should be noted that for the sake of brevity of description, only the differences from the above embodiment are described in the variant embodiment.

[0070] In one embodiment, the method further includes: if there are sentences in the training corpus text whose topics are unknown topics, constructing a new topic name and corresponding semantic elements based on the sentences; adding the sentences, the new topic name and corresponding semantic elements to the topic-keyword dictionary to obtain a topic-keyword dictionary after adding the new topic; and continuously updating the topic-keyword dictionary after adding the new topic based on the remaining sentences in the training corpus text whose topics are unknown topics until all sentences are clustered to obtain a new topic-keyword dictionary.

[0071] Among them, the unknown topic can be a topic different from the topics already existing in the constructed topic-keyword dictionary.

[0072] In this embodiment, if there are sentences in the training corpus text whose topics are unknown topics, a new topic name and corresponding semantic elements can be constructed based on the sentences, and the sentences, the new topic name and corresponding semantic elements can be added to the topic-keyword dictionary to obtain a new topic-keyword dictionary. Then, the new topic-keyword dictionary can be continuously updated according to the remaining sentences in the training corpus text whose topics are unknown topics until all sentences are clustered to obtain the final topic-keyword dictionary.

[0073] Exemplarily, the remaining sentences and topics in the training corpus text are taken out to construct a new topic and a new dictionary, and then the LDA is used to train a new model until all sentences can be classified. For example Figure 8As shown, the unknown clusters are z2 - z4. Among them, z2 is a phrase rather than a sentence, so it can be temporarily ignored. z2 is an artificial risk discovery process, which is inconsistent with the risk mainly caused by a certain physical mutation phenomenon resulting in structural collapse, so it can be disregarded. z4 is a structural description that depicts the physical characteristics of failure. This characteristic is a detailed portrait of failure, is related to the risk mode, and is an attribute description of failure. Therefore, another semantic framework for risk can be defined, and the name of the constructed theme is "failure feature description", including the following attributes (object, specification, length, failure location, failure length distribution, failure plane description, failure space description). Figure 10 This is a schematic diagram of the theme-keyword dictionary after adding a new theme provided by the embodiment of the present invention, as Figure 10 shown. The last row in the table is the newly added sentence, the theme to which the sentence belongs, and the corresponding multi-tuple of the theme. It can be seen from the table that the semantic elements of "failure feature description" and "causal relationship" are completely different, but both are related to the entire risk description mode and are the content of risk natural language research. These processed and proofread corpora, dictionaries, etc. all need to be stored in the database and saved as important business knowledge.

[0074] In this embodiment, the text is split into sentences, and the sentences are used as the documents of LDA. For the sentences with known themes, according to the theme dictionary, a sentence composed of a theme-keyword multi-tuple with the same semantics is added to the original sentence, which is equivalent to weighting the distribution of the original sentence and strengthening the keyword weights in the sentence. Then, LDA clustering calculation is performed on the transformed text. According to the calculation results, keyword analysis is carried out on the sentences that are clustered together with the known theme sentences. In the clustered theme-keyword sequence, the words contained in the sentence are taken as the keyword multi-tuple from the front to the back, and then a sentence composed of the theme-keyword multi-tuple is added to the corresponding sentence, and the second LDA calculation is entered; after calculating several times and finding no new sentences with known patterns, the clustering of sentences with unknown themes is checked. First, a theme and its keywords are determined according to the sentences in the cluster, then a sentence composed of the theme-keyword multi-tuple is added to the corresponding sentence, and the theme-keyword multi-tuple is added to the dictionary, and then LDA clustering is performed until all sentences have corresponding themes.

[0075] In one embodiment, the method further includes: obtaining an evaluation corpus text; evaluating the theme-keyword dictionary based on the evaluation corpus text to obtain an evaluation result; if the evaluation result does not meet the evaluation criteria, obtaining a new training corpus text and continuing to update the theme-keyword dictionary until the evaluation result corresponding to the theme-keyword dictionary meets the evaluation criteria.

[0076] Among them, the evaluation result can be the coverage of text recognition, and the coverage refers to the degree of consistent theme coverage. The evaluation criteria can be defined according to the actual situation.

[0077] In this embodiment, the theme-keyword dictionary can be evaluated by the evaluation corpus text. If the evaluation result does not meet the evaluation criteria, new training corpus text is obtained to continue updating the theme-keyword dictionary until the evaluation result corresponding to the theme-keyword dictionary meets the evaluation criteria.

[0078] Exemplarily, as the processed literature corpus increases, the newly added corpus becomes less and less, and the knowledge of the entire business will continue to converge. With the accumulation of the theme-keyword multi-tuple dictionary, the exhaustion of a business domain can be achieved, which can improve the speed, automation degree, and work efficiency of pipeline network risk identification. For example, through the analysis of 997 accident analysis reports, a dictionary with 5 themes and a total of 469 records is obtained, which basically represents the total amount of structured knowledge in the pipeline network risk field.

[0079] At this time, new literature can be read in and operations such as splitting, cleaning, and removing stop words can be performed. By querying the theme-keyword dictionary, the themes and keywords of the sentences in the new literature can be determined. It can be queried through the dictionary, or various algorithm models can be considered, such as Conditional Random Field (CRF), Bidirectional Encoder Representations from Transformers (BERT), and the current Large Language Model (LLM). When evaluating, the evaluation criteria are set. For risk analysis literature, the evaluation criteria are set at 60%, that is, 60 out of 100 sentences can identify the known themes and theme-keyword multi-tuples, then the engineering goal can be achieved. This is because although it is literature describing risk analysis, nearly half of the descriptions are still irrelevant to the risk itself. Therefore, the risk theme coverage is not set too high. After more than 1000 fault analyses, the cumulative number of the theme-keyword dictionary is greater than 5 themes and about 500, which can basically cover 60% of the general risk semantic recognition, and there is no need to increase the training workload. This shows that for a specific small topic, its themes and keywords are limited, which provides good data support for the estimation of engineering workload.

[0080] After traditional LDA clustering, no knowledge is retained. In this embodiment, by introducing a known "topic-element tuple" dictionary and combining the topic extraction ability of the LDA algorithm, the topics of the sentences and elements of new documents are identified, and the conversion from unstructured sentences to structured ones can be completed. If the dictionary needs to be updated, the text can also be used as corpus to modify the sentences and repeat the LDA operation. This can not only discover new combinations of element tuples from known patterns, but also discover new topic patterns from the remaining topic-keywords, thereby predicting the number of corresponding sentence patterns in the entire corpus, estimating the workload of natural language processing, and improving the controllability of the project.

[0081] Suppose after obtaining the topic-keyword dictionary and then needing to study an engineering project. If 10 topics z are classified from the content of the engineering project, and there is only 1 topic in the topic-keyword dictionary, then this project should be positioned as a new project; if there are 8 topics in the topic-keyword dictionary and only 2 topics are unknown, then this project can be defined as a project to fill in the gaps, and 2 topics can be added to the original dictionary, so that professionals do not need to study the entire project.

[0082] Embodiment 2

[0083] Figure 11 FIG. 2 is a schematic structural diagram of a device for constructing a topic-keyword dictionary provided in Embodiment 2 of the present invention. This device is applicable to constructing a topic-keyword dictionary in the corresponding field to study engineering projects through the topic-keyword dictionary. The device can be implemented by software and / or hardware and is generally integrated on an electronic device.

[0084] As Figure 11 shown, the device includes:

[0085] An acquisition module 210, configured to acquire a training corpus text;

[0086] A clustering module 220, configured to perform Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on the topic-keyword dictionary to be trained, and obtain a clustering result, where the clustering result at least includes the topic corresponding to each sentence in the training corpus text;

[0087] An update module 230, configured to update the topic-keyword dictionary to be trained based on the sentences in the training corpus text whose topics are known topics, and obtain an updated topic-keyword dictionary.

[0088] This embodiment provides a device for constructing a thematic keyword dictionary, including: an acquisition module for acquiring a training corpus text; a clustering module for performing Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on a to-be-trained theme-keyword dictionary to obtain a clustering result, where the clustering result at least includes the theme corresponding to each sentence in the training corpus text; and an update module for updating the to-be-trained theme-keyword dictionary based on the sentences in the training corpus text whose themes are known themes to obtain an updated theme-keyword dictionary. Studying engineering projects through the theme-keyword dictionary solves the problem in the prior art that information in the corresponding field cannot be mined and the type of engineering projects cannot be studied based on the mined information.

[0089] Further, the acquisition module 210 includes:

[0090] Acquire the original corpus data, and after splitting, word segmenting, and removing stop words from the original corpus data, obtain the initial training corpus;

[0091] Construct a to-be-trained theme-keyword dictionary based on known themes;

[0092] Transform the sentences in the initial training corpus that have the same theme as the to-be-trained theme-keyword dictionary to obtain transformed sentences;

[0093] Update the initial training corpus based on the transformed sentences to obtain the training corpus text.

[0094] Further, transforming the sentences in the initial training corpus that have the same theme as the to-be-trained theme-keyword dictionary to obtain transformed sentences includes:

[0095] Find out the sentences in the initial training corpus that have the same theme as the to-be-trained theme-keyword dictionary;

[0096] For the sentences with the same theme, generate new sentences based on the semantic elements corresponding to the theme in the to-be-trained theme-keyword dictionary;

[0097] Use the new sentences and the sentences with the same theme as the transformed sentences corresponding to the sentences with the same theme.

[0098] Further, the clustering module 220 includes:

[0099] Construct a distribution matrix based on the training corpus text, where the distribution matrix includes a sentence-theme matrix and a theme-word matrix;

[0100] Based on the preset number of topic categories and the distribution matrix, perform LDA clustering on the training corpus text through the to-be-trained topic-keyword dictionary to obtain a clustering result.

[0101] Further, the update module 230 includes:

[0102] For sentences in the training corpus text with known topics, determine whether the matching degree between the sentence and the keywords of the corresponding known topic reaches a preset threshold;

[0103] If so, construct a corresponding keyword multi-tuple for the sentence;

[0104] Add the keyword multi-tuple corresponding to the sentence to the to-be-trained topic-keyword dictionary, and add the sentence as a sample to the labeled samples in the to-be-trained topic-keyword dictionary to obtain an updated topic-keyword dictionary.

[0105] Further, the device further includes:

[0106] If there are sentences in the training corpus text with unknown topics, construct a new topic name and corresponding semantic elements based on the sentence;

[0107] Add the sentence, the new topic name and the corresponding semantic elements to the topic-keyword dictionary to obtain a topic-keyword dictionary after adding a new topic;

[0108] Continue to update the topic-keyword dictionary after adding a new topic based on the remaining sentences with unknown topics in the training corpus text until all sentences are clustered, and obtain a new topic-keyword dictionary.

[0109] Further, the device further includes:

[0110] Obtain an evaluation corpus text;

[0111] Evaluate the topic-keyword dictionary based on the evaluation corpus text to obtain an evaluation result;

[0112] If the evaluation result does not meet the evaluation criteria, obtain a new training corpus text and continue to update the topic-keyword dictionary until the evaluation result corresponding to the topic-keyword dictionary meets the evaluation criteria.

[0113] The above-mentioned device for constructing a topic-keyword dictionary can execute the method for constructing a topic-keyword dictionary provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0114] Embodiment III

[0115] Figure 12FIG. 0 shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0116] As Figure 12 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0117] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0118] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for constructing a subject keyword dictionary.

[0119] In some embodiments, the method for constructing the subject keyword dictionary can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the processor 11, one or more steps of the method for constructing the subject keyword dictionary described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for constructing the subject keyword dictionary by any other suitable means (e.g., by means of firmware).

[0120] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0122] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0125] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0126] It should be understood that various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0127] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for constructing a dictionary of topic keywords, characterized in that The method includes: Obtaining a training corpus text; Performing Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on a to-be-trained topic-keyword dictionary to obtain a clustering result, where the clustering result at least includes the topic corresponding to each sentence in the training corpus text; Updating the to-be-trained topic-keyword dictionary based on the sentences in the training corpus text whose topics are known topics to obtain an updated topic-keyword dictionary.

2. The method according to claim 1, wherein The obtaining of the training corpus text includes: Obtaining original corpus data, and performing splitting, word segmentation, and stop word removal on the original corpus data to obtain an initial training corpus; Determining known topics, and constructing a to-be-trained topic-keyword dictionary based on the known topics; Transforming the sentences in the initial training corpus that have the same topics as those in the to-be-trained topic-keyword dictionary to obtain transformed sentences; Updating the initial training corpus based on the transformed sentences to obtain a training corpus text.

3. The method according to claim 2, wherein The transforming of the sentences in the initial training corpus that have the same topics as those in the to-be-trained topic-keyword dictionary to obtain transformed sentences includes: Searching for the sentences in the initial training corpus that have the same topics as those in the to-be-trained topic-keyword dictionary; For the sentences with the same topics, generating new sentences based on the semantic elements corresponding to the topics in the to-be-trained topic-keyword dictionary; Taking the new sentences and the sentences with the same topics as the transformed sentences corresponding to the sentences with the same topics.

4. The method according to claim 1, characterized in that, The performing of Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on a to-be-trained topic-keyword dictionary to obtain a clustering result includes: Constructing a distribution matrix based on the training corpus text, where the distribution matrix includes a sentence-topic matrix and a topic-word matrix; Performing LDA clustering on the training corpus text based on a preset number of topic categories and the distribution matrix through the to-be-trained topic-keyword dictionary to obtain a clustering result.

5. The method according to claim 1, wherein The updating of the to-be-trained topic-keyword dictionary based on the sentences in the training corpus text whose topics are known topics to obtain an updated topic-keyword dictionary includes: For the sentences in the training corpus text whose topics are known topics, determining whether the matching degree between the sentence and the keywords of the corresponding known topic reaches a preset threshold; If so, constructing a corresponding keyword multi-tuple for the sentence; Adding the keyword multi-tuple corresponding to the sentence to the to-be-trained topic-keyword dictionary, and adding the sentence as a sample to the labeled samples in the to-be-trained topic-keyword dictionary to obtain an updated topic-keyword dictionary.

6. The method according to claim 1, wherein The method further includes: If there are sentences in the training corpus text whose topics are unknown topics, constructing a new topic name and corresponding semantic elements based on the sentences; Adding the sentences, the new topic name, and the corresponding semantic elements to the topic-keyword dictionary to obtain a topic-keyword dictionary with a new topic added. Continue to update the topic-keyword dictionary after adding new topics based on the sentences with unknown topics among the remaining topics in the training corpus text until all sentences are clustered, and obtain a new topic-keyword dictionary.

7. The method according to claim 1, characterized in that, The method further includes: Obtain an evaluation corpus text; Evaluate the topic-keyword dictionary based on the evaluation corpus text to obtain an evaluation result; If the evaluation result does not meet the evaluation criteria, obtain a new training corpus text and continue to update the topic-keyword dictionary until the evaluation result corresponding to the topic-keyword dictionary meets the evaluation criteria.

8. An apparatus for constructing a subject keyword dictionary, characterized in that The device includes: An acquisition module for acquiring a training corpus text; A clustering module for performing Latent Dirichlet Allocation (LDA) clustering on the training corpus text based on the topic-keyword dictionary to be trained, and obtaining a clustering result, where the clustering result at least includes the topic corresponding to each sentence in the training corpus text; An update module for updating the topic-keyword dictionary to be trained based on the sentences with known topics in the training corpus text to obtain an updated topic-keyword dictionary.

9. An electronic device, characterized in that, The device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for constructing the topic-keyword dictionary according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the method for constructing the topic-keyword dictionary according to any one of claims 1-7 when executed.