A Method for Constructing an Oilfield Environmental Protection and Safety Standard Knowledge Graph
By building a knowledge map for environmental protection and safety standards in oilfields, the problems of traditional backwardness in standard acquisition methods and data fragmentation in environmental protection and safety management of oilfield enterprises have been solved, and the rapid and accurate acquisition and unified management of knowledge have been achieved.
Patent Information
- Application Number
- CN202310952152.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-07-31
AI Technical Summary
Oilfield enterprises face the problems of traditional backward standard acquisition methods, fragmented data and poor unified integration capabilities in environmental protection and safety management, which makes it difficult to quickly and accurately acquire valuable knowledge.
A method for building a knowledge graph for environmental protection and safety standards in oil fields is proposed. By pre-processing the knowledge text data of environmental protection and safety standards in oil fields, building a knowledge model, using the BERT model for joint extraction of entity relationships, building a data layer of the knowledge graph, and instantiating it with a relational database and graph database.
The construction of the knowledge graph in the field of oil field environmental protection and safety has been realized, and knowledge query, retrieval, entity recognition, relationship extraction and model incremental training has been supported, which has improved the unified management and rapid acquisition of standard knowledge.
Smart Images

Figure CN117131201B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining, and particularly relates to a method for constructing an oilfield environmental protection and safety standard knowledge graph. Background Art
[0002] The digital transformation of standards has become a strategic task for the development of key industries in China in the new era, and is of great significance for enhancing the security of China's industrial development and participating in global market competition.
[0003] Since the development of the oil and gas industry to date, an industrial system with more than a dozen specialties such as oil exploration, development, oil and gas gathering and transportation, and storage and transportation has been formed. With the rapid development of China's oil and gas industry, the enterprise scale has been continuously increasing, and the difficulty of oil and gas exploration and development has also been continuously increasing. The safety situation of oilfield enterprises is relatively serious, facing unprecedented challenges and competitions. At the same time, more and more safety, health, and environmental problems have emerged. The safety and environmental protection problems in aspects such as the objects faced in oil and gas exploration and development, the required technical conditions, new processes, and new technology applications are becoming increasingly prominent.
[0004] The safety and environmental protection issues of oilfields are two important issues in oilfield production, and the society also pays great attention. The relevant departments of oilfields are strengthening the understanding of these two aspects of issues, using correct strategies, and taking effective methods. The daily inspections of oilfield safety and environmental protection professionals include the on-site management of major risk sources, key and vital parts, and direct operation links, etc., with a focus on highlighting the implementation of safety and environmental protection quality measures in links such as offshore, well control, hydrogen sulfide, hazardous chemicals, environmental protection, and crowded places; inspecting the safety and environmental protection quality supervision of contractors, carriers, etc. by each unit, and inspecting the on-site construction management of contractors. The following difficulties exist during the inspections:
[0005] First, the number of current oilfield standards and regulations is huge, and the degree of digitization is not high. The standard systems in use in Shengli Oilfield include a standard dynamic management system, a standard formulation and revision system, a Shengli Oilfield standard query system, a technical supervision management platform, etc. The historical standard data of each platform is stored independently, and the standard specifications and formats are not unified, which brings inconvenience to the unified use and centralized management of standards.
[0006] Second, the standard acquisition method is relatively traditional and backward, and the uncertainty risk of subjective human activities on the use of standards is relatively high. There are many safety risk points in oilfields, which are widely distributed and the operation sites are scattered. The environmental protection and safety management of oilfield enterprises is a typical knowledge-intensive task. However, this knowledge is scattered in various materials, such as standards and specifications, construction organization design plans, safety technical plans, accident investigation reports, and other various technical and management materials. A large amount of professional knowledge support is required in the process of oilfield environmental protection and safety management. Due to the large amount of data, various types, and wide sources of these materials, it is difficult to quickly and accurately obtain valuable knowledge from them. However, there is still a lack of a unified standard semantic knowledge base in China, and the ability to integrate industry standards is relatively poor. Therefore, a method for extracting oilfield environmental protection and safety standard knowledge and constructing a knowledge graph is needed. Summary of the Invention
[0007] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary section is not a comprehensive review, nor is it intended to identify key / important elements or delineate the scope of protection of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.
[0008] The present invention proposes a method for constructing an oilfield environmental protection and safety standard knowledge graph, which is improved in that it includes:
[0009] (1) Preprocess the text data in the oilfield environmental protection and safety standard knowledge to construct an oilfield environmental protection and safety knowledge model;
[0010] (2) Construct a schema layer through the oilfield environmental protection and safety knowledge model;
[0011] (3) Use the pre-trained model BERT to train the schema layer to obtain a joint extraction training model for oilfield environmental protection and safety standard knowledge entity relationships;
[0012] (4) Input the single sentences in the validation set into the joint extraction training model for oilfield environmental protection and safety standard knowledge entity relationships, obtain the entity relationship attributes of each single sentence, and perform entity attribute relationship recognition and extraction;
[0013] (5) Predict the relationship triples of the entity attribute relationships after recognition and extraction, and construct the data layer of the knowledge graph;
[0014] (6) Store the oilfield environmental protection and safety data source knowledge in a relational database, and combine the schema layer with the data layer of the knowledge graph in a graph database to obtain the instantiation of the knowledge graph.
[0015] Preferably, step (1) includes using artificially constructed semantic rules, templates, and dictionaries to perform entity extraction, attribute extraction, and relationship extraction on semi-structured and unstructured data in the use case table. Entities E = oilfield business dimensions are extracted according to specific things. Based on the Mysql relational database and the knowledge ontology theory, and adopting the thesaurus organization method, a triple core data model applicable to oilfield enterprises is studied, namely standardized object - style - index, which constitutes the oilfield environmental protection and safety knowledge model.
[0016] Furthermore, the method for preprocessing the data includes removing stop words, punctuation marks, and spelling errors.
[0017] Preferably, step (3) includes
[0018] (3-1) Add the table name to the limited scope institution and expand the data so that the data quantities of the four types of exploration and development, surface engineering, public works, and offshore are of the same order of magnitude;
[0019] (3-2) Mix the four types of data and divide them into three sets of train, dev, and test in a ratio of 6:2:2 for training and testing.
[0020] Furthermore, the ratio division classification model uses the BERT model to encode the text.
[0021] Preferably, in step (4), entity attribute relationship recognition and extraction include
[0022] (4-1) Obtain the decoded corpus;
[0023] (4-2) Perform part-of-speech tagging on the corpus;
[0024] (4-3) Construct a dependency relationship path and annotate the path according to the dependency annotation table;
[0025] (4-4) Extract the core words;
[0026] (4-5) Construct a verb-object structure and find the subject and object as two entities with the core word as the relationship.
[0027] Preferably, in step (5), predicting the relationship triple includes
[0028] (5-1) Use a deep learning model to predict the relationship triple of each text statement, and use accuracy and recall metrics to evaluate the performance of the model. Use a neural network model to encode the text statement;
[0029] (5-2) Use the encoded vector to represent each text statement, and use a classifier to predict the relationship of the relationship triple of each text statement;
[0030] (5-3) Use the softmax function to map the vector to the categories of relational triples, and use the cross-entropy loss function to train the model.
[0031] Further, encoding the vector includes:
[0032] Set initial values:
[0033]
[0034] where the probability δ t (i) belonging to state i at time t, the hidden state serial number ψ t (i) of state i at time t, the confusion matrix b i ;
[0035] Recursive calculation:
[0036]
[0037] where the duration T of the entire time series, the possible number of states N, the sequence length k, and the transition matrix a of the hidden state ij ;
[0038]
[0039] where the arg max function finds the parameter when the probability ψ t (i) takes the maximum value;
[0040] Predict the optimal state sequence:
[0041]
[0042] Further, word segmentation of the encoded vector includes:
[0043] Cut the string with word segmentation from left to right into w1, w2,..., w m ; Calculate the probability of the current word and the previous word:
[0044]
[0045] where the string has m words and the previous n related words, 1 ≤ n ≤ N;
[0046] Calculate the cumulative probability value of this word:
[0047]
[0048] Retain the large cumulative probability until the end of the string.
[0049] Further, the oilfield operations include exploration and development, surface engineering, utilities, and offshore operations.
[0050] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0051] The present invention acquires and cleans data in the field of environmental protection and safety of pilot oilfields, and constructs a multi-source corpus. Taking textbooks, industry standards and specifications in the field of environmental protection and safety of pilot oilfields as data sources, and using the data collected in the National Standards Library for data supplementation, so as to obtain semi-structured data and unstructured text data of knowledge, and construct a multi-source corpus.
[0052] Based on the theoretical research of the knowledge graph construction method, the present invention designs the overall architecture of the knowledge graph construction in the field of oilfield environmental protection and safety. According to entity extraction and relationship extraction. In the process of ontology and index extraction, machine learning is applied to analyze the part-of-speech of the literature to complete the construction of knowledge triples; and the triple data is stored as specific files according to the information type and stored in the database by a program.
[0053] The present invention realizes the construction of the knowledge graph in the field of environmental protection and safety of pilot oilfields through processes such as data acquisition, information extraction, and knowledge storage, and realizes the knowledge query function, retrieval, entity recognition, relationship extraction, and model incremental training of the knowledge graph system in the field of environmental protection and safety of pilot oilfields.
[0054] The present invention conducts a structured analysis on the structure of safety and environmental protection standard documents in the oilfield field, and constructs a standard knowledge triple data model in the field of oilfield environmental protection and safety, that is, ontology (standardized object) - style (standard paragraph structure) - standard index. Among them, synonyms and hyponym relationships need to be established for both the ontology and the style, and the standard index also includes index items, index values, measurement units, limitation classes, etc., so as to realize the fragmented analysis of the literature and realize the representation of standard knowledge in the field of oilfield environmental protection and safety.
[0055] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0057] Figure 1 is a schematic flowchart of a method for constructing a knowledge graph of oilfield environmental protection and safety standards shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The following description and the drawings fully illustrate specific embodiments of the present invention, enabling those skilled in the art to practice them. The embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations can vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. The scope of the embodiments of the present invention includes the entire scope of the claims and all available equivalents of the claims. In this document, the embodiments may be individually or collectively referred to by the term "invention" merely for convenience, and if more than one invention is actually disclosed, it is not intended to automatically limit the scope of the application to any single invention or inventive concept. In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method or device comprising a series of elements includes not only those elements but also other elements not expressly listed. The various embodiments in this document are described in a progressive manner, with each embodiment highlighting the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the structures, products, etc. disclosed in the embodiments, since they correspond to the parts disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0059] The present invention will be further described below with reference to the drawings and embodiments:
[0060] The present application discloses a data mining method for safety standard knowledge based on the knowledge ontology theory, including constructing an oilfield standard professional ontology according to a standard glossary and oilfield professional vocabulary; obtaining entity extraction corpus for the oilfield standard professional ontology based on a bidirectional long short-term memory network model; inputting the entity extraction corpus into a deep neural network entity relationship extraction model to implement relationship extraction and output oilfield semantic relationships; and completing the construction of entity spatial relationship knowledge data mining according to the output triple of oilfield semantic and spatial relationship information. The present invention provides decision-making services for major issues such as mining the potential value of knowledge ontology entity knowledge big data by mining various oilfield information and data and integrating the standard semantic knowledge base and industry standards capabilities uniformly.
[0061] Embodiment 1
[0062] As Figure 1 shown, the present invention provides a method for constructing an oilfield environmental protection and safety standard knowledge graph, including the following steps:
[0063] Step 1: Preprocess the Chinese text data of oilfield environmental protection and safety standards knowledge, and construct an oilfield environmental protection and safety knowledge model;
[0064] Step 2: Construct a pattern layer according to the oilfield environmental protection and safety knowledge model;
[0065] Step 3: Use the pre-trained model BERT to train the pattern layer to obtain a joint extraction training model for oilfield environmental protection and safety standard knowledge entity relationships;
[0066] Step 4: Input the single sentences in the validation set into the joint extraction training model for oilfield environmental protection and safety standard knowledge entity relationships, and obtain the entity relationship attributes of each single sentence through a scoring function; use sparse multi-label cross-entropy decoding to identify and extract entity attribute relationships; where the validation set is a data set used to verify and evaluate the model performance. In the task of oilfield environmental protection and safety standard knowledge, the validation set is a part of the data extracted from real data, used to verify the model and evaluate the accuracy and generalization ability of the model in this task.
[0067] Step 5: Predict the relationship triples of the entity attribute relationships after identification and extraction, and construct the data layer of the knowledge graph;
[0068] Step 6: Store the oilfield environmental protection and safety data source knowledge using a relational database, and combine the pattern layer with the data layer of the knowledge graph in a graph database to obtain the instantiation of the knowledge graph.
[0069] In the above technical solution, the method for preprocessing the data includes removing stop words, punctuation marks, and spelling mistakes.
[0070] In the above technical solution, in Step 1, the oilfield environmental protection and safety knowledge model uses artificially constructed semantic rules, templates, and dictionaries to perform entity extraction, attribute extraction, and relationship extraction on semi-structured and unstructured data in the use case table, extracts entities E = oilfield business dimensions (such as exploration and development, surface engineering, utility engineering, and offshore, etc.) according to specific things, based on the Mysql relational database, based on the knowledge ontology theory, and adopts organizational methods such as thesaurus to study the triple core data model applicable to oilfield enterprises, that is, standardized object - style - index, to form the oilfield environmental protection and safety knowledge model.
[0071] When constructing the standard knowledge ontology in the field of oilfield environmental protection and safety, it is necessary to expand and refine the relationships between concepts according to the specific conditions of relevant standards in the field of oilfield environmental protection and safety. The relationships between concepts of relevant standards in the field of oilfield environmental protection and safety of the present invention are divided into two categories: general relationships and custom relationships. Combining the ontology division framework table constructed in this report and the standards classified according to the standard knowledge system in the field of oilfield environmental protection and safety, the relationships between concepts of relevant standards in the field of oilfield environmental protection and safety are divided.
[0072] In the above technical solution, in step 3, the table name is added to the limited scope mechanism to expand the data so that the order of magnitude of the four types of data is the same. The four types of data are mixed and divided into three sets of train, dev, and test according to the ratio of 6:2:2 for training and testing. The classification model first encodes the text using the BERT model.
[0073] Among them, the limited scope mechanism is a data processing method used to filter or screen certain specific table names in the dataset. This mechanism selects only the data samples that meet the requirements of specific table names by using the table name as a limiting condition.
[0074] Among them, the four types are the data of the above exploration and development (type A), surface engineering (type B), public works (type C), and offshore (type D). These types can be defined according to different table names, and each type has different characteristics and sample numbers.
[0075] The 768-dimensional vector output by the BERT model is passed into a fully connected network with one hidden layer, and the dropout strategy is used. The output layer has 4 nodes, which are activated using the softmax function. The output 4-dimensional vector represents the classification result, that is, the possibility under the four classifications. During the training process, cross-entropy is used as the loss function for the classification task.
[0076] The parameter settings during training are as follows:
[0077] Parameter Name Parameter Value Note train_batch_size 32 max_seq_length 128 Maximum Sequence Length num_train_epochs 3 warmup_proportion 0.1 Learning Rate Warmup Proportion learning_rate 5e-5
[0078] The final training result can reach 95% accuracy on the validation set.
[0079] In the above technical solution, the entity attribute relationship recognition and extraction method includes.
[0080] (1) Obtain the decoded corpus;
[0081] (2) Perform part-of-speech tagging on the corpus;
[0082] (3) Construct a dependency relationship path and annotate the path according to the dependency annotation table;
[0083] (4) Extract core words;
[0084] (5) Construct a subject-object structure and use the core word as the relationship to find the subject and object as two entities.
[0085] In the above technical solutions, the corpus is obtained from a variety of sources, including professional field corpus and network text. The goal of these tasks is to perform structured analysis and semantic understanding of natural language text to extract part-of-speech information, relationship information, grammatical structure, etc. Therefore, the corpus data provides the basic information required for the task, and through the processing and analysis of the corpus, more structured information can be obtained to support subsequent semantic parsing and reasoning.
[0086] In the above technical solution, the relationship triplet prediction method includes using a deep learning model to predict the relationship triplet of each text sentence, and using indicators such as accuracy and recall to evaluate the performance of the model, using a neural network model to encode the text sentence, and then using the encoded vector to represent each text sentence, using a classifier to predict the relationship of the relationship triplet of each text sentence, using a softmax function to map the vector to the category of the relationship triplet, and then using a cross entropy loss function to train the model.
[0087] A structured analysis is conducted on the structure of standard documents on oilfield safety and environmental protection, and a ternary metadata model of standard knowledge in the field of oilfield environmental protection and safety is constructed, namely, ontology (standardized object)-style (standard paragraph structure)-standard indicators. Both ontology and style need to establish synonyms and hierarchical relationships, and standard indicators also include indicator items, indicator values, measurement units, and restriction categories, thereby realizing document fragmentation analysis and the representation of standard knowledge in the field of oilfield environmental protection and safety.
[0088] Example 2
[0089] The triple data corresponding to the ontology (standardized object) - style (standard paragraph structure) - standard indicator is "flame-retardant clothing - storage - storage conditions", and the disclosure content corresponding to this indicator is "the storage conditions of flame-retardant clothing are that it must not be placed together with corrosive items, the storage place should be dry and ventilated, avoid direct sunlight, and the packaging should be at least 20 cm away from the wall and the ground to prevent rat bites, insect infestation, and mildew."
[0090] Example 3
[0091] Numerical fragmented knowledge (taking GB2811-2019 "Head Protection Safety Helmets" as an example): the triple data corresponding to the ontology (standardized object) - style (standard paragraph structure) - standard indicator is "safety helmets - quality (excluding accessories) - relative error between the actual quality of the product and the marked quality". The corresponding disclosure content of this indicator is "the relative error between the actual quality of the safety helmet product and the marked quality should not be greater than 5%".
[0092] The present application also includes a structure of an electronic device. At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage, etc. Of course, the electronic device may also include hardware required for other services.
[0093] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0094] The memory is used to store the program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.
[0095] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a data mining device for safety standard knowledge based on knowledge ontology theory at the logical level. The processor executes the program stored in the memory and is specifically used to execute any of the aforementioned oilfield environmental safety standard knowledge extraction and knowledge graph construction methods.
[0096] The present invention can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general-purpose processor, including a central processing unit, a network processor (Network Processor, NP), etc.; it can also be a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0097] An embodiment of the present application also proposes a computer-readable storage medium, which stores one or more programs, and the one or more programs include instructions. When the instructions are executed by an electronic device including multiple application programs, they execute any of the aforementioned oilfield environmental protection and safety standards knowledge extraction and knowledge graph construction methods.
[0098] It should be understood that the above scheme has presented the foregoing description of specific exemplary embodiments of the present invention for the purpose of illustration and description. It is not intended to exclude or limit the present invention to the precise form disclosed, and it is obvious that many modifications and changes are feasible in view of the above teachings. Select and describe exemplary embodiments to explain certain principles of the present invention and their practical application, so that other technical personnel in the field can make or use various exemplary embodiments of the present invention, and various alternatives and modifications thereof. Its purpose is that the scope of the present invention will be defined by the claims attached to the present invention and their equivalents.
[0099] It is understandable that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present invention. However, the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also regarded as the protection scope of the present invention.
Claims
1. A method for constructing an oilfield environmental protection and safety standard knowledge graph, characterized in that Including: (1) Preprocess the Chinese text data of oilfield environmental protection and safety standard knowledge to construct an oilfield environmental protection and safety knowledge model; (2) Construct a pattern layer through the oilfield environmental protection and safety knowledge model; (3) Use the pre-trained model BERT to train the pattern layer to obtain an oilfield environmental protection and safety standard knowledge entity relationship joint extraction training model; (4) Input single sentences in the validation set into the oilfield environmental protection and safety standard knowledge entity relationship joint extraction training model, obtain the entity relationship attributes of each single sentence, and perform entity attribute relationship recognition and extraction; (5) Predict relationship triples based on the recognized and extracted entity attribute relationships, and construct the data layer of the knowledge graph; (6) Use the Mysql relational database to store the oilfield environmental protection and safety data source knowledge, and combine the pattern layer with the data layer of the knowledge graph in the graph database to obtain the instantiation of the knowledge graph; Among them, step (1) includes using artificially constructed semantic rules, templates, and dictionaries to perform entity extraction, attribute extraction, and relationship extraction on semi-structured and unstructured data respectively. Entities are extracted according to specific things, E = oilfield business dimension. Based on the Mysql relational database and the knowledge ontology theory, using the thesaurus organization method, study the triple core data model applicable to oilfield enterprises, that is, standardized object - style - index, to form the oilfield environmental protection and safety knowledge model; The predicted relationship triples in step (5) include: (5-1) Use a deep learning model to predict the relationship triples of each text statement, use accuracy and recall metrics to evaluate the performance of the model, and use a neural network model to encode the text statements; (5-2) Use the encoded vectors to represent each text statement, and use a classifier to predict the relationships of the relationship triples of each text statement; (5-3) Use the softmax function to map the vectors to the categories of relationship triples, and use the cross-entropy loss function to train the model; Encoding the vectors includes: Setting initial values: ; Among them, the probability of belonging to state i at time t , the hidden state serial number of state i at time t , confusion matrix ; Recursive calculation: ; Among them, for the duration T of the entire time series, there are possible state numbers N, sequence length k, and the transition matrix of hidden states ; ; Among them, The function calculates the probability of the parameter when the maximum value is taken; Predicting the optimal state sequence: 。 2. The method for constructing an oilfield environmental protection and safety standard knowledge graph according to claim 1, characterized in that The method for preprocessing the data includes removing stop words, punctuation marks, and spelling mistakes.
3. A method for constructing an oilfield environmental protection and safety standard knowledge graph according to claim 1, characterized in that, Step (3) includes (3-1) Add the table name to the limited scope institution and expand the data so that the data quantities of the four types of exploration and development, surface engineering, public works, and offshore are of the same order of magnitude; (3-2) Mix the four types of data and divide them into three sets of train, dev, and test according to the ratio of 6:2:2 for training and testing.
4. A method for constructing an oilfield environmental protection and safety standard knowledge graph according to claim 3, characterized in that The classification model for the ratio division uses the BERT model to encode the text.
5. A method for constructing an oilfield environmental protection and safety standard knowledge graph according to claim 1, characterized in that, The entity attribute relationship recognition and extraction in step (4) includes (4-1) Obtain the decoded corpus; (4-2) Perform part-of-speech tagging on the corpus; (4-3) Construct a dependency relationship path and annotate the path according to the dependency annotation table; (4-4) Extract the core words; (4-5) Construct a start-object structure, and use the core word as the relationship to find the subject and object as the two entities.
6. The method for constructing an oilfield environmental protection and safety standard knowledge graph according to claim 1, wherein, Tokenizing the encoded vectors includes: Split the string with word segmentation from left to right into ; Calculate the probability of the current word and the previous word: ; Among them, there are m string words and the first n related words, ; Calculating the cumulative probability value of the word: ; Retain the large cumulative probability until the end of the string.
7. A method for constructing an oilfield environmental protection and safety standard knowledge graph according to claim 1, characterized in that The oilfield operations include exploration and development, surface engineering, utilities, and offshore operations.
Citation Information
Patent Citations
Vertical domain knowledge graph construction method and system
CN113177124A
Transmission solution generation method and system based on transmission knowledge graph
CN114218406A