Knowledge graph-based oil and gas field surface engineering case library construction method and system
By introducing knowledge graph technology into the oil and gas field ground engineering case library, the shortcomings of the traditional case library in data quality, update difficulties and semantic understanding have been solved, efficient construction, update and retrieval are achieved, knowledge reasoning capabilities are provided, and intelligent engineering decision-making is supported.
Patent Information
- Application Number
- CN202311520806.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-23
AI Technical Summary
The traditional oil and gas field ground engineering case library has defects in data quality, difficulty in updating, lack of semantic understanding and reasoning capabilities, and limitations in professional fields, and is difficult to meet the needs of intelligent and cross-domain applications.
The knowledge graph-based method is used to build a case library for oil and gas field ground engineering. By pre-processing, classification, intelligent dismantling and knowledge graph construction of case files, it realizes automated storage and efficient retrieval, and has the ability to reason.
It improves the construction efficiency and update speed of the case library, enhances the credibility and practicality of the data, provides more efficient retrieval and rich knowledge reasoning capabilities, and supports a wider range of engineering practices and decision-making support.
Smart Images

Figure CN120030168A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent engineering technology, and in particular to a method for constructing an oil and gas field surface engineering case library based on a knowledge graph. Background Art
[0002] Oil and gas surface engineering is a necessary link in the development and production of oil and gas fields. It is an important aspect to achieve efficient development, reflect development effects and economic and technological levels, and is an important means to reduce investment control costs and improve development benefits. However, the variety of oil and gas fields, complex development objects, advanced development technology, and the need for coordination and cooperation among multiple departments have increased the complexity of oil and gas field surface engineering technology and operation procedures.
[0003] Establishing an oil and gas field surface engineering case library can assist engineering practice decision-making and improve the level of standardized operation and information management of oil and gas field surface engineering. The construction of a traditional oil and gas field surface engineering case library mainly includes: case data collection and organization, data storage and indexing, usually using a relational database (such as MySQL) or a document database (such as MongoDB) for storage.
[0004] The traditional construction method of oil and gas field surface engineering case library has the following major defects: (1) Data quality issues: The data sources of traditional engineering case libraries are wide, including literature, reports, expert experience, etc., and the data quality varies. Some data may be erroneous, incomplete or inaccurate, which will affect the credibility and practicality of the case library. (2) Difficulty in data updating: Traditional engineering case library construction methods are mostly manual, requiring manual collection, organization and updating of data. When faced with a large number of data updates and changes, the workload of maintaining and updating the case library is large, and omissions and delays are prone to occur. This makes it impossible for the data in the case library to reflect the latest engineering practices and technological developments in a timely manner. (3) Lack of semantic understanding and reasoning capabilities: Traditional engineering case libraries usually simply store and retrieve case information, lacking deep semantic understanding and reasoning capabilities. For example, it is difficult for traditional case libraries to understand the associations and patterns between cases, and cannot provide higher-level knowledge discovery and knowledge reasoning, which leads to limited application scenarios and effects. (4) Professional field limitations: Traditional engineering case libraries are usually built for specific industries or fields, and there are problems with professional field limitations. This limits the scope of applicability of case libraries. For projects involving multiple disciplines or across industries, traditional case libraries often cannot provide comprehensive and diverse case support.
[0005] With the development of technologies such as artificial intelligence and knowledge graphs, building a more intelligent, automated, and cross-domain engineering case library has become a development direction. Such a case library can perform semantic understanding, automatic update, efficient retrieval and reasoning, and is widely used in engineering practice and decision support. Summary of the invention
[0006] The present invention provides a method and system for constructing an oil and gas field surface engineering case library based on a knowledge graph to solve the problems that the establishment of a traditional engineering case library is time-consuming and labor-intensive, the case library is difficult to update, and related knowledge cannot be provided.
[0007] The present invention is achieved through the following technical solutions:
[0008] A first aspect of the present invention provides a method for constructing an oil and gas field surface engineering case library based on a knowledge graph, comprising:
[0009] Preprocessing the oil and gas field surface engineering case files to obtain basic information of the engineering case files;
[0010] Classifying the engineering case files according to the basic information, and establishing a storage level for the engineering case files, wherein the storage level is used to store the engineering case files in an engineering case library according to the storage level;
[0011] Intelligently disassemble the engineering case file according to the information disassembly item, obtain the information disassembly item attribute value corresponding to the engineering case file, and form the storage content including the engineering case file and the information disassembly item attribute value;
[0012] Establishing a knowledge graph of oil and gas field surface engineering cases based on the information disassembly items, wherein each entity in the knowledge graph corresponds to an information disassembly item;
[0013] Integrate the stored content and the knowledge graph to build an oil and gas field surface engineering case library based on the knowledge graph.
[0014] The present invention integrates the reasoning ability of the knowledge graph into the traditional oil and gas field surface engineering case library, efficiently classifies and stores engineering case files through basic information, and forms the associated content of case files, information disassembly items and information disassembly item attribute values through intelligent information disassembly, and can realize efficient retrieval through key information. At the same time, classification, information disassembly and storage are all completed automatically, which improves the efficiency of case library construction and efficient update. A knowledge graph of oil and gas field surface engineering cases is established based on information disassembly items, so that the case library has the ability of knowledge reasoning. Each entity in the knowledge graph corresponds to an information disassembly item, so that knowledge reasoning can be performed with the information disassembly item as the origin to obtain all related information in the case library. Provide more efficient retrieval and rich retrieval results to provide decision support for engineering practice.
[0015] In one embodiment, the oil and gas field surface engineering case file is preprocessed to obtain basic information of the engineering case file, including:
[0016] Identify the engineering case file using optical character recognition technology to obtain the engineering case text;
[0017] Perform keyword recognition on the engineering case text to obtain the basic information of the engineering case file.
[0018] In one implementation, classify the engineering case files according to the basic information, and establish the storage hierarchy of the engineering case files, including:
[0019] Classify the engineering cases into the first level according to the basic information of the engineering case files as single project cases and typical construction cases;
[0020] Classify the single project cases and typical construction cases into the second level according to important processes, special environments, and risk operations;
[0021] Under the second-level classification, match the basic information of the engineering case with the existing subclasses under each second level, and classify the engineering case into the corresponding subclass according to the matching result. If the basic information of the engineering case does not match all subclasses, a new subclass is established according to the basic information, and the engineering case is classified into the new subclass;
[0022] Use the classification hierarchy of the engineering case as the storage hierarchy of the engineering case.
[0023] In one implementation, the information decomposition items include: case overview, key parameters, operation time, construction drawing files, construction resources, construction measures, HSE measures, process record data, and picture images, and the attribute values of the information decomposition items are the corresponding contents of the information decomposition items.
[0024] In one implementation, establish an oil and gas field surface engineering case knowledge graph based on the information decomposition items, including:
[0025] S401. Match the engineering case text with the information decomposition items using a rule-based entity recognition method to obtain named entities;
[0026] S402. According to the named entities obtained in step S401, determine a limited number of entity relationships manually to form an entity relationship table;
[0027] S403. Match the entity relationships in the entity relationship table with the case text to generate a template in the form of a five-tuple;
[0028] S404. Perform clustering based on the similarity of each template, and average the templates belonging to the same class to obtain a new template;
[0029] S405, calculating the similarity between the new template and the entity relationship in the entity relationship table, discarding the templates with similarity less than a threshold, forming a new entity relationship with the templates with similarity greater than or equal to the threshold, and adding the templates to the entity relationship table;
[0030] S406. Repeat steps S403 to S405 until all engineering case texts are processed, and generate an oil and gas field surface engineering case knowledge graph based on the entity relationship table.
[0031] In one embodiment, clustering is performed based on the similarity of each template, including:
[0032] Convert each quintuple template <Left+Entity 1+Middle+Entity 2+Right> into vector form:
[0033] X=(L,T 1 , M, T 2 , R)
[0034] Among them, T 1 , T 2 is the word vector of entity 1 and entity 2, L is the word vector on the left of entity 1, M is the word vector between entity 1 and entity 2, R is the word vector on the right of entity 2, and the similarity of any two five-tuple templates i and j is expressed as:
[0035]
[0036] Among them, the vector of template i is expressed as The vector representation of template j is w 1 、w 2 、w 3 is the weight, and w 2 Greater than w 1 and w 3 .
[0037] In one implementation, before step S401, the method further includes:
[0038] The engineering case text is processed in a regularized manner and converted into a standard structured text, including:
[0039] S11, filtering and cleaning punctuation marks based on regular expression judgment to obtain continuous natural language text;
[0040] S12, segmenting continuous natural language text to obtain a vocabulary sequence with semantic rationality and integrity;
[0041] S13. By analyzing the dependency relationship between words in a sentence, the syntactic structure information of the words is captured, and the syntactic structure information of the sentence is represented by a tree structure to obtain a standard structured text.
[0042] A second aspect of the present invention provides a system for building a case library for oil and gas field surface engineering based on a knowledge graph, the system comprising:
[0043] A preprocessing module is used to preprocess the oil and gas field surface engineering case files to obtain basic information of the engineering case files;
[0044] A case classification module, used to classify the engineering case files according to the basic information, and establish a storage level for the engineering case files, wherein the storage level is used to store the engineering case files in an engineering case library according to the storage level;
[0045] An intelligent disassembly module, used to intelligently disassemble the engineering case file according to the information disassembly item, obtain the information disassembly item attribute value corresponding to the engineering case file, and form a storage content including the engineering case file and the information disassembly item attribute value;
[0046] A knowledge graph generation module, used to establish a knowledge graph of oil and gas field surface engineering cases based on the information disassembly items, wherein each entity in the knowledge graph corresponds to an information disassembly item;
[0047] The case library generation module integrates the stored content and the knowledge graph to construct an oil and gas field surface engineering case library based on the knowledge graph.
[0048] A third aspect of the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for constructing a surface engineering case library for oil and gas fields according to any one of the embodiments of the present invention is implemented.
[0049] A fourth aspect of the present invention provides a computer storage medium having computer instructions stored thereon, the computer instructions being executable by a processor to implement the method for constructing an oil and gas field surface engineering case library according to any embodiment of the present invention.
[0050] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0051] By preprocessing the collected engineering case files, engineering case files from different sources and in different formats are standardized and formatted to facilitate subsequent information extraction, analysis and other processing, thereby improving text processing efficiency.
[0052] Automatically classify case files based on basic information to build a storage hierarchy for case files, achieve semi-structuring of engineering case data, improve data processing and classification efficiency, and provide specific traceability through basic information, which facilitates subsequent classification updates and case library maintenance.
[0053] The attribute values of information disassembly items obtained through intelligent information disassembly are stored together with the case files, forming a storage content including the original files of the engineering case and the attribute values of the corresponding information disassembly items, so as to facilitate the retrieval of key information. The original file information can be obtained through key information retrieval, and the database stores more comprehensive information.
[0054] The information disassembly item establishes the knowledge graph of oil and gas field surface engineering cases, so that the case library has the ability of knowledge reasoning. Each entity in the knowledge graph corresponds to an information disassembly item. Therefore, knowledge reasoning can be performed with the information disassembly item as the origin to obtain all related information in the case library.
[0055] Since the associated information of the original case file, basic information, information decomposition items, and information decomposition item attribute values is formed in advance, the knowledge network based on the knowledge graph can search the case library more flexibly, and different search instructions can be associated with the entities in the knowledge graph for knowledge reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative work. In the drawings:
[0057] Figure 1 It is a flow chart of a method for constructing an oil and gas field surface engineering case library based on a knowledge graph according to an embodiment of the present invention;
[0058] Figure 2 This is a flowchart of a method for establishing a knowledge graph of oil and gas field surface engineering cases according to an embodiment of the present invention.
[0059] Figure 3 The figure is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0060] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.
[0061] An embodiment of the present invention provides a method for constructing an oil and gas field surface engineering case library based on a knowledge graph, which is applicable to the construction of an engineering case library in the field of oil and gas field surface engineering, is beneficial to improving the construction efficiency of the case library and the knowledge reasoning ability, and provides decision-making support for oil and gas field surface engineering construction.
[0062] As Figure 1 shown, Figure 1 is a flowchart of a method for constructing an oil and gas field surface engineering case library based on a knowledge graph. The case library construction method includes:
[0063] Step S1: Preprocess the oil and gas field surface engineering case file to obtain the basic information of the engineering case file.
[0064] The data sources of the oil and gas field surface engineering case library are extensive, including literature, reports, expert experience, etc. The data quality is uneven, and some data may be incorrect, incomplete or inaccurate. This will affect the credibility and practicability of the case library. Therefore, it is first necessary to preprocess the collected case library text, remove redundant and irrelevant information, convert it into a unified text format for processing, extract useful basic information through natural language processing technology, and extract and locate key information, so as to classify and organize unstructured data.
[0065] The oil and gas field surface engineering case library can be collected manually, or directly from the cases in the existing database, or crawled from the Internet websites related to oil and gas field development.
[0066] In an optional embodiment, the above preprocessing method for the engineering case file includes first using optical character recognition technology to recognize the engineering case file to obtain the engineering case text; then performing keyword recognition on the engineering case text to obtain the basic information of the engineering case file.
[0067] The same engineering case file usually includes various forms such as text, pictures, and tables, which causes difficulties in automatic recognition and extraction of useful information. Therefore, first, the engineering case file is converted into machine-readable text through optical character recognition technology (OCR). OCR can extract the text information in images and tables. The basic processing process includes image preprocessing, text line detection, single-character segmentation, single-character recognition, and post-processing. The collected engineering case text can be converted into a unified PDF format and then OCR recognition is performed. After obtaining the text information of the engineering case, the key information in the text is located through keyword recognition technology, that is, the basic information of the engineering case file.
[0068] According to the characteristics of oil and gas field surface engineering, basic information includes: project name, project number, project location, start time, acceptance time, transfer time, whether it is digitally handed over, project pictures, oil and gas field type, engineering type or special environment, etc. Text matching is performed by pre-setting keywords corresponding to basic information to quickly extract basic information, and the basic information is associated with the engineering case file for storage, which is easy to retrieve.
[0069] In an optional embodiment, the preprocessing of the engineering case file also includes converting the impure, disordered, and non-standard natural language text into a regular, easy-to-process, and standard structured text, thereby performing keyword recognition, semantic recognition, information extraction, and other processing on the structured text. The steps of text regularization include:
[0070] S11, filtering and cleaning punctuation marks based on regular expression judgment to obtain continuous natural language text;
[0071] S12, segmenting continuous natural language text to obtain a vocabulary sequence with semantic rationality and integrity;
[0072] S13. By analyzing the dependency relationship between words in a sentence, the syntactic structure information of words is captured, and a tree structure is used to represent the syntactic structure information of the sentence to obtain a standard structured text.
[0073] The characters such as commas, periods, quotation marks, etc. in the text represent the pauses and connections of sentences, and have no actual meaning in semantic analysis. The present invention adopts a regular matching method to match and filter the input text for punctuation marks to obtain a continuous natural language text. The existing word segmentation tools such as Jieba, LTP, etc. can be used for word segmentation.
[0074] The present invention preferably uses a word segmentation method based on string matching to segment continuous natural language texts, matches the entries in the sentence to be segmented with the words in the corpus according to a certain scanning method, and then returns the corresponding results. The word segmentation method based on string matching matches the string to be matched with the words in the dictionary with the help of a dictionary. The algorithm is simple and easy to implement, the word segmentation speed is fast, and it is suitable for processing large quantities of text. Combined with professional dictionaries in the oil and gas field field, the word segmentation results are more accurate.
[0075] Syntactic analysis is an important part of language comprehension. The present invention adopts dependency parsing (DEP) to obtain sentence structure and semantic relationship between words in a sentence, and converts natural language text into structured text with semantics for subsequent processing. Specifically, existing dependency parsing tools such as LTP, Stanford Parser, etc. can be used.
[0076] Step S2: classify the engineering case files according to the basic information, and establish a storage level for the engineering case files, wherein the storage level is used to store the engineering case files in an engineering case library according to the storage level.
[0077] According to the operation type of oil and gas field surface engineering, it can be hierarchically classified according to different operation characteristics. According to the basic information corresponding to the engineering case files, they can be divided into corresponding categories to build a multi-dimensional engineering case storage classification coordinate system.
[0078] In one implementation, all cases are first classified into single project cases and typical construction cases according to the typicality of the project cases. Distinguishing between single project cases and typical construction cases can analyze typical operation information in a targeted manner, and also facilitate the subsequent maintenance and update of the case library.
[0079] In single project cases and typical construction cases, the second level classification is carried out according to the characteristics of the operations into three categories: important processes, special environments, and risky operations.
[0080] Special environments include: construction in winter, construction in rainy season, construction in high temperature, construction at night, construction in places containing hydrogen sulfide, construction in typhoon weather, major exhibitions, construction during holidays, emergency construction during water and power outages, etc.
[0081] Important processes include assessment and procedures, non-destructive testing, anti-corrosion and thermal insulation, container manufacturing and installation, purging and pressure testing, commissioning and commissioning, pile foundation, deep foundation pit, blasting, etc.
[0082] Risky operations include: hot work, earth-moving work, circuit breaking work, high-altitude work, equipment maintenance work, blind plate plugging work, pipeline and equipment opening work, temporary power work, excavation work, entering confined space work, lifting work, high-voltage power work, climbing work, refrigeration and air-conditioning work, drilling driller work, etc.
[0083] Based on the basic information obtained, the above categories are matched by the basic information and the files are classified into corresponding categories as the storage level.
[0084] Since there are various types of surface engineering operations in oil and gas fields, the above are only examples and do not include all types. To ensure classification efficiency and comprehensiveness, the present invention automatically classifies based on basic information recognition.
[0085] In one embodiment of the present invention, under the second-level classification, the basic information of the engineering case is matched with the existing subclasses under each second level, and the engineering case is classified into the corresponding subclass according to the matching result. If the basic information of the engineering case does not match all subclasses, a new subclass is established based on the basic information, and the engineering case is classified into the new subclass.
[0086] The classification level of the engineering case is used as the storage level of the engineering case.
[0087] Based on the basic information, automatic classification and automatic creation of new classifications can be achieved, and case files can be classified into corresponding categories more comprehensively and accurately, thereby improving the efficiency and accuracy of classification. It should be noted that the hierarchical classification method based on the present invention classifies engineering cases in multiple dimensions. The same case file may appear in different categories at different levels. In this way, the association information between different ground operations can be fully retained, which is convenient for the subsequent construction of the engineering case knowledge graph, so that the case library has higher reasoning ability.
[0088] Classification is carried out based on basic information. The correlation between the original case files, basic information and classification levels makes the classification process traceable, which facilitates the subsequent adjustment and update of the classification results.
[0089] Step S3: intelligently disassemble the engineering case file according to the information disassembly items to obtain the information disassembly item attribute values corresponding to the engineering case file, and form storage content including the engineering case file and the information disassembly item attribute values.
[0090] The information breakdown items are more specific information of the engineering case, and are also detailed knowledge content that provides reference and decision-making support for construction, including: case overview, key parameters, operation time, construction drawings, construction resources, construction measures, HSE measures, process record data and pictures, etc.
[0091] According to the information decomposition items in the matching case file, the attribute values corresponding to the information decomposition items, that is, the specific content, are obtained.
[0092] The original case file, basic information, information disassembly items, and information disassembly item attribute values are associated and stored in the database to store the complete and key information of the case file.
[0093] Furthermore, the information decomposition items can be divided into basic information decomposition items and detailed information decomposition items according to single project cases and typical construction cases.
[0094] The basic information breakdown items include: case overview, key parameters and operation time;
[0095] The detailed information breakdown items include: construction drawings, construction resources, construction measures, HSE measures, process record data and picture images.
[0096] Step S4: establishing a knowledge graph of oil and gas field surface engineering cases based on the information decomposition items, wherein each entity in the knowledge graph corresponds to an information decomposition item.
[0097] Taking the information disassembled items as the target entities, entity relationships are extracted from the engineering case text to obtain the relationship network between each entity, thereby establishing the knowledge graph of the oil and gas field surface engineering cases. Since each entity corresponds to the information disassembled item, the information disassembled item can be used as the origin for knowledge reasoning to obtain all related information in the case library. And because the associated content of the original case file, basic information, information disassembled item, and information disassembled item attribute values has been formed and stored in the database in advance, through the knowledge network of the knowledge graph, the case library can be retrieved more flexibly.
[0098] In one embodiment of the present invention, as Figure 2 shown is the method flow chart for establishing the knowledge graph of the oil and gas field surface engineering cases based on the information disassembled items, including:
[0099] S401. Using a rule-based entity recognition method to match the engineering case text with the information disassembled items to obtain named entities.
[0100] S402. According to the named entities obtained in step S401, a limited number of entity relationships are determined manually to form an entity relationship table.
[0101] S403. Matching the entity relationships in the entity relationship table with the case text to generate a template in the form of a five-tuple.
[0102] The form of the five-tuple template is <Left + entity 1 + Middle + entity 2 + Right>. The elements in the template can be set with corresponding word lengths according to the actual situation. Left is the word vector of len1 words on the left of entity 1, Middle is the lexical vector between entity 1 and entity 2, Right is the word vector of len2 words on the right of entity 2, and len1 and len2 are the word lengths, which are selected according to the actual situation. For example, len1 = 2 and len2 = 3.
[0103] S404. Clustering based on the similarity of each template, and averaging the templates belonging to the same class to obtain a new template.
[0104] Specifically, the Snowball algorithm is used for similarity calculation. The templates with similarity greater than the threshold are clustered. After clustering, the average value of each class is taken to obtain a representative template, which is added to the storage template relationship tuple library. The threshold is selected according to the actual situation from 0.7 to 0.9. In this embodiment, 0.85 is taken.
[0105] S405. Calculating the similarity between the new template and the entity relationships in the entity relationship table, discarding the templates with similarity less than the threshold, and forming new entity relationships with the templates with similarity greater than or equal to the threshold, and adding them to the entity relationship table.
[0106] Furthermore, the similarity calculation process is as follows:
[0107] Convert each quintuple template <Left+Entity 1+Middle+Entity 2+Right> into a vector form:
[0108] X=(L,T 1 , M, T 2 ,R) (1)
[0109] Among them, T 1 , T 2 is the word vector of entity 1 and entity 2, L is the word vector on the left of entity 1, M is the word vector between entity 1 and entity 2, R is the word vector on the right of entity 2, and the similarity of any two five-tuple templates i and j is expressed as:
[0110]
[0111] In the above formula, the vector of template i is expressed as The vector representation of template j is If or or The similarity between the target template and the comparison template is 0; if and The target template and the comparison template have the same entity type, and the similarity between them is calculated as SIM = w 1 L i L j +w 2 M i M j +w 3 R i R j , w 1 、w 2 、w 3 is the weight, and w 2 Greater than w 1 and w 3 Since the intermediate word vector has a greater impact on similarity, the weight w is generally 2 Set to maximum weight value.
[0112] S406. Repeat steps S403 to S405 until all engineering case texts are processed, and generate an oil and gas field surface engineering case knowledge graph based on the entity relationship table.
[0113] In this embodiment, a small number of entity relationships are manually determined to form a template, which is matched with the sentences in the case text that include the entities in the template. More entity relationships are extracted to form a new template, and the new template is screened. The templates that meet the requirements are added to the entity relationship table to expand the knowledge graph network until all texts are processed. A knowledge graph is established according to the entity relationship table.
[0114] For the update of the knowledge graph, when new case text is obtained, the above-mentioned steps S403 - S405 of entity relationship extraction are performed on the new text, that is, one update of the knowledge graph is completed. Correspondingly, when new case text is obtained, steps S1 - S4 are performed, that is, one update of the engineering case library is completed.
[0115] Further, in this embodiment, before establishing the oil and gas field surface engineering case knowledge graph based on the information decomposition items, it also includes a preprocessing step of converting the engineering case text into standard structured text, and performing entity extraction on the structured text. The preprocessing steps are as described in the above embodiment and will not be elaborated here.
[0116] Step S5: Integrate the warehousing content and the knowledge graph to construct an oil and gas field surface engineering case library based on the knowledge graph.
[0117] Through steps S1 - S3, the associated warehousing content of the case original file, basic information, information decomposition items, and information decomposition item attribute values has been formed. Through step S4, a knowledge graph network with information decomposition items as entities has been established. According to the warehousing level corresponding to the case file in step S2, an oil and gas field surface engineering case library based on the knowledge graph is constructed.
[0118] In one implementation manner, the above-mentioned oil and gas field surface engineering case library is displayed in a visual manner, a retrieval entry is provided through a visual interface, and retrieval tag items are set according to the basic information and information decomposition items. The associated information is queried in the case library according to the retrieval instruction, and based on the reasoning ability of the knowledge graph, relevant operation information is associated through entity relationships to provide more comprehensive reference information and provide decision support for engineering personnel.
[0119] In the second aspect of the present invention, a system for constructing an oil and gas field surface engineering case library based on a knowledge graph is provided, including:
[0120] A preprocessing module for preprocessing the oil and gas field surface engineering case file to obtain the basic information of the engineering case file;
[0121] A case classification module for classifying the engineering case file according to the basic information, establishing the warehousing level of the engineering case file, and the warehousing level is used to warehouse the engineering case file into the engineering case library according to the warehousing level;
[0122] An intelligent disassembly module, used to intelligently disassemble the engineering case file according to the information disassembly item, obtain the information disassembly item attribute value corresponding to the engineering case file, and form a storage content including the engineering case file and the information disassembly item attribute value;
[0123] A knowledge graph generation module, used to establish a knowledge graph of oil and gas field surface engineering cases based on the information disassembly items, wherein each entity in the knowledge graph corresponds to an information disassembly item;
[0124] The case library generation module integrates the stored content and the knowledge graph to construct an oil and gas field surface engineering case library based on the knowledge graph.
[0125] A third aspect of the present invention provides an electronic device, such as Figure 3 As shown, Figure 3 FIG. 4 is a schematic diagram of the structure of an electronic device of the present invention, wherein the electronic device comprises a processor 40, a memory 41, an input device 42, an output device 43 and a communication device 44; the number of processors 40 in the computer device can be one or more. Figure 3 The processor 40, the memory 41, the input device 42 and the output device 43 in the electronic device can be connected by a bus or other means. Figure 3 The example of connecting through bus is taken in the following.
[0126] The memory 41 is a computer-readable storage medium that can be used to store software programs, computer executable programs, and modules. The processor 40 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 41, so as to implement the method for constructing an oil and gas field surface engineering case library based on a knowledge graph according to any of the above embodiments of the present invention.
[0127] The memory 41 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 41 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 41 may further include a memory remotely arranged relative to the processor 40, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0128] The input device 42 can be used to receive oil and gas field case file data; the output device 43 is used to output data processing results.
[0129] In a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method for constructing a knowledge graph-based oil and gas field surface engineering case library according to any of the above embodiments of the present invention is implemented. The storage medium may be a ROM / RAM, a magnetic disk, an optical disk, etc.
[0130] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for constructing a case library of oil and gas field surface engineering based on knowledge graph, It is characterized in that The method comprises: Preprocessing the oil and gas field surface engineering case files to obtain basic information of the engineering case files; Classifying the engineering case files according to the basic information, and establishing a storage level for the engineering case files, wherein the storage level is used to store the engineering case files in an engineering case library according to the storage level; Intelligently disassemble the engineering case file according to the information disassembly item, obtain the information disassembly item attribute value corresponding to the engineering case file, and form the storage content including the engineering case file and the information disassembly item attribute value; Establishing a knowledge graph of oil and gas field surface engineering cases based on the information disassembly items, wherein each entity in the knowledge graph corresponds to an information disassembly item; Integrate the stored content and the knowledge graph to build an oil and gas field surface engineering case library based on the knowledge graph.
2. The case library construction method according to claim 1, It is characterized in that The preprocessing of the oil and gas field surface engineering case file to obtain basic information of the engineering case file includes: Using optical character recognition technology to identify the engineering case file to obtain the engineering case text; Keyword recognition is performed on the engineering case text to obtain basic information of the engineering case file.
3. The case library construction method according to claim 1, It is characterized in that The step of classifying the engineering case files according to the basic information and establishing a storage level for the engineering case files includes: According to the basic information of the engineering case file, the engineering case is classified into single engineering cases and typical construction cases at the first level; Classify the single project cases and typical construction cases according to the second level of important processes, special environments, and risky operations; Under the second level classification, the basic information of the engineering case is matched with the existing subclasses under each second level, and the engineering case is classified into the corresponding subclass according to the matching result. If the basic information of the engineering case does not match any subclass, a new subclass is established according to the basic information, and the engineering case is classified into the new subclass; The classification level of the engineering case is used as the storage level of the engineering case.
4. The case library construction method according to claim 1, It is characterized in that The information disassembly items include: case overview, key parameters, operation time, construction drawing files, construction resources, construction measures, HSE measures, process record data and picture images, and the attribute value of the information disassembly item is the content corresponding to the information disassembly item.
5. The case library construction method according to claim 1, It is characterized in that The establishing of the oil and gas field surface engineering case knowledge graph based on the information decomposition items includes: S401, matching the engineering case text with the information decomposition item using a rule-based entity recognition method to obtain a named entity; S402, according to the named entities obtained in step S401, manually determine a limited number of entity relationships to form an entity relationship table; S403. Match the entity relationships in the entity relationship table with the case text to generate templates in the form of five-tuples; S404. Cluster based on the similarity of each template, and calculate the average value of the templates belonging to the same class to obtain a new template; S405. Calculate the similarity between the new template and the entity relationships in the entity relationship table, discard the templates with similarity less than the threshold, and form new entity relationships with templates with similarity greater than or equal to the threshold and add them to the entity relationship table; S406. Repeat steps S403 - S405 until all engineering case texts are processed, and generate an oil and gas field surface engineering case knowledge graph according to the entity relationship table.
6. The case library construction method according to claim 5, wherein, the clustering based on the similarity of each template includes: Convert each five-tuple template <Left + entity 1 + Middle + entity 2 + Right> into a vector form: X=(L,T 1 ,M,T 2 ,R) Among them, T 1 , T 2 is the word vector of entity 1 and entity 2, L is the word vector on the left of entity 1, M is the word vector between entity 1 and entity 2, R is the word vector on the right of entity 2, and the similarity of any two five-tuple templates i and j is expressed as: Among them, the vector of template i is expressed as The vector representation of template j is w 1 、w 2 、w 3 is the weight, and w 2 Greater than w 1 and w 3 .
7. The case library construction method according to claim 5, wherein, Before step S401, the method further includes: Regularize the engineering case text and convert it into a standard structured text, including: S11. Based on regular expression determination, screen and clean punctuation marks to obtain continuous natural language text; S12. Segment the continuous natural language text to obtain a lexical sequence with semantic rationality and integrity; S13. By analyzing the dependency relationship between words in a sentence, capture the syntactic structure information of words and use a tree structure to represent the syntactic structure information of the sentence to obtain a standard structured text.
8. An oil and gas field surface engineering case library construction system based on a knowledge graph, wherein, the system includes: A preprocessing module for preprocessing oil and gas field surface engineering case files to obtain basic information of the engineering case files; A case classification module for classifying the engineering case files according to the basic information, establishing an inbound hierarchy for the engineering case files, and the inbound hierarchy is used to store the engineering case files in the engineering case library according to the inbound hierarchy; An intelligent disassembling module for intelligently disassembling the engineering case files according to information disassembly items to obtain information disassembly item attribute values corresponding to the engineering case files, and forming inbound content including the engineering case files and information disassembly item attribute values; A knowledge graph generation module for establishing an oil and gas field surface engineering case knowledge graph based on the information disassembly items, and each entity in the knowledge graph corresponds to an information disassembly item; A case library generation module for integrating the inbound content and the knowledge graph to construct an oil and gas field surface engineering case library based on the knowledge graph.
9. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the case library construction method according to any one of claims 1 to 7.
10. A computer storage medium, wherein, The storage medium stores computer instructions, and the computer instructions can be executed by a processor to implement the case library construction method according to any one of claims 1 to 7.