Construction method of intelligent auxiliary analysis system for corrosion failure of petroleum pipe

By building an intelligent auxiliary analysis system for corrosion failure of petroleum pipes, and using deep learning technology to automatically analyze text data of corrosion failure of petroleum pipes, the problem of relying on expert experience in the existing technology is solved, and the accuracy and efficiency of analysis are improved.

CN119940057APending Publication Date: 2025-05-06CHINA NAT PETROLEUM CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311443698.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology relies on expert experience in petroleum pipe failure analysis, which makes it difficult for people who have just engaged in failure analysis to accurately analyze the causes of failure, affecting the formulation of failure prevention measures.

Method used

By constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes, the candidate relationship recognition layer, word vector conversion layer, bidirectional LSTM layer, CRF layer, loss compensation layer and model output layer are used to automatically identify and analyze the text data of corrosion failure of petroleum pipes to generate failure conclusions.

Benefits of technology

It improves the accuracy and efficiency of corrosion failure analysis of petroleum pipes, reduces the dependence on manual experience, and can handle the oil pipe corrosion failure analysis text more quickly and accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940057A_ABST
    Figure CN119940057A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method of an intelligent auxiliary analysis system for corrosion failure of a petroleum pipe. Sequentially setting a candidate relation recognition layer, a word vector conversion layer, a bidirectional LSTM layer, a CRF layer, a loss compensation layer and a model output layer; inputting the petroleum pipe corrosion failure analysis text training data set into a candidate relation identification layer, comparing with a plurality of preset entity relation types, and determining entities with relations; entity vectors obtained through conversion of entities with the relation are input into a bidirectional LSTM layer for training, trained samples are input into a CRF layer, and the conditional probability of a labeling sequence is determined; inputting the labeling sequence of which the conditional probability meets the condition into a loss compensation layer; outputting the compensated labeling sequence through a model output layer; and constructing the petroleum pipe corrosion failure intelligent auxiliary analysis system by taking the entities in the output sequence as nodes and taking the entity relationship in the output sequence as edges. Petroleum pipe failure analysis can be intelligently assisted, and the accuracy of failure reason analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of failure analysis, and in particular to a method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes. Background Art

[0002] Oil and gas field pipes such as oil well pipes, surface pipelines, and fracturing manifolds are used in large quantities and have high costs. Once a failure accident occurs, it will not only cause the oil and gas wells to stop transmission and production, causing huge economic losses, but also cause huge environmental pollution and casualties due to oil and gas leakage caused by corrosion perforation. At present, the basis of failure control of oil pipes is failure analysis, that is, through failure analysis, the cause of failure is clarified, and suggestions are put forward to avoid failure accidents. There are many reasons for the failure of oil pipes, and corrosion-related causes account for more than 70%. Failure analysis is a multidisciplinary discipline. In addition to relying on various analysis methods, the conclusion of the failure cause depends largely on the experience of the person conducting the failure analysis, that is, expert experience. However, for those who have just engaged in failure analysis, due to the limitations of professional knowledge, on-site conditions and failure analysis experience, the failure conclusions drawn are often incomplete or wrong, which brings great difficulties to the formulation of failure prevention measures. Summary of the invention

[0003] In order to solve the above problems, the inventors have made the present invention, and provide a method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum tubular goods through specific implementation methods.

[0004] In a first aspect, an embodiment of the present invention provides a method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum tubular goods, comprising the following steps:

[0005] Set up the candidate relationship recognition layer, word vector conversion layer, bidirectional LSTM layer, CRF layer, loss compensation layer and model output layer in sequence;

[0006] The text in the oil pipe corrosion failure analysis text training data set is input into the candidate relationship recognition layer, and compared with the preset multiple entity relationship types to determine the entities with relationships;

[0007] Input the entities with relationships into the word vector conversion layer to obtain the corresponding entity vectors;

[0008] The entity vector is input into the bidirectional LSTM layer for training, and the trained samples are input into the CRF layer to determine the conditional probability of the labeled sequence;

[0009] Input the labeled sequence whose conditional probability meets the preset conditions into the loss compensation layer for compensation;

[0010] Output the compensated annotation sequence through the model output layer;

[0011] Taking the entities in the output sequence as nodes and the entity relationships in the output sequence as edges, an intelligent auxiliary analysis system for corrosion failure of oil pipes is constructed.

[0012] Specifically, establishing a text training dataset for oil pipe corrosion failure analysis includes the following steps:

[0013] Collect oil pipe corrosion failure analysis literature and reports to build a text dataset;

[0014] According to entity type, the professional vocabulary extracted from the text dataset is classified to obtain a professional corpus dataset;

[0015] Combined with the entity types in the professional corpus dataset, the text in the text dataset is annotated to obtain the training dataset.

[0016] Specifically, setting the entity relationship type includes the following steps:

[0017] According to the corrosive medium, material, corrosion morphology and corrosion product XRD, a variety of entity relationship types are set, including carbon dioxide corrosion entity relationship, oxygen corrosion entity relationship, hydrogen sulfide stress corrosion cracking relationship, chloride stress corrosion cracking relationship, under-scale corrosion relationship and galvanic corrosion relationship.

[0018] Specifically, the multiple entity relationship types include at least one of the following entity relationship types:

[0019] In the carbon dioxide corrosion entity relationship, the corrosive medium is CO2 and produced fluid, the material is carbon steel or alloy steel, the corrosion morphology is ulcer-like or pit-like, and the corrosion product XRD corresponds to FeCO3;

[0020] In the oxygen corrosion entity relationship, the corrosive medium is O2 and produced fluid, the material is carbon steel, alloy steel or stainless steel, the corrosion morphology is elliptical or continuous or yellow-brown, and the corrosion product XRD corresponds to Fe3O4, Fe2O3, and FeOOH;

[0021] In the hydrogen sulfide stress corrosion cracking relationship, the corrosive medium is H2S and produced fluid, the material is carbon steel or alloy steel or stainless steel, the corrosion morphology is cracks, and the corrosion product XRD corresponds to FeS;

[0022] In the chloride stress corrosion cracking relationship, the corrosive medium is Cl-, the material is stainless steel, the corrosion morphology is cracks, and the corrosion product XRD corresponds to FeCl3;

[0023] In the under-scale corrosion relationship, the corrosive medium is high-mineralization produced fluid, the material is carbon steel, alloy steel or stainless steel, the corrosion morphology is scale layer or continuous corrosion pits, and the corrosion product XRD corresponds to BaSO4 and CaCO3.

[0024] Specifically, the text in the oil pipe corrosion failure analysis text training data set is input into the candidate relationship recognition layer, and compared with multiple preset entity relationship types to determine the entities with relationships, including the following steps:

[0025] The annotated text in the oil pipe corrosion failure analysis text training dataset is input into the candidate relationship recognition layer, and the sentences are extracted according to the period through the semantic extractor.

[0026] Generate a sequence by sliding the window, combine entities in each sequence in pairs to obtain a first candidate relationship, and filter out entity combinations in the first candidate relationship that are not included in the preset entity relationship type by using a filter to obtain a second candidate relationship;

[0027] Enter the two entity vectors corresponding to each entity combination in the second candidate relationship into the following formula:

[0028]

[0029] Among them, λ and θ are two entity vectors corresponding to the entity combination, N represents the number of entity vectors in the sequence, t is the starting value of the entity vector, i and j represent the position numbers of the entity vectors λ and θ in the sequence respectively;

[0030] When Compare(λ|θ) is greater than 1, it is determined that the two entities corresponding to λ and θ have a relationship. When Compare(λ|θ) is less than 1, it is determined that the two entities corresponding to λ and θ do not have a relationship.

[0031] Specifically, the entity vector is input into the bidirectional LSTM layer for training, and the trained samples are input into the CRF layer to determine the conditional probability of the labeled sequence, including the following steps:

[0032] The entity vector is extracted through one-dimensional convolution, pooling, residual network, and mask mechanism to extract data features, and input into the bidirectional LSTM layer for training. The trained samples are input into the CRF layer, and the following formula is used to predict the conditional probability that the labeled sequence is y when the input observation sequence is x:

[0033]

[0034] Among them, P(y|x) represents the conditional probability that the labeled sequence is y when the input observation sequence is x, n represents the number of observation sequences, and x i represents the observation sequence value, y l represents the value of the label sequence, μ i For x i The corresponding conditioning factor, μ l for y l The corresponding conditioning factor, w l represents the weight of the labeled sequence, r iis the weight of the observation sequence, i and l represent the positions of x and y in the sequence respectively, and k is the similarity coefficient.

[0035] Specifically, the labeled sequence whose conditional probability meets the preset conditions is input into the loss compensation layer for compensation, which includes the following steps:

[0036] The labeled sequences with conditional probabilities greater than 90% are input into the loss compensation layer to perform conditional compensation on the sequences with few samples in the labeled sequences. The compensation model is an entropy balance compensation model based on the residual sequence:

[0037]

[0038] Among them, y represents the labeled sequence, λ is the residual compensation factor, N represents the number of sequences, the subscript i of y and λ represents the sequence position number corresponding to y and λ, and e represents a natural constant.

[0039] In a second aspect, an embodiment of the present invention provides an intelligent auxiliary analysis system for corrosion failure of petroleum pipes, comprising:

[0040] The scene input module is used to input the on-site situation text report into the candidate relationship recognition layer, compare it with the preset multiple entity relationship types, and determine the entities with relationships; input the entities with relationships into the word vector conversion layer to obtain the corresponding entity vector; input the entity vector into the trained bidirectional LSTM layer, and input the sequence output by the bidirectional LSTM layer into the CRF layer to determine the conditional probability of the labeled sequence; input the labeled sequence whose conditional probability meets the preset conditions into the loss compensation layer for compensation;

[0041] The failure conclusion output module is used to output the compensated annotation sequence through the model output layer to obtain the corrosion failure conclusion of the oil pipe.

[0042] Based on the same inventive concept, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the aforementioned method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes is implemented.

[0043] Based on the same inventive concept, an embodiment of the present invention provides a computer storage medium, wherein the computer storage medium stores computer executable instructions, and when the computer executable instructions are executed, the aforementioned method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes is implemented.

[0044] The beneficial effects of the above technical solution provided by the embodiment of the present invention include at least:

[0045] It is possible to build an intelligent auxiliary analysis system for oil pipe corrosion failure, intelligently assist oil pipe failure analysis, improve the text processing speed of oil pipe corrosion failure analysis, improve the accuracy of failure cause analysis, and reduce the dependence of failure analysis on manual work.

[0046] Other features and advantages of the present invention will be described in the following description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.

[0047] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0049] Figure 1 This is a flow chart of a method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum tubular goods in an embodiment of the present invention;

[0050] Figure 2 Schematic diagram of the candidate relationship identification layer to the CRF layer in an embodiment of the present invention;

[0051] Figure 3 The figure is a schematic diagram of the structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0053] In order to solve the problems existing in the prior art, an embodiment of the present invention provides a method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes.

[0054] The embodiment of the present invention provides a method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes, the process of which is as follows: Figure 1 As shown, the following steps are included:

[0055] Step S1: Set the candidate relationship recognition layer, word vector conversion layer, bidirectional LSTM layer, CRF layer, loss compensation layer and model output layer in sequence.

[0056] Step S2: Input the text in the oil pipe corrosion failure analysis text training data set into the candidate relationship recognition layer, compare it with the preset multiple entity relationship types, and determine the entities with existing relationships.

[0057] Establishing a text training dataset for oil pipe corrosion failure analysis includes the following steps:

[0058] Collect literature and reports on oil pipe corrosion failure analysis and establish a text data set. Specifically, collect publicly published literature and existing report materials in the field of oil pipe corrosion failure analysis, establish an oil pipe corrosion failure text data set, and build a distributed storage platform to classify and store the text. For example, the original text is obtained from relevant paper searches, network searches, and oil-related website searches such as "Analysis of the Causes of Submarine Oil Pipeline Fractures", "Corrosion Failure Analysis of Metal Pipelines", "Corrosion Failure Analysis of a Certain Oil Pipeline in Changqing Oilfield", "Analysis of the Causes of Failure of Oil Gathering Pipelines", and "Failure Analysis of Metal Pipelines in High Temperature Environments". At the same time, all text data is stored in the database for distributed storage.

[0059] According to the entity type, the professional vocabulary extracted from the text data set is classified to obtain a professional corpus data set. Specifically, the professional vocabulary in the text data of oil pipe corrosion failure is extracted and classified. The professional vocabulary involved is divided into 9 entity types, including: sample name, corrosion medium, corrosion morphology, corrosion type, corrosion location, material, service time, corrosion product spectrum, and corrosion product XRD. Then, by writing code to call the python toolkit and combining the Jieba word segmentation tool, text data preprocessing work is carried out, including cleaning, noise reduction, and useless information removal.

[0060] Combined with the entity types in the professional corpus dataset, the text in the text dataset is annotated to obtain the training dataset. Specifically, referring to the NLPCC-KQBA public dataset format, the BIO annotation strategy is adopted to complete the entity annotation of the pre-processed oil pipe corrosion failure text. The annotated text entities are distributed and stored in the form of .ann files to form a training dataset for the construction of the intelligent auxiliary analysis system for the cause of oil pipe corrosion. Taking the entity category "failed sample" as an example, the beginning part of the entity related to the failed sample is represented by "B-EXAM", the remaining entity content is annotated as "I-EXAM", and the remaining components in the sentence are all annotated as "O". The oil pipe experimental dataset is cited as an example to illustrate the process of manual annotation of the named entity recognition dataset. The sample question content is "The main reason for the fracture is the unreasonable combination of the pumping rod" and is annotated according to the NLPCC-KQBA public dataset format standard.

[0061] Setting the entity relationship type includes the following steps:

[0062] According to the corrosive medium, material, corrosion morphology and corrosion product XRD, a variety of entity relationship types are set, including carbon dioxide corrosion entity relationship, oxygen corrosion entity relationship, hydrogen sulfide stress corrosion cracking relationship, chloride stress corrosion cracking relationship, under-scale corrosion relationship and galvanic corrosion relationship.

[0063] The multiple entity relationship types include at least one of the following entity relationship types:

[0064] In the carbon dioxide corrosion entity relationship, the corrosive medium is CO2 and produced fluid, the material is carbon steel or alloy steel, the corrosion morphology is ulcer-like or pit-like, and the corrosion product XRD corresponds to FeCO3;

[0065] In the oxygen corrosion entity relationship, the corrosive medium is O2 and produced fluid, the material is carbon steel, alloy steel or stainless steel, the corrosion morphology is elliptical or continuous or yellow-brown, and the corrosion product XRD corresponds to Fe3O4, Fe2O3, and FeOOH;

[0066] In the hydrogen sulfide stress corrosion cracking relationship, the corrosive medium is H2S and produced fluid, the material is carbon steel or alloy steel or stainless steel, the corrosion morphology is cracks, and the corrosion product XRD corresponds to FeS;

[0067] In the chloride stress corrosion cracking relationship, the corrosive medium is Cl-, the material is stainless steel, the corrosion morphology is cracks, and the corrosion product XRD corresponds to FeCl3;

[0068] In the under-scale corrosion relationship, the corrosive medium is high-mineralization produced fluid, the material is carbon steel, alloy steel or stainless steel, the corrosion morphology is scale layer or continuous corrosion pits, and the corrosion product XRD corresponds to BaSO4 and CaCO3.

[0069] The format of the above entity relationship type is:

[0070] The entity relationship of CO2 corrosion is: corrosive medium (CO2 + produced fluid) × material (carbon steel, alloy steel) [produced] corrosion morphology (ulcer-like, pit-like) [generated] corrosion product XRD (FeCO3);

[0071] The entity relationship of oxygen corrosion is: corrosive medium (O2 + produced liquid) × material (carbon steel, alloy steel, stainless steel) [produced] corrosion morphology (elliptical, continuous flake, yellow-brown) [generated] corrosion product XRD (Fe3O4, Fe2O3, FeOOH);

[0072] The relationship between hydrogen sulfide stress corrosion cracking is: corrosive medium (H2S + produced fluid) × material (carbon steel, alloy steel, stainless steel) [produced] corrosion morphology (cracks) [generated] corrosion product XRD (FeS);

[0073] The relationship between chloride stress corrosion cracking is: corrosive medium (Cl-) × material (stainless steel) [produced] corrosion morphology (cracks) [generated] corrosion product XRD (FeCl3);

[0074] The relationship between under-scale corrosion is: corrosive medium (high-mineralization produced fluid) × material (carbon steel, alloy steel, stainless steel) [produces] corrosion morphology (scale layer, continuous lamellar corrosion pits) [generates] corrosion product XRD (BaSO4, CaCO3).

[0075] The text in the oil pipe corrosion failure analysis text training data set is input into the candidate relationship recognition layer, and compared with multiple preset entity relationship types to determine the entities with relationships, including the following steps:

[0076] The annotated text in the oil pipe corrosion failure analysis text training data set is input into the candidate relationship recognition layer, and the sentences are extracted according to the period through the semantic extractor. Specifically, the annotated text entity .ann file is input into the first layer of the neural network model, and the semantic extractor is used to divide each paragraph into sentences according to the period. The format of the divided sentences is Where C represents the text, λ and θ are the fields in the sequence, and their subscripts represent the positions in the sequence.

[0077] A sequence is generated by sliding a window, and the entities in each sequence are combined in pairs to obtain a first candidate relationship. The entity combinations in the first candidate relationship that are not included in the preset entity relationship type are filtered out by a filter to obtain a second candidate relationship. Specifically, sentences are generated according to a sliding window with a window size of 2 and a stride of 1. The corresponding entities in each sentence are combined in pairs to form candidate relationship 1, and then the undefined relationships are filtered out by a filter. After filtering, candidate relationship 2 is formed and compared with the above entity relationship type. The relationship comparison algorithm model is Among them, λ and θ are two entity vectors in the candidate relationship, N represents the number of entity vectors in the text sequence, t is the starting value of the entity vector, and i and j represent the positions of the entity vectors λ and θ in the sequence respectively. If there is such a relationship, it is set to 1, otherwise it is set to 0.

[0078] Enter the two entity vectors corresponding to each entity combination in the second candidate relationship into the following formula:

[0079]

[0080] Among them, λ and θ are two entity vectors corresponding to the entity combination, N represents the number of entity vectors in the sequence, t is the starting value of the entity vector, i and j represent the position numbers of the entity vectors λ and θ in the sequence respectively;

[0081] When Compare(λ|θ) is greater than 1, it is determined that the two entities corresponding to λ and θ have a relationship. When Compare(λ|θ) is less than 1, it is determined that the two entities corresponding to λ and θ do not have a relationship.

[0082] Step S3: Input the entities with relationships into the word vector conversion layer to obtain the corresponding entity vectors. Specifically, the selected entities with relationships are input into the word vector conversion layer, and the labeled relationship entities are converted into vectors based on the improved word2vec algorithm.

[0083] Step S4: Input the entity vector into the bidirectional LSTM layer for training, and input the trained samples into the CRF (Conditional Random Field) layer to determine the conditional probability of the labeled sequence.

[0084] like Figure 2 As shown, the entity vector is input into the bidirectional LSTM layer for training, and the trained samples are input into the CRF layer to determine the conditional probability of the labeled sequence, including the following steps:

[0085] The entity vector is extracted through one-dimensional convolution, pooling, residual network, and mask mechanism to extract data features, and input into the bidirectional LSTM layer for training. The trained samples are input into the CRF layer, and the following formula is used to predict the conditional probability that the labeled sequence is y when the input observation sequence is x:

[0086]

[0087] Among them, P(y|x) represents the conditional probability that the labeled sequence is y when the input observation sequence is x, n represents the number of observation sequences, and x i represents the observation sequence value, y l represents the value of the label sequence, μ i For x i The corresponding conditioning factor, μ l for y l The corresponding conditioning factor, w l represents the weight of the labeled sequence, r i is the weight of the observation sequence, i and l represent the positions of x and y in the sequence respectively, and k is the similarity coefficient.

[0088] Figure 2In the figure, BOS (Basic object storage) represents the basic entity stored, Mask represents the mask mechanism, V represents the basic entity stored, for example, y1, y2 and other vectors V(y1), V(y2) generated by the mask mechanism, S of the CRF layer represents the start vector value (Start value) after being processed by the bidirectional LSTM layer, E represents the emission score (Emission score), B represents the starting position of the entity vector (Beginning of a noun phrase chunk), I represents the middle position of the entity vector (Inside of a noun phrase chunk), and O represents a non-entity (Outside of a noun phrase chunk).

[0089] The above-mentioned bidirectional LSTM layer (BiLSTM) analyzes the embedding of each character received and predicts the probability of each character being in a specific entity category. Using bidirectional LSTM, the bidirectional network can capture both positive and negative information, making the use of text information more comprehensive and effective. A linear layer is added after the bidirectional LSTM layer to project the hidden layer output results generated by the bidirectional LSTM layer to an interval with the meaning of the specific entity category label feature defined by the expression. The learning rate lr of the bidirectional LSTM layer is 1e-3, the number of epoch iterations is 100, the word vector dimension is 3000, and the hidden vector dimension is 256.

[0090] The CRF layer takes the state scores of the bidirectional LSTM layer as input and outputs the maximum possible predicted annotation sequence that meets the annotation transfer constraints. The maximum possible annotation sequence probability prediction model is an optimized Markov conditional vector field. The sequence output by the bidirectional LSTM layer satisfies the Markov distribution. According to the structural characteristics of the pipe corrosion failure entity, the conditional probability when the input observation sequence X takes the value x and the output annotation sequence Y takes the value y is calculated.

[0091] Step S5: Input the labeled sequence whose conditional probability meets the preset conditions into the loss compensation layer for compensation; output the compensated labeled sequence through the model output layer; and construct an intelligent auxiliary analysis system for corrosion failure of oil pipes with entities in the output sequence as nodes and entity relationships in the output sequence as edges.

[0092] The labeled sequence whose conditional probability meets the preset conditions is input into the loss compensation layer for compensation, which includes the following steps:

[0093] The labeled sequences with conditional probabilities greater than 90% are input into the loss compensation layer to perform conditional compensation on the sequences with few samples in the labeled sequences. The compensation model is an entropy balance compensation model based on the residual sequence:

[0094]

[0095] Among them, y represents the labeled sequence, λ is the residual compensation factor, N represents the number of sequences, the subscript i of y and λ represents the sequence position number corresponding to y and λ, and e represents a natural constant.

[0096] The sequences with a probability greater than 90% after training in the CRF layer enter the loss compensation layer. The purpose is to conditionally compensate for the sequences with few samples in the labeled sequence to prevent uneven sample classification.

[0097] After the above entity recognition and relationship recognition, the nodes and edges of the auxiliary system are determined, and the intelligent auxiliary analysis system framework of the pipe corrosion cause can be constructed with entities as nodes and relationships as edges. Based on the data set formed by manual annotation, the knowledge graph is constructed and displayed using the graph database; data is obtained by calling the interface in the MySQL database, and nodes and relationships are constructed through the graph database using Cypher statements.

[0098] In the above method of this embodiment, an intelligent auxiliary analysis system for oil pipe corrosion failure can be constructed to intelligently assist oil pipe failure analysis, improve the text processing speed of oil pipe corrosion failure analysis, improve the accuracy of failure cause analysis, and reduce the dependence of failure analysis on manual work.

[0099] Those skilled in the art can change the above sequence without departing from the protection scope of the present disclosure.

[0100] Another embodiment of the present invention provides an intelligent auxiliary analysis system for corrosion failure of petroleum pipes, comprising:

[0101] The scene input module is used to input the on-site situation text report into the candidate relationship recognition layer, compare it with the preset multiple entity relationship types, and determine the entities with relationships; input the entities with relationships into the word vector conversion layer to obtain the corresponding entity vector; input the entity vector into the trained bidirectional LSTM layer, and input the sequence output by the bidirectional LSTM layer into the CRF layer to determine the conditional probability of the labeled sequence; input the labeled sequence whose conditional probability meets the preset conditions into the loss compensation layer for compensation;

[0102] The failure conclusion output module is used to output the compensated annotation sequence through the model output layer to obtain the corrosion failure conclusion of the oil pipe.

[0103] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0104] In this embodiment, it is possible to intelligently assist in the failure analysis of oil pipes, improve the text processing speed of the oil pipe corrosion failure analysis, improve the accuracy of the failure cause analysis, and reduce the dependence of the failure analysis on manual work.

[0105] Based on the same inventive concept, an embodiment of the present invention provides an electronic device, whose structure is as follows Figure 3 As shown, it includes: a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the aforementioned method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes is implemented.

[0106] Based on the same inventive concept, an embodiment of the present invention provides a computer storage medium, wherein the computer storage medium stores computer executable instructions, and when the computer executable instructions are executed, the aforementioned method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes is implemented.

[0107] Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention shall still fall within the patent scope of the present invention.

Claims

1. A method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes, characterized in that: The following steps are involved: Set up the candidate relationship recognition layer, word vector conversion layer, bidirectional LSTM layer, CRF layer, loss compensation layer and model output layer in sequence; The text in the oil pipe corrosion failure analysis text training data set is input into the candidate relationship recognition layer, and compared with the preset multiple entity relationship types to determine the entities with relationships; Input the entities with relationships into the word vector conversion layer to obtain the corresponding entity vectors; The entity vector is input into the bidirectional LSTM layer for training, and the trained samples are input into the CRF layer to determine the conditional probability of the labeled sequence; Input the labeled sequence whose conditional probability meets the preset conditions into the loss compensation layer for compensation; Output the compensated annotation sequence through the model output layer; Taking the entities in the output sequence as nodes and the entity relationships in the output sequence as edges, an intelligent auxiliary analysis system for corrosion failure of oil pipes is constructed.

2. The method according to claim 1, characterized in that Establishing a text training dataset for oil pipe corrosion failure analysis includes the following steps: Collect oil pipe corrosion failure analysis literature and reports to build a text dataset; According to entity type, the professional vocabulary extracted from the text dataset is classified to obtain a professional corpus dataset; Combined with the entity types in the professional corpus dataset, the text in the text dataset is annotated to obtain the training dataset.

3. The method according to claim 1, characterized in that Setting the entity relationship type includes the following steps: According to the corrosive medium, material, corrosion morphology and corrosion product XRD, a variety of entity relationship types are set, including carbon dioxide corrosion entity relationship, oxygen corrosion entity relationship, hydrogen sulfide stress corrosion cracking relationship, chloride stress corrosion cracking relationship, under-scale corrosion relationship and galvanic corrosion relationship.

4. The method according to claim 1, characterized in that The multiple entity relationship types include at least one of the following entity relationship types: In the carbon dioxide corrosion entity relationship, the corrosive medium is CO2 and produced fluid, the material is carbon steel or alloy steel, the corrosion morphology is ulcer-like or pit-like, and the corrosion product XRD corresponds to FeCO3; In the oxygen corrosion entity relationship, the corrosive medium is O2 and produced fluid, the material is carbon steel, alloy steel or stainless steel, the corrosion morphology is elliptical or continuous or yellow-brown, and the corrosion product XRD corresponds to Fe3O4, Fe2O3, and FeOOH; In the hydrogen sulfide stress corrosion cracking relationship, the corrosive medium is H2S and produced fluid, the material is carbon steel or alloy steel or stainless steel, the corrosion morphology is cracks, and the corrosion product XRD corresponds to FeS; In the chloride stress corrosion cracking relationship, the corrosive medium is Cl-, the material is stainless steel, the corrosion morphology is cracks, and the corrosion product XRD corresponds to FeCl3; In the under-scale corrosion relationship, the corrosive medium is high-mineralization produced fluid, the material is carbon steel, alloy steel or stainless steel, the corrosion morphology is scale layer or continuous corrosion pits, and the corrosion product XRD corresponds to BaSO4 and CaCO3.

5. The method according to claim 1, characterized in that The text in the oil pipe corrosion failure analysis text training data set is input into the candidate relationship recognition layer, and compared with multiple preset entity relationship types to determine the entities with relationships, including the following steps: The annotated text in the oil pipe corrosion failure analysis text training dataset is input into the candidate relationship recognition layer, and the sentences are extracted according to the period through the semantic extractor. Generate a sequence by sliding the window, combine entities in each sequence in pairs to obtain a first candidate relationship, and filter out entity combinations in the first candidate relationship that are not included in the preset entity relationship type by using a filter to obtain a second candidate relationship; Enter the two entity vectors corresponding to each entity combination in the second candidate relationship into the following formula: Among them, λ and θ are two entity vectors corresponding to the entity combination, N represents the number of entity vectors in the sequence, t is the starting value of the entity vector, i and j represent the position numbers of the entity vectors λ and θ in the sequence respectively; When Compare(λ|θ) is greater than 1, it is determined that the two entities corresponding to λ and θ have a relationship. When Compare(λ|θ) is less than 1, it is determined that the two entities corresponding to λ and θ do not have a relationship.

6. The method according to claim 1, characterized in that The entity vector is input into the bidirectional LSTM layer for training, and the trained samples are input into the CRF layer to determine the conditional probability of the labeled sequence, including the following steps: The entity vector is extracted through one-dimensional convolution, pooling, residual network, and mask mechanism to extract data features, and input into the bidirectional LSTM layer for training. The trained samples are input into the CRF layer, and the following formula is used to predict the conditional probability that the labeled sequence is y when the input observation sequence is x: Among them, P(y|x) represents the conditional probability that the labeled sequence is y when the input observation sequence is x, n represents the number of observation sequences, and x i represents the observation sequence value, y l represents the value of the label sequence, μ i For x i The corresponding conditioning factor, μ l for y l The corresponding conditioning factor, w l represents the weight of the labeled sequence, r i is the weight of the observation sequence, i and l represent the positions of x and y in the sequence respectively, and k is the similarity coefficient.

7. The method according to claim 1, characterized in that The labeled sequence whose conditional probability meets the preset conditions is input into the loss compensation layer for compensation, which includes the following steps: The labeled sequences with conditional probabilities greater than 90% are input into the loss compensation layer to perform conditional compensation on the sequences with few samples in the labeled sequences. The compensation model is an entropy balance compensation model based on the residual sequence: Among them, y represents the labeled sequence, λ is the residual compensation factor, N represents the number of sequences, the subscript i of y and λ represents the sequence position number corresponding to y and λ, and e represents a natural constant.

8. An intelligent auxiliary analysis system for corrosion failure of petroleum pipes, characterized in that: include: The scene input module is used to input the on-site situation text report into the candidate relationship recognition layer, compare it with the preset multiple entity relationship types, and determine the entities with relationships; input the entities with relationships into the word vector conversion layer to obtain the corresponding entity vector; input the entity vector into the trained bidirectional LSTM layer, and input the sequence output by the bidirectional LSTM layer into the CRF layer to determine the conditional probability of the labeled sequence; input the labeled sequence whose conditional probability meets the preset conditions into the loss compensation layer for compensation; The failure conclusion output module is used to output the compensated annotation sequence through the model output layer to obtain the corrosion failure conclusion of the oil pipe.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, the method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes as claimed in any one of claims 1 to 7 is implemented.

10. A computer storage medium, characterized in that: The computer storage medium stores computer executable instructions, which, when executed, implement the method for constructing an intelligent auxiliary analysis system for corrosion failure of petroleum pipes as claimed in any one of claims 1 to 7.