Oil field exploration and development long text segmentation method combined with large model modeling
Through joint large-scale modeling and cross-attention mechanisms, the problem of low accuracy and generalization in long text segmentation in oilfield exploration and development is solved, and efficient and accurate long text segmentation and irrelevant information filtering are achieved.
Patent Information
- Application Number
- CN202311458499.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
The existing technology is difficult to accurately segment the original long text of oil field exploration and development, mainly because the general text segmentation method lacks the necessary understanding of oil field exploration and development expertise and does not have irrelevant data filtering functions, resulting in low accuracy and generalization.
The joint large modeling method is adopted to obtain the key information of the original text through serialized modeling and semantic correlation, and to capture the closeness between text sentences in combination with the cross attention mechanism, automatically segment the long corpus and filter professionally irrelevant content.
Long text segmentation of oil field exploration and development with low complexity and high algorithm efficiency is realized, adapting to a variety of application scenarios, and improving the accuracy and generalization of segmentation.
Smart Images

Figure CN119940366A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of oilfield data processing, and in particular relates to a long text segmentation method for oilfield exploration and development combined with large model modeling. Background Art
[0002] Oilfield exploration and development involves a lot of professional knowledge, a wide range of areas, and a large scale. Therefore, in order to train a suitable large oil and gas natural language model, it is necessary to first complete the organization of its corpus. It is worth noting that the manual organization process is huge and time-consuming, so it is an inevitable trend to automatically segment the unorganized professional knowledge through efficient analysis and calculation to achieve professional corpus mining.
[0003] Early research works have been done by quantifying the lexical cohesion within text segments (Masao Utiyama et al., 2001; Fernando Llopis et al., 2002). Since it is difficult to define and quantify accurately, lexical cohesion is usually approximated by counting the number of word repetitions. Deep learning serialization modeling provides further room for improvement in text segmentation. Omri Koshorek et al. proposed a method for document segmentation using hierarchical Bi-LSTMs in 2018. To improve generalization, Jing Li et al. introduced a model based on attention mechanism for document segmentation in 2018, and achieved advanced results in chapter segmentation. Yizhong Wang et al. proposed the CRF-BiLSTM method in 2018, using ELMO pre-trained embedding. Charuta Pethe et al. proposed a method based on Bert combined with logical constraints in 2020, and achieved the best results in the novel segmentation task. Michal Lukasik et al. proposed a method based on Transformer combined with multi-attention mechanism in 2020.
[0004] Based on the above methodology, those skilled in the art have also made many attempts. For example, the patent document with the title of patent: A method for detecting outstanding safety problems in oil fields and the application number of CN201910305672.6 records the following technical solution: The present invention relates to a method for detecting outstanding safety problems in oil fields. The method collects a large number of cases of oil field safety problems to establish a corpus; then selects certain texts from the corpus to establish a training sample set, trains the texts in the training sample set, and establishes an oil field safety outstanding problem detection model; uses the oil field safety outstanding problem detection model to predict the outstanding safety problems in the oil field to be tested, calculates the probability values of each topic corresponding to the document to be tested, and selects the topic with the largest probability value as the prediction result of the document to be tested, which is the prediction result of the outstanding safety problems in the oil field to be tested. The detection method uses known data to train the manager prediction model. When using it, it only needs to input the outstanding safety problems in the oil field to be tested into the prediction model. The operation process is simple, and more importantly, the requirements for the staff are low, and the prediction results are less affected by the operator.
[0005] However, after further research, the inventors found that the existing technologies including the above patent documents are difficult to accurately segment the original long text of oilfield exploration and development. The main reason is that the segmentation method of general text lacks the necessary understanding of oilfield exploration and development professional knowledge and has no irrelevant data filtering function, resulting in low accuracy and generalization. Therefore, it is urgent for those skilled in the art to develop a method for segmenting long text of oilfield exploration and development to solve the above technical problems existing in the prior art. Summary of the invention
[0006] The present invention provides a method for segmenting long texts for oilfield exploration and development by using a joint large model. The method obtains key information of the original text through serialization modeling and semantic relevance; then, the closeness between text sentences is captured by combining the cross-attention mechanism with the expressive power of the general joint large model. On the basis of semantic atomization, the long corpus is automatically segmented and professional irrelevant content is automatically filtered, so it has low complexity and high algorithm efficiency, and can meet the needs of professional long text segmentation work in various application scenarios.
[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions: A long text segmentation method for oilfield exploration and development combined with large model modeling includes the following steps: Step 1: Process the long text corpus to obtain a sample set; Step 2: In the process of general knowledge modeling, use the sentence vectors extracted from the natural language big model as the understanding features of the general language; Step 3: In the process of professional knowledge modeling, the network structure is built based on the standard Transformer structure; and semantic relevance operations are added to the standard Transformer structure to further mine the association relationship in the sample set text after QKV feature projection; Step 4: Use the cross-attention mechanism to interact with features, mine key information, and extract enhanced features to provide technical support for text segmentation and irrelevant information filtering; Step 5: Design corresponding loss functions for text segmentation and irrelevant information filtering tasks, and perform model training; The training model obtained after training can be used to segment long texts for oil field exploration and development after deployment.
[0008] Preferably, before step 1, the method further comprises the following steps: Step 0: Filter the original data; Filter out long text corpora containing professional knowledge in oilfield exploration and development.
[0009] Preferably, the step 1 can be specifically described as: Step 11: Segment the long text corpus; mark the short texts formed after segmentation with the corresponding positions in the original document; Step 12: Analyze the short text obtained in step 11, distinguish the three parts of meaningful professional concepts, independent sentences, and invalid expressions, and generate a sample set.
[0010] Preferably, the step 11 can be specifically described as: Step 111: Using a text segmentation model, segment the long text corpus into a number of short texts; Step 112: manually correcting the short text formed after segmentation; Step 113: Using a string matching algorithm, mark and match the short texts formed after segmentation with corresponding positions in the original document.
[0011] Preferably, the specific representation form of the general knowledge modeling in step 2 is: Fea llm =LLM hidden (C) (1); In the formula, LLM hidden Represents the hidden layer output of the natural language model, Fea llm Represents the final hidden layer features.
[0012] Preferably, the specific representation form of the professional knowledge modeling in step 3 is: Fea up-q =D_Conv5×1 (Fea q ) (2) Fea up-k =D_Conv 5×1 (Fea k ) (3) Fea up-v =D_Conv 5×1 (Fea v ) (4) Where D_Conv represents the convolution operation based on row vectors, and the width used is 5; Fea q 、Fea k 、Fea v They represent the QKV projection features of the sample set text features; Fea up-q 、Fea up-k 、Fea up-v They represent the features after semantic relevance operation.
[0013] Preferably, the specific representation of text segmentation in step 4 is: Fea hidden =MHCA(W q (Fea llm ),W k (Fea long ),W v (Fea long )) (5) Where MHCA represents the cross attention mechanism network; W q , W k , W v Represents feature projection; Fea hidden Represents the final long text feature.
[0014] Preferably, in step 5, corresponding loss functions are designed for the text segmentation and irrelevant information filtering tasks, and the specific expressions are: L seg =-Plog(P′)-(1-P)log(1-P′) (6) L filter =-ylog(y)-(1-y)log(1-y′) (7) In the formula, P represents the true value, and P' represents the predicted value; when P = 1, it means that the current position needs to be separated; when P = 0, it means that the current position does not meet the segmentation conditions; y represents the true value, and y' represents the predicted value; when y=1, it means that the current position needs to be separated; when y=0, it means that the current position does not meet the segmentation conditions; The loss functions described in equations (6) and (7) are combined, which is specifically expressed as: L loss =αL seg +βL filter (8) Among them, L loss Represents the overall loss value of text segmentation and irrelevant information filtering, α and β represent the weights of the loss caused by text segmentation and the loss caused by irrelevant information filtering in modeling.
[0015] The invention provides a long text segmentation method for oilfield exploration and development by joint large model modeling. The long text segmentation method for oilfield exploration and development by joint large model modeling comprises the following steps: step 1: processing a long text corpus to obtain a sample set; step 2: in a general knowledge modeling process, using a natural language large model to take sentence vectors extracted therefrom as understanding features of the general language; step 3: in a professional knowledge modeling process, building a network structure based on a standard Transformer structure; and adding a semantic relevance operation to the standard Transformer structure, so as to further mine the association relationship in the sample set text after QKV feature projection; step 4: performing feature interaction through a cross-attention mechanism, mining out key information, and extracting enhanced features, so as to provide technical support for text segmentation and irrelevant information filtering; step 5: designing corresponding loss functions for text segmentation and irrelevant information filtering tasks, and performing model training.
[0016] A long text segmentation method for oilfield exploration and development combined with a large model modeling having the above-mentioned step characteristics has at least the following technical advantages over the prior art: (1) The oilfield exploration and development long text segmentation method provided by the present invention combines multiple technologies such as natural language large model, exploration and development professional knowledge modeling, and cross-attention mechanism, so that the method can adapt to various scenarios.
[0017] (2) The long text segmentation method for oilfield exploration and development provided by the present invention, based on the standard Transformer structure serialization modeling, combines semantic relevance operations to more accurately model the long text of oilfield exploration and development. The above algorithm can work more effectively in different scenarios.
[0018] (3) The method for segmenting long texts for oilfield exploration and development based on the joint large-scale modeling provided by the present invention can ensure the completion of the segmentation of long texts for oilfield exploration and development while also having the characteristics of low complexity and high algorithm efficiency. Therefore, it is particularly suitable for the segmentation of long texts for oilfield exploration and development in actual working scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0020] In the following drawings: Figure 1 A schematic diagram of the flow of the long text segmentation method for oil field exploration and development combined with large model modeling provided by the present invention; Figure 2 An example image after long text is segmented into short text for oil field exploration and development. DETAILED DESCRIPTION
[0021] The present invention provides a method for segmenting long texts for oilfield exploration and development by using a joint large model. The method obtains key information of the original text through serialization modeling and semantic relevance; then, the closeness between text sentences is captured by combining the cross-attention mechanism with the expressive power of the general joint large model. On the basis of semantic atomization, the long corpus is automatically segmented and professional irrelevant content is automatically filtered, so it has low complexity and high algorithm efficiency, and can meet the needs of professional long text segmentation work in various application scenarios.
[0022] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Embodiment 1
[0023] The present invention provides a long text segmentation method for oil field exploration and development combined with large model modeling, such as Figure 1 As shown, the following steps are included: Step 1: Process the long text corpus to obtain a sample set.
[0024] As a more preferred implementation, step 1 can be further specifically described as: Step 11: Segment the long text corpus; mark the short texts formed after segmentation with the corresponding positions in the original document.
[0025] Step 12: Analyze the short text obtained in step 11, distinguish the three parts of meaningful professional concepts, independent sentences, and invalid expressions, and generate a sample set.
[0026] The purpose of analyzing the short texts in step 12 is to distinguish each short text into three parts: meaningful professional concepts, independent sentences, and invalid expressions. Among them, the meaningful professional concepts and (usually) independent sentences in the short text are meaningful contents for the present invention; while the invalid expressions are meaningless contents that can be ignored. Specifically, Figure 2 As shown, Figure 2 An example of segmenting a long text of oilfield exploration and development into short texts is shown.
[0027] After completing step 1, proceed to step 2.
[0028] Step 2: In the process of general knowledge modeling, use the natural language big model to extract the sentence vectors as the understanding features of the general language.
[0029] It is worth noting that the reason why the natural language big model is used in the general knowledge modeling process is that the natural language big model, as a general knowledge dictionary, contains a large amount of general knowledge corpus in its training corpus, and has good understanding and expression capabilities. Therefore, when used as a feature extractor to obtain text key information and context information, gradient back propagation may not be performed during training. Among them, the more preferred specific representation of general knowledge modeling in step 2 is: Fea llm =LLM hidden (C) (1); In the formula, LLM hidden Represents the hidden layer output of the natural language model, Fea llm Represents the final hidden layer features.
[0030] After completing step 2, proceed to step 3.
[0031] Step 3: In the process of professional knowledge modeling, the network structure is built based on the standard Transformer structure; and semantic relevance operations are added to the standard Transformer structure to further explore the association relationship in the sample set text after QKV feature projection.
[0032] As a preferred embodiment of the present invention, the specific representation form of the professional knowledge modeling in step 3 is: Fea up-q =D_Conv 5×1 (Fea q ) (2) Fea up-k =D_Conv 5×1 (Fea k ) (3) Fea up-v=D_Conv 5×1 (Fea v ) (4) Where D_Conv represents the convolution operation based on row vectors, and the width used is 5; Fea q 、Fea k 、Fea v They represent the QKV projection features of the sample set text features; Fea up-q 、Fea up-k 、Fea up-v They represent the features after semantic relevance operation.
[0033] After completing step 3, proceed to step 4.
[0034] Step 4: Use the cross-attention mechanism to interact with features, mine key information, and extract enhanced features to provide technical support for text segmentation and irrelevant information filtering.
[0035] It is worth noting that the large natural language model, as a general knowledge dictionary, has the ability to express and understand general knowledge, while the modeling built by professional knowledge requires professional knowledge and the ability to learn long texts. Taking this as an opportunity, the cross-attention mechanism is used to interact with features in text segmentation, to mine key information, and ultimately provide technical support for text segmentation and irrelevant information filtering.
[0036] Among them, as a more preferred implementation of the present invention, the specific representation form of text segmentation in step 4 is: Fea hidden =MHCA(W q (Fea llm ),W k (Fea long ),W v (Fea long )) (5) Where MHCA represents the cross attention mechanism network; W q , W k , W v Represents feature projection; Fea hidden Represents the final long text feature.
[0037] After completing step 4, proceed to step 5.
[0038] Step 5: Design corresponding loss functions for text segmentation and irrelevant information filtering tasks, and perform model training; The training model obtained after training can be used to segment long texts for oil field exploration and development after deployment.
[0039] It should be noted that in the process of segmenting the long text of oilfield exploration and development expertise, the two tasks of text segmentation and irrelevant information filtering use multi-task recognition algorithms, so the cross entropy classification algorithm is used to judge the above two tasks: Specifically, as a preferred implementation of the present invention, in step 5, corresponding loss functions are designed for the text segmentation and irrelevant information filtering tasks, and the specific expressions are: L seg =-Plog(P′)-(1-P)log(1-P′) (6) L filter =-ylog(y)-(1-y)log(1-y′) (7) In the formula, P represents the true value, and P' represents the predicted value; when P = 1, it means that the current position needs to be separated; when P = 0, it means that the current position does not meet the segmentation conditions; y represents the true value, and y' represents the predicted value; when y=1, it means that the current position needs to be separated; when y=0, it means that the current position does not meet the segmentation conditions; The loss functions described in equations (6) and (7) are combined, which is specifically expressed as: L loss =αL seg +βL filter (8) Among them, L loss Represents the overall loss value of text segmentation and irrelevant information filtering, α and β represent the weights of the loss caused by text segmentation and the loss caused by irrelevant information filtering in modeling.
[0040] At this point, the entire process of obtaining the training model after training is completed. Based on this training model (deployment), the long text segmentation work for oilfield exploration and development can be carried out. Embodiment 2
[0041] Embodiment 2 includes all the technical features of Embodiment 1. In addition, Embodiment 2 is further limited to the following technical features: Among them, as a more preferred embodiment of the present invention, step 11 can be further specifically described as: Step 111: Using a text segmentation model, segment the long text corpus into a number of short texts; Step 112: manually correcting the short text formed after segmentation; Step 113: Using a string matching algorithm, mark and match the short texts formed after segmentation with corresponding positions in the original document.
[0042] It is worth noting that through the above-mentioned processing steps, a semi-automatic processing method is used to divide the long text corpus into several short texts, thereby realizing the preprocessing of the long text corpus. Embodiment 3
[0043] Embodiment 3 includes all the technical features of Embodiment 1. In addition, Embodiment 3 is further limited to the following technical features: The present invention provides a method for segmenting long text in oil field exploration and development by combining large model modeling. Figure 1 As shown, before step 1, the following steps are also included: Step 0: Filter the original data; Filter out long text corpora containing professional knowledge in oilfield exploration and development.
[0044] It is worth noting that step 0 is used to perform a preliminary screening of the original data to determine whether there is a long text corpus containing professional knowledge in the original data, so as to provide assistance for subsequent steps. I will not go into details here.
[0045] The invention provides a long text segmentation method for oilfield exploration and development by joint large model modeling. The long text segmentation method for oilfield exploration and development by joint large model modeling comprises the following steps: step 1: processing a long text corpus to obtain a sample set; step 2: in a general knowledge modeling process, using a natural language large model to take sentence vectors extracted therefrom as understanding features of the general language; step 3: in a professional knowledge modeling process, building a network structure based on a standard Transformer structure; and adding a semantic relevance operation to the standard Transformer structure, so as to further mine the association relationship in the sample set text after QKV feature projection; step 4: performing feature interaction through a cross-attention mechanism, mining out key information, and extracting enhanced features, so as to provide technical support for text segmentation and irrelevant information filtering; step 5: designing corresponding loss functions for text segmentation and irrelevant information filtering tasks, and performing model training.
[0046] A long text segmentation method for oilfield exploration and development combined with a large model modeling having the above-mentioned step characteristics has at least the following technical advantages over the prior art: (1) The oilfield exploration and development long text segmentation method provided by the present invention combines multiple technologies such as natural language large model, exploration and development professional knowledge modeling, and cross-attention mechanism, so that the method can adapt to various scenarios.
[0047] (2) The long text segmentation method for oilfield exploration and development provided by the present invention, based on the standard Transformer structure serialization modeling, combines semantic relevance operations to more accurately model the long text of oilfield exploration and development. The above algorithm can work more effectively in different scenarios.
[0048] (3) The method for segmenting long texts for oilfield exploration and development based on the joint large-scale modeling provided by the present invention can ensure the completion of the segmentation of long texts for oilfield exploration and development while also having the characteristics of low complexity and high algorithm efficiency. Therefore, it is particularly suitable for the segmentation of long texts for oilfield exploration and development in actual working scenarios.
[0049] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A long text segmentation method for oilfield exploration and development combined with large model modeling, characterized in that: The steps include: Step 1: Process the long text corpus to obtain a sample set; Step 2: In the process of general knowledge modeling, use the sentence vectors extracted from the natural language big model as the understanding features of the general language; Step 3: In the process of professional knowledge modeling, the network structure is built based on the standard Transformer structure; and semantic relevance operations are added to the standard Transformer structure to further mine the association relationship in the sample set text after QKV feature projection; Step 4: Use the cross-attention mechanism to interact with features, mine key information, and extract enhanced features to provide technical support for text segmentation and irrelevant information filtering; Step 5: Design corresponding loss functions for text segmentation and irrelevant information filtering tasks, and perform model training. The trained model obtained after training can be deployed to segment long texts for oilfield exploration and development.
2. According to the method for segmenting long texts in oilfield exploration and development by combining large model modeling as described in claim 1, it is characterized in that: Before step 1, the following steps are also included: Step 0: Filter the original data; Filter out long text corpora containing professional knowledge in oilfield exploration and development.
3. The method for segmenting long texts for oilfield exploration and development by combining large model building according to claim 1, characterized in that: The step 1 can be specifically described as: Step 11: Segment the long text corpus; mark the short texts formed after segmentation with the corresponding positions in the original document; Step 12: Analyze the short text obtained in step 11, distinguish the three parts of meaningful professional concepts, independent sentences, and invalid expressions, and generate a sample set.
4. The method for segmenting long texts for oilfield exploration and development by combining large model building according to claim 1, characterized in that: The step 11 can be specifically described as: Step 111: Using a text segmentation model, segment the long text corpus into a number of short texts; Step 112: manually correcting the short text formed after segmentation; Step 113: Using a string matching algorithm, mark and match the short texts formed after segmentation with corresponding positions in the original document.
5. The method for segmenting long texts in oilfield exploration and development by combining large model building according to claim 1, characterized in that: The specific representation of general knowledge modeling in step 2 is: Fea llm =LLM hidden (C) (1); In the formula, LLM hidden Represents the hidden layer output of the natural language model, Fea llm Represents the final hidden layer features.
6. The method for segmenting long texts in oilfield exploration and development by combining large model building according to claim 1, characterized in that: The specific representation of professional knowledge modeling in step 3 is: Fea up-q =D_Conv 5×1 (Fea q ) (2) Fea up-k =D_Conv 5×1 (Fea k ) (3) Fea up-v =D_Conv 5×1 (Fea v ) (4) Where D_Conv represents the convolution operation based on row vectors, and the width used is 5; Fea q 、Fea k 、Fea v They represent the QKV projection features of the sample set text features; Fea up-q 、Fea up-k 、Fea up-v They represent the features after semantic relevance operation.
7. The method for segmenting long text in oilfield exploration and development by combining large model building according to claim 1, characterized in that: The specific representation of text segmentation in step 4 is: Whoa hidden =MHCA(W q (What llm ),W k (What long ),W v (What long )) (5) Where MHCA represents the cross attention mechanism network; W q , W k , W v Represents feature projection; Fea hidden Represents the final long text feature.
8. The method for segmenting long texts in oilfield exploration and development by combining large model building according to claim 1, characterized in that: In step 5, corresponding loss functions are designed for the text segmentation and irrelevant information filtering tasks, and the specific expressions are: THE seg =-Plog(P′)-(1-P)log(1-P′) (6) L filter =-ylog(y)-(1-y)log(1-y′) (7) In the formula, P represents the true value, and P' represents the predicted value; when P = 1, it means that the current position needs to be separated; when P = 0, it means that the current position does not meet the segmentation conditions; y represents the true value, and y' represents the predicted value; when y=1, it means that the current position needs to be separated; when y=0, it means that the current position does not meet the segmentation conditions; The loss functions described in equations (6) and (7) are combined, which is specifically expressed as: L loss =αL seg +βL filter (8) Among them, L loss Represents the overall loss value of text segmentation and irrelevant information filtering, α and β represent the weights of the loss caused by text segmentation and the loss caused by irrelevant information filtering in modeling.
Citation Information
Patent Citations
Oil field safety outburst problem detection method
CN110046664A