Earth science geological entity relation extraction method, medium and equipment

Through the first-stage training of neural network models combined with the second-stage optimization of PPO algorithm, the adaptability and local optimization problems in geological entity relationship extraction are solved, and more efficient model training and more accurate geological entity relationship extraction are achieved.

CN120354909AActive Publication Date: 2025-07-22CHINA UNIV OF GEOSCIENCES (WUHAN)

Patent Information

Application Number
CN202510854944.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The traditional geological entity relationship extraction method relies on manually defining features and lacks adaptability and generalization capabilities, resulting in low accuracy and poor robustness in complex semantic scenarios, and deep learning models lack adaptability during training, making them prone to local optimality.

Method used

The neural network model is used for one-stage training, combined with the PPO algorithm for two-stage optimization. By defining the reward function and the advantage function, limiting the policy gradient update in real time, building a joint loss function, realizing the transformation from supervised learning to reinforcement learning, and dynamically adjusting the model parameters.

Benefits of technology

The adaptability and accuracy of the model are improved, the instability and overfitting risks during the training process are reduced, the efficiency and accuracy of relationship extraction are improved, and the semantic diversity and conceptual drift of different geological scenarios are adapted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354909A_ABST
    Figure CN120354909A_ABST
Patent Text Reader

Abstract

The invention provides a geology entity relationship extraction method, medium and equipment, and relates to the technical field of entity relationship extraction, and the method comprises the steps: dividing geology data, entities and labels into a training set and a test set; constructing a neural network model for geological entity relation extraction, and inputting the training set into the model for first-stage training; performing second-stage training on the model after the first-stage training by using the training set, taking the input of the second-stage training model as a state and the output as action probability distribution, setting a reward function and a dominant function, and constructing a PPO model to optimize the second-stage training model; constructing a joint loss function according to the loss function of the first-stage training and the loss function of the PPO model during the second-stage training; and testing the model after the two-stage training by using a test set, and performing earth scientific geological entity relation extraction by using the tested model. According to the method, the problems of instability and local optimum in training are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of entity relation extraction, and in particular to a method, medium, and device for extracting geological entity relations in earth science. Background Art

[0002] With the rapid development of the geoscience field, a large number of unstructured earth science-related literatures have emerged continuously, which contain rich geological entity information and complex entity relations. Traditional geological entity relation extraction methods usually adopt rule-based or supervised learning methods. The performance of these methods highly depends on manually defined features and rules, lacking adaptability and generalization ability, and often showing problems such as low accuracy and poor generalization in dealing with complex and changeable semantic scenarios. Therefore, researchers have proposed to use deep learning models for entity relation extraction in order to improve the extraction performance and reduce the degree of manual intervention.

[0003] In recent years, deep learning methods represented by pre-trained language models such as BERT and SciBERT have made remarkable progress in entity relation extraction tasks. However, these methods usually adopt a fixed loss function during the training process. Once the parameters are determined in the initial training stage, it is very difficult for the model strategy to be further dynamically adjusted according to the actual extraction effect, resulting in a lack of adaptability in the training process, being easily trapped in local optima, and limiting the accuracy and robustness of relation extraction.

[0004] The Proximal Policy Optimization (PPO) algorithm has been successfully applied in the field of reinforcement learning. It realizes a more stable and efficient policy optimization process by restricting the step size of policy updates and avoiding drastic fluctuations between the old and new policies. However, the traditional PPO algorithm has not been directly applied to entity relation extraction tasks and needs to be further improved to be applicable to complex natural language processing scenarios. Summary of the Invention

[0005] The purpose of the present invention is to propose a method for extracting geological entity relations in earth science to solve the instability and local optimum problems in the training process of neural network models, including the following steps: S1. Obtain earth science geological data, extract entities from it and generate labels, and divide the geological data, entities, and labels into a training set and a test set; S2. Construct a neural network model for geological entity relation extraction, input the training set into the neural network model for the first-stage training, and obtain the model after the first-stage training; S3. Use the training set to perform second-stage training on the model after the first-stage training to obtain a second-stage trained model. Take the input of the second-stage trained model as the state, the output as the action probability distribution, and set the reward function and the advantage function. Construct a Proximal Policy Optimization (PPO) model to optimize the second-stage trained model; S4. Construct a joint loss function based on the loss function of the first-stage training and the loss function of the PPO model during the second-stage training; S5. Use the test set to test the model after the second-stage training, and use the tested model for earth science geological entity relationship extraction.

[0006] Further, the reward function of the PPO model is as follows: , where, denotes the reward function, x represents the input of the second-stage trained model, y represents the true label of the output of the second-stage trained model, denotes the normalized score of the second-stage trained model for y, denotes the normalized score of the second-stage trained model for c, and c represents the set of true labels that are not the output of the second-stage trained model.

[0007] Further, the advantage function of the PPO model is as follows: , , , where, denotes the advantage of the i-th sample relative to the baseline, denotes the reward value of the i-th sample, b represents the baseline function, denotes the decay coefficient, denotes the average batch reward, and N represents the batch size.

[0008] Further, the joint loss function is as follows: , where, denotes the joint loss, denotes the loss of the first-stage training, denotes the loss of the PPO model during the second-stage training, denotes the hyperparameter for adjusting the weight.

[0009] Further, the loss function of the first-stage training is: , where, denotes that the i-th input is the label of the j-th output class, represents the probability that the i-th input is the j-th output category, N is the input batch size, and C is the total number of categories.

[0010] Furthermore, during the two-stage training, the loss function of the Proximal Policy Optimization model is: , , , , , where, represents taking the expectation, represents the policy loss, represents the value function loss, represents the entropy regularization term, and represent the weight coefficients, represents the probability ratio, represents the parameters of the neural network for the two-stage training, represents the advantage function, represents the clipping operation, which limits to the interval inside, represents the value function, represents the reward function, represents the probability ratio of the new policy to select the action in the state , represents the probability ratio of the old policy to select the action in the state and

[0011] The present invention also proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned method for extracting the relationship of geoscience geological entities is implemented.

[0012] The present invention also proposes an electronic device, including a processor and a memory, the processor is interconnected with the memory, wherein, the memory is used to store a computer program, the computer program includes computer-readable instructions, and the processor is configured to call the computer-readable instructions to execute the above-mentioned method for extracting the relationship of geoscience geological entities.

[0013] The beneficial effects brought by the technical solution provided by the present invention are: The present invention first conducts a first-stage training through a neural network model, uses labeled data for supervised learning, and quickly obtains good entity recognition and relationship classification effects; uses the training model of the first stage for the second-stage training, and introduces the PPO algorithm to further optimize the strategy, realizing the transformation from supervised learning to reinforcement learning. By defining a reward function and an advantage function applicable to the relationship extraction task, and leveraging the unique clipping objective mechanism of PPO, the amplitude of the policy gradient update is restricted in real time, avoiding overly drastic parameter changes during the training process and reducing the risk of overfitting. In addition, the adaptive policy optimization based on the reward signal can achieve more efficient model migration between different documents and different geological scenarios, solve the problems of semantic diversity and concept drift that are difficult to handle by traditional methods, and can dynamically adjust the model parameters according to the actual extraction feedback, thereby continuously improving the self-adaptability and accuracy of the model during the extraction process. The present invention combines the initial robustness of supervised learning with the self-adaptability of the reinforcement learning stage, improves the relationship extraction efficiency, effectively reduces the instability during the training process, overcomes the local optimum problem, and better adapts to the intelligent analysis requirements of geological big data in the field of earth science. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flowchart of the method for extracting geological entity relationships in the earth science in the embodiment of the present invention; Figure 2 is a block diagram of an electronic device in an exemplary embodiment of the embodiment of the present invention; Figure 3 is the change of the loss in the first-stage training and the second-stage training using the Bert-base-uncased model as the neural network model in the embodiment of the present invention; Figure 4 is the change of the loss in the first-stage training and the second-stage training using the Bert-base-uncased model as the neural network model in the embodiment of the present invention; Figure 5 is the change of the loss in the first-stage training using the Roberta-large model as the neural network model in the embodiment of the present invention; Figure 6 is the change of the loss in the second-stage training using the Roberta-large model as the neural network model in the embodiment of the present invention; Figure 7 is the change of the loss in the first-stage training using the Albert-large-v1 model as the neural network model in the embodiment of the present invention; Figure 8 is the change of the loss in the second-stage training using the Albert-large-v1 model as the neural network model in the embodiment of the present invention. Detailed implementation manners

[0015] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0016] The flowchart of the method for extracting geological entity relationships in the embodiments of the present invention is as Figure 1 , and specifically includes the following steps: S1. Obtain geological data in earth science, extract entities therefrom and generate entity relationship annotation labels, and divide the geological data, entities and labels into a training set and a test set.

[0017] Entities in the geological data in earth science are accurately extracted through a geological dictionary and the opinions of geological experts, ensuring the professionalism and accuracy of the entities. The geological dictionary such as "Geological Dictionary" is a comprehensive geological dictionary, including five major disciplines of atmospheric science, geography, geology, geophysics and ocean science. Each discipline consists of a general introduction and branch directions. The whole dictionary constructs a discipline knowledge tree according to the knowledge system of the discipline, selects entries from the discipline knowledge tree, determines the level, and arranges the order. The labels are generated based on geological rule sets and expert opinions, ensuring the high quality and reliability of the data labels. The entity relationships in the data set are evenly distributed, ensuring the representativeness of the data.

[0018] Transform the data into a data format suitable for neural networks, laying a foundation for subsequent model training and evaluation. Finally, the data set is divided into training samples and test samples. The training samples are used for model training, and the test samples are used to verify the effect of the model, ensuring the generalization ability of the model.

[0019] S2. Construct a neural network model for extracting geological entity relationships, input the training set into the neural network model for the first-stage training, and perform supervised learning using the labels to obtain the model after the first-stage training.

[0020] The neural network models for extracting geological entity relationships can be: BERT-base-uncased, RoBERTa-large and ALBERT-large-v1. These pre-trained models are all based on the Transformer architecture and have powerful language understanding capabilities. BERT-base-uncased provides deep language representations through bidirectional context learning. RoBERTa-large performs stronger in multiple tasks through improved training strategies. ALBERT-large-v1 reduces computational resource consumption through parameter sharing and matrix factorization while maintaining high efficiency.

[0021] By training these models to adapt to specific task requirements, such as entity relationship extraction, the accuracy and generalization ability of the models are further improved. The hyperparameters of the experiment are set, including the maximum sequence length (128), batch size (8), number of training epochs (3), and the use of the AdamW optimizer and learning rate (2e-5). During the training process, first, the data of each batch is passed into the model for feedforward calculation. The model generates predicted logits based on the input sentences and entities and calculates the loss value. Then, the gradients of each model parameter are calculated through backpropagation, and the weights are updated using the optimizer. The training loss of each epoch is recorded and visualized to track the convergence of the model. After training is completed, the performance of the model is evaluated on the test set, and the accuracy, recall, and F1 score are calculated to verify the effectiveness of the model. The entire process continuously optimizes the model through feedforward and backpropagation, gradually improving the performance of the model in the task of earth science geological entity relationship extraction.

[0022] The loss function for the first-stage training of the neural network model is: , where indicates that the i-th input is the label of the j-th output class, indicates that the i-th input is the probability of the j-th output class, N is the input batch size, and C is the total number of classes.

[0023] S3. Use the training set to perform second-stage training on the model after the first-stage training to obtain the second-stage training model. Take the input of the second-stage training model as the state and the output as the action probability distribution, and set the reward function and advantage function. The policy represents the probability distribution of which actions the second-stage training model takes under a given state (input). Construct a Proximal Policy Optimization (PPO) model to optimize the second-stage training model.

[0024] The construction of the Proximal Policy Optimization (PPO) model includes the following steps: 1. Define the state, action, and reward function.

[0025] Take the input of the second-stage training model as the state of the Proximal Policy Optimization (PPO) model, the output of the second-stage training model as the action probability distribution, and the reward function is as follows: , where represents the reward function. When the value is large, it means that the model's confidence in the correct class is much higher than that in any wrong class; x represents the input of the second-stage training model, y represents the true label of the output of the second-stage training model, represents the normalized score of the second-stage training model for y, Represents the normalized score of the two-stage training model for c, where c represents the set of true labels output by the non-two-stage training model.

[0026] 2. Calculate the advantage function to evaluate the pros and cons of a certain action relative to the average performance in a state (advantage function). The advantage function is as follows: , , , where, represents the advantage of the i-th sample relative to the baseline, represents the reward value of the i-th sample, b represents the baseline function, and the initial value is R; represents the decay coefficient, represents the average batch reward, and N represents the batch size.

[0027] 3. Calculate the probability ratio: , where, represents the probability ratio, represents the parameters of the neural network for two-stage training, represents the probability ratio of the new policy to select action in state , represents the probability ratio of the old policy to select action in state .

[0028] 4. Policy update: If the update is too large (exceeding the clipping range ), it will be penalized to ensure that the update amplitude is moderate, neither too conservative nor too radical. The policy loss is: , where, represents the policy loss, represents taking the expectation, represents the advantage function, represents the clipping operation, which limits to the interval .

[0029] 5. Value function update, optimize the value function with the following loss function: , where, represents the value function loss, represents the value function, represents the reward function.

[0030] 6. The loss function for constructing the proximal policy optimization model during the two-stage training is as follows: , where, represents the loss of the proximal policy optimization model during the two-stage training, and represent the weight coefficients. Repeat the above steps: Through multiple rounds of iteration, gradually optimize the policy until convergence.

[0031] S4. According to the loss function of the first-stage training and the loss function of the proximal policy optimization model during the two-stage training, construct a joint loss function. The joint loss function is as follows: , where, represents the joint loss, represents the loss of the first-stage training, represents the loss of the proximal policy optimization model during the two-stage training, represents the hyperparameter for adjusting the weight.

[0032] S5. Use the test set to test the model after the two-stage training, and use the tested model for earth science geological entity relationship extraction.

[0033] Set different iterative training cycles. After each iterative training cycle is completed, test the model after the two-stage training through test samples, and calculate the loss data; adjust the structural parameters of the network model through the calculation of the loss function.

[0034] In an exemplary embodiment, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor, implements the above-mentioned earth science geological entity relationship extraction method.

[0035] Please refer to Figure 2 , in an exemplary embodiment, there is further provided an electronic device, including at least one processor, at least one memory, and at least one communication bus.

[0036] Wherein, a computer program is stored on the memory, and the computer program includes computer-readable instructions. The processor calls the computer-readable instructions stored in the memory through the communication bus to execute the above-mentioned earth science geological entity relationship extraction method.

[0037] To verify the effectiveness of the method of the present invention, a data set containing more than 15,000 pieces of data is obtained for the extraction of earth science geological entity relationships, and the hyperparameters of the experiment are set, including the maximum sequence length (128), batch size (8), number of training cycles (3), and the use of the AdamW optimizer and learning rate (2e-5 ).

[0038] For the results of the first-stage training and first-stage evaluation using the Bert-base-uncased model as the neural network model, refer to Table 1. For the results of the second-stage training and second-stage evaluation using the Bert-base-uncased model as the neural network model, refer to Table 2. For the changes in the losses of the first-stage training and second-stage training using the Bert-base-uncased model as the neural network model, refer to Figure 3 and Figure 4 .

[0039] For the results of the first-stage training and first-stage evaluation using the Roberta-large model as the neural network model, refer to Table 3. For the results of the second-stage training and second-stage evaluation using the Roberta-large model as the neural network model, refer to Table 4 respectively. For the changes in the losses of the first-stage training and second-stage training using the Roberta-large model as the neural network model, refer to Figure 5 and Figure 6 .

[0040] For the results of the first-stage training and first-stage evaluation using the Albert-large-v1 model as the neural network model, refer to Table 5. For the results of the second-stage training and second-stage evaluation using the Albert-large-v1 model as the neural network model, refer to Table 6 respectively. For the changes in the losses of the first-stage training and second-stage training using the Albert-large-v1 model as the neural network model, refer to Figure 7 and Figure 8 .

[0041] Among them, the number of training rounds in the first stage and the second stage is selected three times, and the loss of each time is calculated. The evaluation metrics in the first stage and the second stage are F1 score, Recall, and Accuracy.

[0042] Table 1

[0043] Table 2

[0044] It can be clearly seen from Table 1 and Table 2 that after the second-stage optimization: the F1 has increased by 1.78%, the Recall has increased by 2.61%, the Accuracy has increased by 2.76%, and the initial loss of the second-stage training is significantly lower than the loss at the end of the first stage (decreased from 0.0653 to 0.0261).

[0045] Table 3

[0046] Table 4

[0047] It can be clearly seen from Table 3 and Table 4 that after the second-stage optimization: F1 increased by 1.28%, Recall increased by 2.29%, Accuracy increased by 0.86%, and the initial loss of the second-stage training was significantly lower than the loss at the end of the first stage (decreased from 0.1297 to 0.0518).

[0048] Table 5

[0049] Table 6

[0050] It can be clearly seen from Table 5 and Table 6 that after the second-stage optimization: F1 increased by 2.21%, Recall increased by 1.32%, Accuracy increased by 2.14%, and the initial loss of the second-stage training was significantly lower than the loss at the end of the first stage (decreased from 0.0820 to 0.0227).

[0051] In summary, (1) The increase in F1 indicates the steady improvement of the overall performance of the model and a more optimized balance between precision and recall; the increase in Recall greatly reduces the missed detections and improves the model's ability to capture key information; the increase in Accuracy indicates a significant enhancement in the credibility of the prediction results and an improvement in the model's generalization ability. (2) The initial loss of the second-stage training is lower than the loss at the end of the first stage, which indicates that the convergence of the model training is better and the parameter optimization is more refined. (3) The stable and significant improvement of the indicators not only proves the effectiveness of continued training but also proves the effectiveness of the parameter strategy, training method, and overfitting prevention strategy adopted in the second stage. (4) The necessity of the second-stage training: When the performance of the first-stage model has approached its limit, continuing the second-stage training usually only brings a small or even stagnant improvement in the indicators; if the training strategy is inappropriate (such as not reasonably setting regularization or adjusting the learning rate), it may even lead to overfitting of the model and a decline in performance. Therefore, the significant performance improvement in the second-stage training of this experiment indicates that the model has not reached its bottleneck state and there is still room for further optimization. In addition, this obvious improvement also clearly verifies the effectiveness of the parameter optimization and strategy design adopted in the second-stage training, reflecting the good generalization ability of the model. In other words, the second-stage training is not just about increasing the number of training rounds but further exploring the data potential through more refined parameter adjustment and optimization strategies (such as learning rate decay, regularization, etc.), more comprehensively demonstrating the performance advantages and optimization space of the model, indicating that the second-stage training is very necessary and effective in this experiment.

[0052] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for extracting relationships of geological entities in geoscience, characterized in that, It includes the following steps: S1. Obtain geoscience geological data, extract entities from it and generate labels, and divide the geological data, entities, and labels into a training set and a test set; S2. Construct a neural network model for geological entity relationship extraction, input the training set into the neural network model for the first-stage training, and obtain the model after the first-stage training; S3. Use the training set to perform the second-stage training on the model after the first-stage training to obtain the second-stage training model. Use the input of the second-stage training model as the state and the output as the action probability distribution, and set the reward function and the advantage function. Construct a proximal policy optimization model to optimize the second-stage training model; S4. Construct a joint loss function according to the loss function of the first-stage training and the loss function of the proximal policy optimization model during the second-stage training; S5. Use the test set to test the model after the second-stage training, and use the tested model to perform geoscience geological entity relationship extraction.

2. The method for extracting geological entity relationships in geoscience according to claim 1, wherein The reward function of the proximal policy optimization model is as follows: , Among them, represents the reward function, x represents the input of the two-stage training model, and y represents the true label output by the two-stage training model. represents the normalized score of the two-stage training model for y. represents the normalized score of the two-stage training model for c, where c represents the set of true labels output by the non-two-stage training model.

3. A method for extracting relationships of geological entities in geoscience, according to claim 1, characterized in that The advantage function of the proximal policy optimization model is as follows: , , , wherein, represents the advantage of the i-th sample relative to the baseline, represents the reward value of the i-th sample, b represents the baseline function, represents the decay coefficient, represents the average batch reward, and N represents the batch size.

4. A method for extracting relationships of geological entities in geoscience, according to claim 1, wherein The joint loss function is as follows: , Among them, represents the combined loss, represents the loss of the first-stage training, represents the loss of the proximal policy optimization model during the second-stage training, represents the hyperparameter for adjusting the weight.

5. The method for extracting the relationship of geological entities in geoscience according to claim 4, characterized in that, The loss function of the first-stage training is: , wherein, indicates that the i-th input is the label of the j-th output category, indicates that the i-th input is the probability of the j-th output category, N is the input batch size, and C is the total number of categories.

6. The method for extracting the relationship of geological entities in geoscience according to claim 4, wherein The loss function of the proximal policy optimization model during the second-stage training is: , , , , , Among them, represents the expectation, represents the policy loss, represents the value function loss, represents the entropy regularization term, and represents the weight coefficient, represents the probability ratio, represents the parameters of the neural network for two-stage training, represents the advantage function, represents the clipping operation that restricts within the interval inside, represents the value function, represents the reward function, represents the probability ratio of the new policy to select action at state , represents the probability ratio of the old policy to select action at state 。 7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the method according to any one of claims 1-6.

8. An electronic device, characterized in that, It includes a processor and a memory, the processor is connected to the memory. Among them, the memory is used to store a computer program, the computer program includes computer-readable instructions, and the processor is configured to call the computer-readable instructions to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method for automatically generating submission demand abstract based on strategy gradient algorithm

    CN111291175A

  • One-to-many energy supplement method based on deep reinforcement learning in wireless sensor network

    CN115314943A

  • Fine-grained Chinese geological entity spatial relationship joint extraction method and system

    CN119474248A

  • Auditing domain knowledge graph ontology framework construction method based on large language model

    CN119647580A

  • Hardware and neural architecture co-search

    US20220019880A1

Cited By

  • Biological environment text named entity recognition method, medium, equipment and product

    CN121659944A

  • Method, medium, device and product for named entity recognition of bio-environmental text

    CN121659944B