Keyword guidance and large language model near-end strategy optimization combined key sentence extraction method
By constructing keyword-key sentence pairs, using joint matching models to evaluate correlation and introduce proximity strategy optimization of KL divergence and state value networks, the problems of strong dependence on labeled data and high training costs in the existing technology are solved, and the effect of key sentence extraction is significantly improved.
Patent Information
- Application Number
- CN202510238007.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-08-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing key sentence extraction methods are highly dependent on labeled data and are costly to train large language models, making it difficult to effectively extract the core content of scientific and technological literature.
Construct keyword-key sentence pairs, use joint matching models to evaluate the correlation and generate reward values, introduce KL divergence to measure the difference between the training model and the reference model, combine the state value network to optimize the near-end strategy, and guide the extraction of key sentences through reinforcement learning.
The effect of key sentence extraction was significantly improved, especially in the scientific and technological literature data sets, the ROUGE-L index was increased by 1.92 percentage points, solving the problem of strong dependence on labeled data and high training cost of large language models, and improving the accuracy and stability of key sentence extraction.
Smart Images

Figure CN120509402A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a key sentence extraction method that combines keyword guidance with proximal strategy optimization of a large language model. Background Art
[0002] Key sentence extraction is a crucial task in natural language processing. Its goal is to extract the most representative sentences from a document or text to summarize the core content. Traditional key sentence extraction methods typically rely on features such as word frequency and sentence position for analysis. These methods often rely on large amounts of annotated data, making them highly dependent on this data.
[0003] The above background information does not necessarily constitute prior art. Summary of the Invention
[0004] The purpose of this application is to provide a key sentence extraction method that combines keyword guidance with proximal strategy optimization of a large language model.
[0005] The technical solution of this application is achieved as follows:
[0006] In a first aspect, embodiments of the present application provide a key sentence extraction method that combines keyword guidance with proximal strategy optimization of a large language model, including:
[0007] Construct keyword-key sentence pairs;
[0008] Use the joint matching model to evaluate their relevance and generate a reward value;
[0009] KL divergence is introduced to measure the difference between the training model and the reference model, and the value score of the current state is estimated by combining the state value network;
[0010] The extraction of key sentences is achieved by optimizing the guidance model through proximal strategy.
[0011] In some embodiments, constructing a keyword-key sentence pair includes:
[0012] Each key sentence text after preprocessing is represented as K = {k 1, k2,...k n}, each keyword is represented by L = {l1, l2, ..., l m}, where k n Indicates the nth character of the key sentence, l m Indicates the mth character of the keyword;
[0013] The pre-trained Chinese BERT model is used to map text into feature vectors, generating vectorized representations of key sentences and keywords respectively.
[0014] Calculate the similarity score between each keyword and key sentence through cosine similarity;
[0015] For each keyword, the key sentence with the highest similarity score that is not repeated with other keywords is selected to form a keyword-key sentence matching pair.
[0016] In some embodiments, evaluating the relevance using a joint matching model to generate a reward value includes:
[0017] Input the generated keyword-key sentence matching pairs into the CEDR-DRMM model and calculate the matching score of each pair;
[0018] All matching scores are normalized and the final reward score is generated by weighted average.
[0019] In some embodiments, introducing KL divergence to measure the difference between the training model and the reference model, and estimating the value score of the current state in combination with the state value network, includes:
[0020] The expected cumulative return in a state is predicted by estimating the value function V(S) of the state. Through optimization training, the state value network minimizes the difference between the value estimate of its output and the actual return.
[0021] The state value network provides a value V(S) for each state s, which represents the cumulative rewards that may be obtained in the future under this state;
[0022] In the learning process, the value function of the current state and the next state is estimated, that is, V(S t ) and V(S t+1 ), and the actual reward R t+1 , calculate the temporal difference error; the temporal difference error is used to update the policy network;
[0023] In reinforcement learning, from a certain state S t Start taking action a t The cumulative return after is defined as
[0024] G t =R t+1 +γR t+2 +γ 2 R t+3 +…
[0025] Among them, G t is the cumulative return starting from time t, y is the discount factor;
[0026] State value function V(S t ) indicates that in state S tThe expected return under the current strategy is defined as
[0027] V(S t )=E[G t |S t ];
[0028] Advantage function δ t Measured in state S t Next select action a t The difference between the actual return and the state value function, δ t =G t -V(S t ).
[0029] In some embodiments, the advantage function δ t It is usually estimated by the time series difference method, and the time series difference error is
[0030] δ t =R t+1 +γV(S t+1 )-V(S t ),
[0031] Among them, R t+1 is the reward value at time t+1, V(S t ) and V(S t+1 ) are the current states S t and the next state S t+1 The value function of .
[0032] In some embodiments,
[0033] The formula of the loss function is
[0034]
[0035] in, is the importance sampling ratio, indicating that action a is executed under the current strategy t The ratio of the probability of executing the action to the probability of executing the action under the old strategy; π θ (a t |S t ) is the probability of training the policy, is the policy probability at the previous time step; is the advantage estimate, which represents the advantage of the current state-action pair, that is, the gap between the actual return and the estimated return; ε is the clipping parameter, which controls the magnitude of the policy change. When r t When (θ) exceeds the range [1-ε,1+ε], it is clamped to the boundary value.
[0036] In some embodiments, KL divergence loss is used to control the difference in strategy between the training strategy and the reference model; the smaller the KL divergence, the more similar the strategy of the training model is to the strategy of the reference model. The formula is:
[0037]
[0038] State value loss is used to optimize the state value function V θ (S t ), which is used to measure the expected value of future rewards obtained by using the strategy in a certain state; the goal of the state value loss is to make the state value function predicted by the model close to the expected return, and the formula is
[0039]
[0040] Among them, V θ (S t ) is the state value function, which means that in state S t Under the current strategy π θ , the expected reward from this state; represents the cumulative reward starting from time step t; E t represents the expected operation at time step t;
[0041] The final loss function formula is
[0042]
[0043] Among them, c1 and c2 are hyperparameters used to balance the weights of various losses.
[0044] In a second aspect, an embodiment of the present application provides a key sentence extraction device that combines keyword guidance with proximal strategy optimization of a large language model, including:
[0045] Construction module, used to construct keyword-key sentence pairs;
[0046] The relevance evaluation module is used to evaluate the relevance using the joint matching model and generate a reward value;
[0047] The value score estimation module is used to introduce KL divergence to measure the difference between the training model and the reference model, and estimate the value score of the current state in combination with the state value network;
[0048] The extraction module is used to extract key sentences by optimizing the guidance model through proximal strategy.
[0049] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory; when the processor executes the running program stored in the memory, the key sentence extraction method combining keyword guidance and large language model proximal strategy optimization as described in any embodiment of the present application is implemented.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the key sentence extraction method combining keyword guidance and large language model proximal strategy optimization as described in any embodiment of the present application is implemented.
[0051] The embodiment of the present application provides a key sentence extraction method that combines keyword guidance with proximal strategy optimization of a large language model, constructs keyword-key sentence pairs, uses a joint matching model to evaluate their relevance, generates a reward value, introduces KL divergence to measure the difference between the training model and the reference model, and estimates the value score of the current state in combination with the state value network. The proximal strategy is optimized to guide the model to achieve key sentence extraction, which largely solves the problems of strong dependence on labeled data and high training cost of large language models. It performs well in key sentence extraction tasks, has a high ROUGE-L index on scientific literature datasets, and significantly improves the effect of key sentence extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A flowchart of a key sentence extraction method combining keyword guidance with proximal strategy optimization of a large language model is shown in an embodiment of the present application.
[0053] Figure 2 A flowchart of using a joint matching model to evaluate relevance and generate a reward value according to one embodiment of the present application is shown.
[0054] Figure 3 A KG-PPO architecture diagram of an embodiment of the present application is shown.
[0055] Figure 4 A DRMM architecture diagram of an embodiment of the present application is shown.
[0056] Figure 5 An embodiment of the present application provides a structural block diagram of a key sentence extraction device that combines keyword guidance with proximal strategy optimization of a large language model.
[0057] Figure 6 A structural block diagram of an electronic device provided by an embodiment of the present application is shown.
[0058] Figure 7 A schematic diagram of a computer-readable storage medium provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0059] In order to further illustrate the features and technical content of the embodiments of this application, the following will further illustrate the technical solutions of this application in conjunction with the accompanying drawings and specific embodiments. The drawings are for reference only and are not intended to limit the scope of the embodiments of this application.
[0060] Unless otherwise defined, all technical and scientific terms used in the examples of this application shall have the same meaning as commonly understood by those skilled in the art. The use of terms is for the purpose of describing the examples of this application only and is not intended to limit this application in any way.
[0061] In the following description, the “some embodiments” mentioned represent a subset of all possible embodiments, but it should be understood that “some embodiments” can be the same or different subsets of all possible embodiments, and can be combined with each other without conflict. It should also be pointed out that if terms such as “first, second, third” appear in the embodiments of the present application, then these terms are only used to distinguish the differences between similar objects and do not represent a specific sort or order. Therefore, terms such as “first, second, third” can be used interchangeably where permitted, and the embodiments of the present application do not have to be executed strictly in the order described.
[0062] In response to the problems in related technologies of key sentence extraction methods that are highly dependent on annotated data and have high training costs for large language models, the embodiment of the present application proposes a key sentence extraction method KG-PPO that combines keyword guidance with proximal strategy optimization of a large language model. The method first constructs keyword-key sentence pairs, and uses a joint matching model to evaluate their relevance, and then weights them to generate reward values; secondly, the KL divergence is introduced to measure the difference between the training model and the reference model, and the value score of the current state is estimated in combination with the state value network; finally, the model is guided by proximal strategy optimization to achieve accurate extraction of key sentences, which largely solves the problems of strong dependence on annotated data and high training costs for large language models. Experimental results show that the method of the embodiment of the present application performs well in the key sentence extraction task, and the ROUGE-L index on the scientific literature dataset is improved by 1.92 percentage points compared with the comparison model, which significantly improves the effect of key sentence extraction.
[0063] refer to Figure 1 As shown, an embodiment of the present application provides a key sentence extraction method that combines keyword guidance with proximal strategy optimization of a large language model, which may include steps S10-S40:
[0064] S10. Construct keyword-key sentence pairs.
[0065] In some embodiments, constructing a keyword-key sentence pair may include:
[0066] Each key sentence text after preprocessing is represented as K = {k1, k2, ...k n}, each keyword is represented by L = {l1, l2, ..., l m}, where k n Indicates the nth character of the key sentence, l m Indicates the mth character of the keyword;
[0067] The pre-trained Chinese BERT model is used to map text into feature vectors, generating vectorized representations of key sentences and keywords respectively.
[0068] Calculate the similarity score between each keyword and key sentence through cosine similarity;
[0069] For each keyword, the key sentence with the highest similarity score that is not repeated with other keywords is selected to form a keyword-key sentence matching pair.
[0070] S20: Use the joint matching model to evaluate the relevance and generate a reward value.
[0071] refer to Figure 2 As shown, in some embodiments, the use of the joint matching model to evaluate the relevance and generate the reward value may include steps S201-S202:
[0072] S201, input the generated keyword-key sentence matching pairs into the CEDR-DRMM model, and calculate the matching score of each pair;
[0073] S202: Normalize all matching scores and generate a final reward score through weighted average.
[0074] S30. KL divergence is introduced to measure the difference between the training model and the reference model, and the value score of the current state is estimated in combination with the state value network.
[0075] S40. Optimize the guidance model through proximal strategy to achieve key sentence extraction.
[0076] In some embodiments, introducing KL divergence to measure the difference between the training model and the reference model and estimating the value score of the current state in combination with the state value network may include:
[0077] The expected cumulative return in a state is predicted by estimating the value function V(S) of the state. Through optimization training, the state value network minimizes the difference between the value estimate of its output and the actual return.
[0078] The state value network provides a value V(S) for each state s, which represents the cumulative rewards that may be obtained in the future under this state;
[0079] In the learning process, the value function of the current state and the next state is estimated, that is, V(S t ) and V(S t+1 ), and the actual reward R t+1 , calculate the temporal difference error; the temporal difference error is used to update the policy network;
[0080] In reinforcement learning, from a certain state S t Start taking action a t The cumulative return after is defined as
[0081] G t =R t+1 +γR t+2 +γ 2 R t+3 +…
[0082] Among them, G t is the cumulative return starting from time t, γ is the discount factor;
[0083] State value function V(S t ) indicates that in state S t The expected return under the current strategy is defined as
[0084] V(S t )=E[G t |S t ];
[0085] Advantage function δ t Measured in state S t Next select action a t The difference between the actual return and the state value function, δ t =G t -V(S t ).
[0086] In some embodiments, the advantage function δ t It is usually estimated by the time series difference method, and the time series difference error is
[0087] δ t =R t+1 +γV(S t+1 )-V(S t ),
[0088] Among them, R t+1 is the reward value at time t+1, V(S t ) and V(S t+1 ) are the current states S t and the next state S t+1 The value function of .
[0089] In some embodiments,
[0090] The formula of the loss function is
[0091]
[0092] in, is the importance sampling ratio, indicating that action a is executed under the current strategy t The ratio of the probability of executing the action to the probability of executing the action under the old strategy; π θ (a t |S t ) is the probability of training the policy, is the policy probability at the previous time step; is the advantage estimate, which represents the advantage of the current state-action pair, that is, the gap between the actual return and the estimated return; ε is the clipping parameter, which controls the magnitude of the policy change. When r t When (θ) exceeds the range [1-ε,1+ε], it is clamped to the boundary value.
[0093] In some embodiments, KL divergence loss is used to control the difference in strategy between the training strategy and the reference model; the smaller the KL divergence, the more similar the strategy of the training model is to the strategy of the reference model. The formula is:
[0094]
[0095] State value loss is used to optimize the state value function V θ (S t ), which is used to measure the expected value of future rewards obtained by using the strategy in a certain state; the goal of the state value loss is to make the state value function predicted by the model close to the expected return, and the formula is
[0096]
[0097] Among them, V θ (S t ) is the state value function, which means that in state S t Under the current strategy π θ , the expected reward from this state; represents the cumulative reward starting from time step t; E t represents the expected operation at time step t;
[0098] The final loss function formula is
[0099]
[0100] Among them, c1 and c2 are hyperparameters used to balance the weights of various losses.
[0101] The overall architecture of the key sentence extraction method of another embodiment of the present application includes four modules: a keyword matching reward module, a state value evaluation module, a KL divergence constraint module and a strategy optimization module. First, the CEDR-DRMM model is used to calculate the weighted average reward value of the keyword and key sentence pairs to measure the relevance score of the text matching. Then, by calculating the KL divergence between the training model and the reference model, the difference between the model generation strategy and the target strategy is evaluated. In this process, the state value function shares the parameters of the training model to obtain the value score of the current state, thereby enhancing the guiding effect on the model decision. Finally, the strategy update step size is constrained by the strategy optimization module to avoid instability caused by excessive strategy changes. The specific model framework is as follows Figure 3 shown.
[0102] Keyword matching reward module
[0103] The core of this module is to construct keyword-keyword matching pairs, input them into the CEDR-DRMM model to obtain matching scores and generate reward values. It mainly consists of two key parts: similarity calculation and reward score calculation.
[0104] Similarity calculation
[0105] The key sentences extracted from the original text by the training model may contain noise and irrelevant information, and directly using them for similarity calculation will reduce its accuracy. To address this problem, the embodiment of the present application preprocesses the sentences, including word segmentation and removing special characters and stop words, to improve the accuracy of similarity calculation. The specific example of sentence preprocessing is shown in Table 1.
[0106] Table 1 Sentence preprocessing examples
[0107]
[0108] For the same scientific literature text, the preset keywords can represent the core content of the text from different thematic perspectives. To this end, the embodiment of the present application takes the preset keywords as the main guide, selects the sentence with the highest similarity to each keyword as the key sentence of the scientific literature text, and accurately extracts the key information that can reflect the text theme. Specifically, each key sentence text after preprocessing is represented as K = {k 1, k2,...k n}, each keyword is represented by L = {l1, l2, ..., l m}, where k n Indicates the nth character of the key sentence, l m Represents the mth character of the keyword. The pre-trained Chinese BERT model maps the text into feature vectors, generating vector representations of key sentences and keywords respectively. The vector calculation process is shown in formula (1):
[0109] E tok =(BERT(Tokens)) tok (1)
[0110] The vector representation generated by the BERT model contains rich contextual information. Based on this, the similarity score between each keyword and the key sentence is calculated by cosine similarity, as shown in formula (2):
[0111] Score=Sim(E K , E L ) (2)
[0112] Here, Sim refers to the cosine similarity calculation. Sort the similarity scores of all keywords and key sentences in descending order.
[0113] Table 2 Examples of keyword-key sentence matching pairs
[0114]
[0115] For each keyword, we select the key sentence with the highest similarity score that does not overlap with other keywords to form a keyword-key sentence matching pair. Specific examples of keyword-key sentence matching pairs are shown in Table 2.
[0116] Reward Points Calculation
[0117] CEDR-DRMM is a model for information retrieval tasks that aims to improve the relevance matching effect between queries and documents by combining contextual information and word embeddings. Among them, DRMM realizes local interaction between queries and documents through a joint deep framework. The core of the model includes matching histograms, feedforward matching networks, and term gating networks, which are used to effectively extract the relevance of each term between keywords and key sentences. Compared with traditional methods, matching histograms can retain more information, thereby improving matching performance. CEDR-DRMM uses the BERT model to vectorize the input information and calculates the matching score through the DRMM model. Its architecture is as follows Figure 4 shown.
[0118] In this embodiment, the generated keyword-keyword matching pairs are input into the CEDR-DRMM model, and the matching score of each pair is calculated. Subsequently, all matching scores are normalized and a final reward score is generated by weighted average. This reward score is used to guide reinforcement learning during the optimization model training process.
[0119] State value assessment module
[0120] The state value network plays a vital role in reinforcement learning. It estimates the value function V(S) of a state to predict the expected cumulative reward in that state. Through continuous optimization training, the state value network aims to minimize the difference between the value estimate and the actual reward it outputs. Specifically, the state value network provides a value V(S) for each state s, which represents the cumulative reward that may be obtained in the future under that state. In the reinforcement learning process, by combining the value function estimate of the current state and the next state, that is, V(S t ) and V(S t+1 ), and the actual reward R t+1 , the temporal difference error can be calculated. This error is used to update the policy network, so that it gradually tends to the strategy that can bring higher returns. The core of the state value network is to extract state features and evaluate potential benefits. Its performance has a direct impact on policy optimization. Because the state value network and the training network need to process the same input state, the feature extraction process of the two is highly similar. Therefore, in the embodiment of the present application, the state value network and the training network share the same model parameters.
[0121] In reinforcement learning, from a certain state S t Start taking action a t The cumulative return after is defined as shown in formula (3):
[0122] G t =R t+1 +γR t+2 +γ 2 R t+3 +… (3)
[0123] Among them, G t is the cumulative return starting from time t, and γ is the discount factor.
[0124] State value function V(S t ) indicates that in state S t The expected return of the current strategy is defined as follows:
[0125] V(S t )=E[G t |S t ] (4)
[0126] Advantage function δ t Measured in state S t Next select action a t The difference between the actual return and the state value function is shown in formula (5):
[0127] δ t =G t -V(S t) (5)
[0128] In practical applications, the advantage function δ t It is usually estimated by the time difference method. The time difference error is shown in formula (6):
[0129] δ t =R t+1 +γV(S t+1 )-V(S t ) (6)
[0130] Among them, R t+1 is the reward value at time t+1, V(S t ) and V(S t+1 ) are the current states S t and the next state S t+1 The advantage function δ is obtained by calculation. t Guide training models to improve decision-making performance.
[0131] KL divergence constraint module
[0132] In the keyword matching reward module, the embodiment of the present application adopts a preset keyword guidance model to select sentences with higher similarity to keywords from the document text as key sentences for extraction. Although using the preset keyword guidance model to extract sentences with higher similarity to keywords as key sentences of scientific and technological document texts can improve the extraction performance, this method still has some potential problems. First, the preset keywords cannot guarantee that all key information is fully covered, and important content may be omitted in individual scientific and technological document texts. Secondly, there is a certain deviation in the similarity calculation between keywords and sentences. Especially when there are synonyms or context changes, the model may find it difficult to accurately capture the actual importance of the sentence.
[0133] Based on the above considerations, this application embodiment adopts the large-scale multilingual pre-training model Qwen developed by Alibaba Cloud, which has powerful text processing and generation capabilities. -72B is used as a reference model to guide the optimization process of the Qwen-32B training model. By calculating the KL divergence, the distribution difference between the training model and the reference model can be quantified, thereby guiding the model to better fit the distribution of real data and ensuring that the generated results are consistent with the output of the reference model. Using the Qwen-72B model as a reference model not only enhances the stability of the training process, but also supplements the information that may be missing in the learning process of the training model with its huge parameter capacity and rich language knowledge, thereby avoiding the occurrence of overfitting. Through this strategy, while maintaining computational efficiency, the performance of the generated text can be optimized, and the generalization ability and robustness of the final training model can be improved. For the output probability distribution of the Qwen-32B and Qwen-72B models, the calculation of the KL divergence is shown in formula (7):
[0134]
[0135] Among them, P(x) represents the probability distribution of the output result x of the training model Qwen-32B, and Q(x) represents the probability distribution of the output result x of the reference model Qwen-72B.
[0136] Strategy Optimization Module
[0137] The core idea of this module is to improve the stability of training by optimizing by limiting the step size of policy updates. The loss function consists of two parts: the original objective function and the objective function with clipping operation, as shown in formula (8):
[0138]
[0139] in, is the importance sampling ratio, indicating that action a is executed under the current strategy t The ratio of the probability of executing the action to the probability of executing the action under the old strategy. Here π θ (a t |S t ) is the probability of training the policy, is the policy probability at the previous time step. is the advantage estimate, which represents the advantage of the current state-action pair, that is, the gap between the actual return and the estimated return; ε is the clipping parameter, which controls the magnitude of the policy change. When r t When (θ) exceeds the range of [1-ε, 1+ε], it is directly limited to the boundary value to prevent the strategy from being updated too much and ensure the stability of the optimization.
[0140] At the same time, the KL divergence loss is used to control the difference in strategy between the training strategy and the reference model. The smaller the KL divergence, the more similar the strategy of the training model is to the strategy of the reference model, which helps to improve the stability of the training process, as shown in formula (9):
[0141]
[0142] Finally, the state value loss is used to optimize the state value function V θ (S t ), which is used to measure the expected value of future rewards obtained by using the strategy in a certain state. The goal of the state value loss is to make the state value function predicted by the model close to the expected return, as shown in formula (10):
[0143]
[0144] Among them, V θ (S t ) is the state value function, which means that in state S t Under the current strategy π θ , the expected reward from this state; represents the cumulative reward starting from time step t; E t represents the desired operation at time step t.
[0145] All loss functions are obtained from this. The ultimate goal is to optimize the policy and value function simultaneously and control the stability of the policy through KL divergence. The final loss function is shown in formula (11):
[0146]
[0147] Among them, c1 and c2 are hyperparameters used to balance the weights of various losses. In the embodiment of the present application, c1 is set to 0.5 and c2 is set to 0.03.
[0148] experiment
[0149] Dataset
[0150] The experiment in the embodiment of this application is based on the Chinese scientific literature dataset of "Chinese Medicinal Materials", which contains 11,400 scientific literature instructions and corresponding 11,400 expert abstracts. In order to ensure the scientificity and objectivity of the model training and evaluation process, the dataset is divided into a training set, a validation set and a test set, of which the training set contains 9,120 articles, the validation set 1,140 articles and the test set 1,140 articles. Each scientific literature text in the dataset consists of three parts: keywords, abstracts and instructions. The sample corpus of scientific literature keyword extraction is shown in Table 3.
[0151] Table 3 Example of scientific literature data
[0152]
[0153]
[0154] The embodiment of the present application uses keyword matching signals as rewards in model training, so it is necessary to construct a keyword-keyword matching data set to provide guidance. Based on 11,400 scientific and technological literature descriptions, 20,000 keywords were extracted therefrom, and a preliminary matching sample of keywords and keyword sentences was constructed. First, the matching degree of keywords and text keyword sentences was calculated using cosine similarity to generate initial keyword-keyword matching pairs and their matching degree scores. Subsequently, after manual correction, the semantic relevance between keywords and keywords was comprehensively analyzed, and all matching keyword sentences and their matching degree scores were revised one by one in combination with the contextual information of the original scientific and technological literature. During the revision process, it was ensured that the matching degree scores were evenly distributed in the interval [0, 1] to optimize the credibility and distribution characteristics of the data samples. Ultimately, a high-quality keyword-keyword matching data set was constructed, as shown in Table 4.
[0155] Baseline Method
[0156] To verify the effectiveness of the KG-PPO model in the key sentence extraction task, this application example selected a variety of existing classic key sentence extraction methods for comparison, covering traditional methods, deep learning models, and enhanced methods of large prediction models. The specific comparison methods include:
[0157] (1) Lead4: It is a position-based key sentence extraction method that generates a summary by preferentially selecting the first three sentences at the beginning of the text, assuming that the leading content usually contains the core information of the text.
[0158] (2) TextRank: It is a key sentence extraction method based on graph ranking. It represents sentences as nodes in a graph, constructs edges based on the similarity between sentences, and iteratively calculates the node importance scores to extract the most representative sentences.
[0159] (3) SummaRuNNer: It is a key sentence extraction method based on recurrent neural networks. It captures the contextual relationship between sentences through sequence modeling and scores their importance based on sentence position, content and global information.
[0160] (4) BERTSUM: It is a key sentence extraction method based on the pre-trained model BERT. It adapts to sentence-level summarization tasks by introducing segment embedding and position embedding, and uses the Transformer structure to encode global context information to identify key sentences.
[0161] (5) LLM-aided: Combining a large language model with a semi-supervised learning framework, it effectively utilizes a small amount of labeled data to improve the ability to extract and generate key information in the dialogue summarization task.
[0162] (6) DiffuSum: It is a large-scale generative key sentence extraction method based on the diffusion model. Through a step-by-step denoising process, it uses a pre-trained large model to directly generate the meaning of the required summary sentence, thereby achieving high-quality, context-consistent key sentence extraction.
[0163] Evaluation metrics and experimental design
[0164] Evaluation indicators and experimental environment
[0165] This experiment uses ROUGE, a commonly used evaluation index in the field of key sentence extraction, to evaluate the quality of generated scientific literature abstracts. The ROUGE index quantifies the quality of abstracts by comparing the n-gram overlap between the generated abstract and the reference abstract. The higher the value, the higher the degree of match between the generated abstract and the reference abstract. This embodiment of the application uses ROUGE-1 (R-1) and ROUGE-2 (R-2) to evaluate the quality of generated scientific literature abstracts. The calculation method is shown in formula (12):
[0166]
[0167] Where N is the number of n-grams; RefSum is the reference summary; C(n-gram) is the number of n-grams in the reference summary; C match (n-gram) is the number of n-grams that appear simultaneously in the extracted key sentence and the reference summary.
[0168] ROUGE-L (RL) is an evaluation method based on the longest common subsequence (LCS) that measures the degree of overlap between the generated summary and the longest common subsequence of the reference summary. ROUGE-L is calculated as shown in formula (13):
[0169]
[0170] Among them, R lcs is the recall rate, P lcs is the accuracy, F lcs is the F1 value, i.e., the ROUGE-L value, and β is a hyperparameter used to adjust the weights of recall and precision.
[0171] This experiment was run on a server running Ubuntu 18.04. The specific experimental environment information is shown in Table 5.
[0172] Table 5 Experimental environment
[0173]
[0174] Experimental parameter configuration
[0175] This experiment is based on the PyTorch deep learning framework, the LLama Factory inference deployment engine, and the TRL (Transformers Reinforcement Learning) reinforcement learning optimization library. The model parameter settings are shown in Table 6.
[0176] Table 6 Experimental parameter settings
[0177]
[0178] Experimental results analysis
[0179] Data Analysis
[0180] Table 7 Experimental results of different models on scientific literature dataset (%)
[0181]
[0182] The experimental comparison results of different models on the "Chinese herbal medicine" scientific literature abstract dataset are shown in Table 7.
[0183] By analyzing Experiments 1 and 2, we can see that traditional methods mainly rely on heuristic rules and graphical models, which are difficult to effectively capture the complex semantic information and contextual associations in scientific literature texts, so their ROUGE scores are relatively low.
[0184] In contrast, the results of Experiments 3, 4, and 7 demonstrate that the KG-PPO method further improves the performance of deep learning methods. Although deep learning methods rely on large-scale training data and complex model structures, their generalization capabilities are limited when faced with highly structured information such as scientific literature text. By introducing a multi-objective joint optimization mechanism based on reinforcement learning, KG-PPO can better handle structured text and significantly improve the accuracy and stability of the model. In all ROUGE indicators, KG-PPO scores higher than SummaRuNNer and BERTSUM, fully demonstrating its advantages in the task of scientific literature text extraction.
[0185] Further analysis of the results of Experiments 5, 6, and 7 reveals that the KG-PPO method still has significant advantages over large language model enhancement methods. Although LLM-aided and DiffuSum integrate the powerful semantic understanding capabilities of large language models, their combination with reinforcement learning and multi-objective optimization is relatively weak, and they fail to fully realize the potential of reinforcement learning in the key sentence extraction task. In contrast, KG-PPO optimizes the training process by introducing KL divergence constraints, weighted averaging of reward values, and guidance from the state-value network, achieving fine-grained adjustment of model parameters and significantly improving its ability to optimize key sentence extraction results. This shows that KG-PPO can more effectively capture key sentence information in text through more precise training optimization and multi-objective reinforcement learning mechanisms, and exhibits higher accuracy and stronger generalization capabilities when handling complex text extraction tasks.
[0186] Ablation experiments
[0187] To verify the effectiveness of the keyword matching reward module, state value assessment module, and KL divergence constraint module in the KG-PPO model, an ablation experiment was conducted in this embodiment of the application. The experiment is described as follows:
[0188] Experiment 1: Remove the keyword matching reward module in KG-PPO;
[0189] Experiment 2: Remove the state value evaluation module and use the separate LoRA fine-tuning method to train the Qwen32B model;
[0190] Experiment 3: Remove the reference model to verify the effect of the KL divergence constraint module.
[0191] The experimental results are shown in Table 8.
[0192] Table 8 Ablation experiment results (%)
[0193]
[0194] In Experiment 1, removing the keyword matching reward module resulted in a significant decrease in the ROUGE score, particularly in the ROUGE-L metric. This result demonstrates that the keyword matching reward module, by weighting the reward signal for the degree of matching between keywords and key sentences, provides precise guidance for model training, enabling the model to focus more on extracting key information, thereby significantly improving extraction performance.
[0195] In Experiment 2, after removing the state value assessment module, ROUGE-L's performance dropped by 2.4 percentage points compared to the KG-PPO model. This result highlights the important role of the state value assessment module in improving model performance. Compared to a single LoRA fine-tuning strategy, the model's extraction performance was significantly improved. This demonstrates that integrating reinforcement learning strategies with global optimization mechanisms can more effectively guide the model's learning process, leading to better extraction performance.
[0196] In Experiment 3, removing the KL divergence constraint module significantly degraded model performance. By calculating the difference between the trained model and the reference model, the KL divergence constraint effectively prevents the model from overfitting the training data, enabling it to generate more accurate key sentences that meet the reference standard. Without this constraint, the model may tend to produce more unstable or inaccurate outputs, resulting in overall performance degradation.
[0197] In summary, the KG-PPO model combines keyword matching rewards, state value assessment, and KL divergence constraints to form a multi-objective joint optimization framework, achieving higher accuracy and stability in key sentence extraction tasks. These modules each play a different role, working together to improve the performance of the overall model.
[0198] The key sentence extraction method (KG-PPO) proposed in the embodiment of the present application, which combines keyword guidance and proximal strategy optimization of a large language model, shows excellent performance in the key sentence extraction task and effectively alleviates the problems of data labeling dependence and high cost of large prediction model training.
[0199] refer to Figure 5 As shown, another embodiment of the present application provides a key sentence extraction device that combines keyword guidance with proximal strategy optimization of a large language model, including:
[0200] Construction module, used to construct keyword-key sentence pairs;
[0201] The relevance evaluation module is used to evaluate the relevance using the joint matching model and generate a reward value;
[0202] The value score estimation module is used to introduce KL divergence to measure the difference between the training model and the reference model, and estimate the value score of the current state in combination with the state value network;
[0203] The extraction module is used to extract key sentences by optimizing the guidance model through proximal strategy.
[0204] In some embodiments, constructing a keyword-key sentence pair includes:
[0205] Each key sentence text after preprocessing is represented as K = {k1, k2, ...kn}, each keyword is represented by L = {l1, l2, ..., l m}, where k n Indicates the nth character of the key sentence, l m Indicates the mth character of the keyword;
[0206] The pre-trained Chinese BERT model is used to map text into feature vectors, generating vectorized representations of key sentences and keywords respectively.
[0207] Calculate the similarity score between each keyword and key sentence through cosine similarity;
[0208] For each keyword, the key sentence with the highest similarity score that is not repeated with other keywords is selected to form a keyword-key sentence matching pair.
[0209] In some embodiments, evaluating the relevance using a joint matching model to generate a reward value includes:
[0210] Input the generated keyword-key sentence matching pairs into the CEDR-DRMM model and calculate the matching score of each pair;
[0211] All matching scores are normalized and the final reward score is generated by weighted average.
[0212] In some embodiments, introducing KL divergence to measure the difference between the training model and the reference model, and estimating the value score of the current state in combination with the state value network, includes:
[0213] The expected cumulative return in a state is predicted by estimating the value function V(S) of the state. Through optimization training, the state value network minimizes the difference between the value estimate of its output and the actual return.
[0214] The state value network provides a value V(S) for each state s, which represents the cumulative rewards that may be obtained in the future under this state;
[0215] In the learning process, the value function of the current state and the next state is estimated, that is, V(S t ) and V(S t+1 ), and the actual reward R t+1 , calculate the temporal difference error; the temporal difference error is used to update the policy network;
[0216] In reinforcement learning, from a certain state S t Start taking action a t The cumulative return after is defined as
[0217] G t =R t+1+γR t+2 +γ 2 R t+3 +…
[0218] Among them, G t is the cumulative return starting from time t, γ is the discount factor;
[0219] State value function V(S t ) indicates that in state S t The expected return under the current strategy is defined as
[0220] V(S t )=E[G t |S t ];
[0221] Advantage function δ t Measured in state S t Next select action a t The difference between the actual return and the state value function, δ t =G t -V(S t ).
[0222] In some embodiments, the advantage function δ t It is usually estimated by the time series difference method, and the time series difference error is
[0223] δ t =R t+1 +γV(S t+1 )-V(S t ),
[0224] Among them, R t+1 is the reward value at time t+1, V(S t ) and V(S t+1 ) are the current states S t and the next state S t+1 The value function of .
[0225] In some embodiments,
[0226] The formula of the loss function is
[0227]
[0228] in, is the importance sampling ratio, indicating that action a is executed under the current strategy t The ratio of the probability of executing the action to the probability of executing the action under the old strategy; π θ (a t |S t ) is the probability of training the policy, is the policy probability at the previous time step; is the advantage estimate, which represents the advantage of the current state-action pair, that is, the gap between the actual return and the estimated return; ε is the clipping parameter, which controls the magnitude of the policy change. When r t When (θ) exceeds the range [1-ε, 1+ε], it is clamped to the boundary value.
[0229] In some embodiments, KL divergence loss is used to control the difference in strategy between the training strategy and the reference model; the smaller the KL divergence, the more similar the strategy of the training model is to the strategy of the reference model. The formula is:
[0230]
[0231] State value loss is used to optimize the state value function V θ (S t ), which is used to measure the expected value of future rewards obtained by using the strategy in a certain state; the goal of the state value loss is to make the state value function predicted by the model close to the expected return, and the formula is
[0232]
[0233] Among them, V θ (S t ) is the state value function, which means that in state S t Under the current strategy π θ , the expected reward from this state; represents the cumulative reward starting from time step t; E t represents the expected operation at time step t;
[0234] The final loss function formula is
[0235]
[0236] Among them, c1 and c2 are hyperparameters used to balance the weights of various losses.
[0237] The embodiment of the present application provides a key sentence extraction device that combines keyword guidance with proximal strategy optimization of a large language model, constructs keyword-key sentence pairs, uses a joint matching model to evaluate their relevance, generates a reward value, introduces KL divergence to measure the difference between the training model and the reference model, and estimates the value score of the current state in combination with the state value network. The proximal strategy optimization guidance model is used to achieve key sentence extraction, which largely solves the problems of strong dependence on labeled data and high training cost of large language models. It performs excellently in key sentence extraction tasks, has a high ROUGE-L index on scientific literature datasets, and significantly improves the effect of key sentence extraction.
[0238] Another embodiment of the present application provides an electronic device, including a processor and a memory; when the processor executes the running program stored in the memory, the method described in any embodiment of the present application is implemented.
[0239] For example, Figure 6 As shown, Figure 6 The following is a schematic diagram of the structure of an electronic device provided as a specific example, including a processor 601, a communication interface 602, a memory 603, and a communication bus 604. The processor 601, the communication interface 602, and the memory 603 communicate with each other via the communication bus 604. The memory 603 is used to store computer programs; the processor 601 is used to execute the programs stored in the memory 603, thereby implementing the methods provided in any embodiment of the present application.
[0240] In the electronic device described above, the communication bus may be a peripheral component interconnect standard bus or an extended industry standard architecture bus. The communication bus may include an address bus, a data bus, and a control bus. The communication interface is used to enable data exchange between the electronic device and other devices.
[0241] Memory 603 may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, memory 603 may include other storage devices located remotely from processor 601. Processor 601 may be a general-purpose processor, such as a central processing unit (CPU) or a network processor (NP); it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other types of programmable logic devices, discrete gate circuits or transistor logic devices, or discrete hardware components.
[0242] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in any embodiment of the present application when the computer program is executed by a processor. Figure 6 As shown, Figure 7 The computer-readable storage medium shown is an optical disc 20 , on which a computer program (ie, a program product) is stored. When the computer program is executed by a processor, the method described in any embodiment of the present application can be implemented.
[0243] The computer-readable storage medium of this embodiment can be any type of medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates multiple media. Common storage media types include magnetic media (such as floppy disks, hard disks, and magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid-state drives (SSDs)).
[0244] In the above embodiments, all or part of the functions can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the functions can be implemented by one or more computer instructions contained in a computer-readable storage medium. When a computer loads and executes these computer program instructions, it can partially or completely perform the processes or functions described in the embodiments of the present invention.
[0245] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions may be stored in a computer-readable storage medium and may be transmitted from one computer-readable storage medium to another via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. For example, computer instructions may be transmitted from a website, computer, server, or data center to another computer, website, server, or data center.
[0246] It should be noted that:
[0247] In the embodiments of the present application, the terms "include", "comprising" and any other variations thereof are intended to indicate a non-exclusive inclusion relationship, meaning that when referring to including certain elements, the process, method, article or apparatus includes not only these elements, but also other elements not explicitly listed, or inherent elements related to the process, method, article or apparatus. Without further qualification, the use of "comprising a..." to express an element does not exclude the possibility that other identical elements may exist in the process, method, article or apparatus that includes the element.
[0248] By the description of the above-mentioned embodiment method, it will be clearly understood by those skilled in the art that these methods can be implemented by software plus necessary general hardware platforms. Of course, it can also be implemented by pure hardware, but in many cases, the former is generally a more optimal implementation method. Based on this understanding, the technical solutions of the embodiments of the present application, or the technical contributions made by the present application, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk, etc.), including several instructions, so that an image display device (such as a mobile phone, a computer, a server, an air-conditioning device or a network device, etc.) executes the various methods described in the embodiments of the present application.
[0249] The above description is a specific implementation of the embodiments of the present application, but does not limit the scope of protection of the present application. Any changes or alternatives that can be conceived by any person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be determined based on the content of the claims.
Claims
1. A key sentence extraction method combining keyword guidance and proximal strategy optimization of a large language model, characterized in that: include: Construct keyword-key sentence pairs; Use the joint matching model to evaluate their relevance and generate a reward value; KL divergence is introduced to measure the difference between the training model and the reference model, and the value score of the current state is estimated by combining the state value network; The extraction of key sentences is achieved by optimizing the guidance model through proximal strategy.
2. The method according to claim 1, characterized in that The construction of keyword-key sentence pairs includes: Each key sentence text after preprocessing is represented as K = {k1, k2, ...k n }, each keyword is represented by L = {l1, l2, ..., l m }, where k n Indicates the nth character of the key sentence, l m Indicates the mth character of the keyword; The pre-trained Chinese BERT model is used to map text into feature vectors, generating vectorized representations of key sentences and keywords respectively. Calculate the similarity score between each keyword and key sentence through cosine similarity; For each keyword, the key sentence with the highest similarity score that is not repeated with other keywords is selected to form a keyword-key sentence matching pair.
3. The method according to claim 1, characterized in that The method of using the joint matching model to evaluate the relevance and generate a reward value includes: Input the generated keyword-key sentence matching pairs into the CEDR-DRMM model and calculate the matching score of each pair; All matching scores are normalized and the final reward score is generated by weighted average.
4. The method according to claim 1, wherein The KL divergence is introduced to measure the difference between the training model and the reference model, and the value score of the current state is estimated by combining the state value network, including: The expected cumulative return in a state is predicted by estimating the value function V(S) of the state. Through optimization training, the state value network minimizes the difference between the value estimate of its output and the actual return. The state value network provides a value V(S) for each state s, which represents the cumulative rewards that may be obtained in the future under this state; In the learning process, the value function of the current state and the next state is estimated, that is, V(S t ) and V(S t+1 ), and the actual reward R t+1 , calculate the temporal difference error; the temporal difference error is used to update the policy network; In reinforcement learning, from a certain state S t Start taking action a t The cumulative return after is defined as G t =R t+1 +γR t+2 +γ 2 R t+3 +… Among them, G t is the cumulative return starting from time t, γ is the discount factor; State value function V(S t ) indicates that in state S t The expected return under the current strategy is defined as V(S t )=E[G t |S t ]; Advantage function δ t Measured in state S t Next select action a t The difference between the actual return and the state value function, δ t =G t -V(S t ).
5. The method according to claim 4, characterized in that The advantage function δ t It is usually estimated by the time series difference method, and the time series difference error is δ t =R t+1 +γV(S t+1 )-V(S t ), Among them, R t+1 is the reward value at time t+1, V(S t ) and V(S t+1 ) are the current states S t and the next state S t+1 The value function of .
6. The method according to claim 5, characterized in that The formula of the loss function is in, is the importance sampling ratio, indicating that action a is executed under the current strategy t The ratio of the probability of executing the action to the probability of executing the action under the old strategy; π θ (a t |S t ) is the probability of training the policy, is the policy probability of the previous time step; is the advantage estimate, which represents the advantage of the current state-action pair, that is, the gap between the actual return and the estimated return; ε is the clipping parameter, which controls the magnitude of the policy change. When r t When (θ) exceeds the range [1-ε, 1+ε], it is clamped to the boundary value.
7. The method according to any one of claims 1 to 6, characterized in that KL divergence loss is used to control the difference between the training strategy and the reference model strategy. The smaller the KL divergence, the more similar the strategy of the training model is to the strategy of the reference model. The formula is: State value loss is used to optimize the state value function V θ (S t ), which is used to measure the expected value of future rewards obtained by using the strategy in a certain state; the goal of the state value loss is to make the state value function predicted by the model close to the expected return, and the formula is Among them, V θ (S t ) is the state value function, which means that in state S t Under the current strategy π θ , the expected reward from this state; represents the cumulative reward starting from time step t; E t represents the expected operation at time step t; The final loss function formula is Among them, c1 and c2 are hyperparameters used to balance the weights of various losses.
8. A key sentence extraction device combining keyword guidance and large language model proximal strategy optimization, characterized in that: include: Construction module, used to construct keyword-key sentence pairs; A relevance evaluation module is used to evaluate the relevance using a joint matching model and generate a reward value; The value score estimation module is used to introduce KL divergence to measure the difference between the training model and the reference model, and estimate the value score of the current state in combination with the state value network; The extraction module is used to extract key sentences by optimizing the guidance model through proximal strategy.
9. An electronic device, characterized in that: It includes a processor and a memory; when the processor executes the running program stored in the memory, it implements the key sentence extraction method combining keyword guidance and large language model proximal strategy optimization as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the key sentence extraction method combining keyword guidance and large language model proximal strategy optimization is implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Feedback-based model training method, keyword extraction method and related equipment
CN116957056A
Reinforcement learning training method and system based on military document and answer similarity
CN119005290A
Reading type examination question generation system and method based on commonsense reasoning
WO2023225858A1
Cited By
Method and device for generating graphic design drawing
CN122049121A