Entity relationship joint extraction method and system for Chinese text in carbon finance field

By introducing a hybrid attention mechanism and zero-initialization training strategy into the Atom-7B large language model, the parameter scale and causal attention mechanism limitations of traditional models in Chinese text relation extraction are resolved, and efficient joint extraction of entity relations is achieved.

CN120654694APending Publication Date: 2025-09-16ZHEJIANG UNIV OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510737914.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In existing technologies, the traditional encoder model has limited parameter scale, making it difficult to break through the low efficiency of knowledge updating, while the causal attention mechanism of the decoder architecture has inherent defects in modeling bidirectional semantic relationships, resulting in poor Chinese text relationship extraction results.

Method used

The right-side attention mechanism and bidirectional attention mechanism are introduced into the Atom-7B large language model to construct a hybrid attention mechanism. A zero-initialization training strategy is adopted to maintain the consistency of the model with the original pre-training distribution, and entity relationship joint extraction is performed through the PFN module.

Benefits of technology

It effectively improves the relationship extraction effect of Chinese text, gives full play to the large-scale knowledge capacity and semantic understanding ability, makes up for the shortcomings of traditional decoders, and improves the adaptability and performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654694A_ABST
    Figure CN120654694A_ABST
Patent Text Reader

Abstract

An entity relationship joint extraction method of a Chinese text in the field of carbon finance comprises the following steps: aiming at the carbon finance Chinese text, obtaining feature representation vectors of text sentences by utilizing an Ato-7B large model, and then obtaining an entity relationship triple in the text in a joint extraction mode through PFN. The invention further provides an entity relation joint extraction system for the Chinese text in the carbon amount field. On the basis of keeping a causal attention mechanism of an Ato-7B large language model, a right attention mechanism and a bidirectional attention mechanism are innovatively introduced, a mixed attention mechanism is constructed, the advantages of Ato-7B in the aspects of scale and knowledge capacity are brought into full play, and the method is suitable for popularization and application. Meanwhile, inherent defects of a traditional decoder in a relation extraction task are made up through a mixed attention mechanism, and the relation extraction effect is effectively improved. According to the method, the performance of extracting the entity relation triad in the Chinese text in the carbon finance field is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular to a method and system for jointly extracting entity relationships from Chinese texts in the field of carbon finance. Background Art

[0002] In the process of constructing the carbon finance knowledge graph, relationship extraction is a core technical link, and its role is to automatically extract structured triples (head entity, relationship, tail entity) from unstructured text. Traditional relationship extraction methods mainly adopt a pipeline processing framework to perform entity recognition and relationship classification tasks in stages. In recent years, joint extraction methods have achieved end-to-end triple extraction by constructing a joint model of entity recognition and relationship classification, showing significant advantages in performance indicators. In the existing technology, PFN (A Partition Filter Network for Joint Entity and Relation Extraction, 2021) proposes a partition filtering network architecture, which promotes the collaborative optimization of named entity recognition and relationship extraction tasks through a deep interaction mechanism, and achieves better extraction performance.

[0003] Currently, there are two main technical approaches in the field of relation extraction: encoder models based on bidirectional attention mechanisms and decoder models based on causal attention. The former maintains a performance advantage in benchmark tests, but is limited by model size and knowledge capacity bottlenecks. Typical encoder models, such as BERT, typically have parameter sizes of no more than hundreds of millions and suffer from inefficient knowledge updates. The latter, represented by the LLaMA large language model (LLM), while possessing a large parameter size and high-density knowledge representation capabilities, its autoregressive generation mechanism requires the use of causal attention, resulting in the model's inability to fully capture the bidirectional semantic dependencies of text.

[0004] The existing technical system faces a dual challenge: on the one hand, traditional encoder models are difficult to break through the limitation of parameter scale, and directly training large-scale encoders requires huge computing resources and massive amounts of labeled data; on the other hand, the causal attention mechanism of the decoder architecture has inherent defects in modeling bidirectional semantic relationships. How to effectively integrate the technical advantages of encoders and decoders to build a relationship extraction model with both bidirectional semantic understanding capabilities and large-scale knowledge capacity has become a key technical problem that needs to be solved in this field. Large models such as LLaMA are suitable for English text, while the Atom-7B large model is an open source large model obtained by pre-training on large-scale Chinese data based on LLaMA2-7B. It has better language understanding and generation capabilities for Chinese text. Atom-7B contains 32 Transformer layers. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this paper provides a method and system for jointly extracting entity relationships from Chinese text in the field of carbon finance. Targeting the Atom-7B large language model, this paper introduces a right-side attention mechanism and a bidirectional attention mechanism while maintaining the causal attention mechanism, creating a hybrid attention mechanism. Furthermore, different attention modes are dynamically integrated to ensure the effectiveness and adaptability of the hybrid attention. Furthermore, a zero-initialization training strategy is employed to effectively maintain consistency with the original pre-trained distribution during the initial stages of model fine-tuning, avoiding potential performance degradation. Finally, the semantic representation generated by the Atom-7B is fed into the relation extraction model (PFN) for joint entity relationship extraction.

[0006] The technical solution adopted by the present invention to solve its technical problem is:

[0007] A method for jointly extracting entity relationships from Chinese text in the field of carbon finance includes the following steps:

[0008] Step 1: Collect Chinese carbon finance text data: Batch crawl raw data of notices and announcements, obtain unstructured text of the entity relationships to be extracted from the specified paragraphs, and a given set of ontology constraints, which includes the relationship name, head entity type, and tail entity type. The labeled Chinese text data in the carbon finance field is divided into a training set D1 and a validation set D2 according to a preset ratio.

[0009] Step 2: Hybrid Attention Enhancement for the Atom-7B Large Language Model: For the Atom-7B large language model, while maintaining the causal attention mechanism, we introduce the right-side attention mechanism and the bidirectional attention mechanism to construct a hybrid attention mechanism. At the same time, we adopt a zero-initialization training strategy and use the Atom-7B large language model to obtain the semantic representation of the text sentences in the training set D1 as the input for Step 3.

[0010] Step 3: Train the entity-relationship joint extraction model: Input the semantic representation output by Atom-7B into the entity-relationship joint extraction model PFN, output the entity and relationship prediction values ​​corresponding to the text sentence, and jointly optimize the model parameters based on the loss gradient;

[0011] Step 4: Repeat steps 2 and 3 until the maximum number of iterations is reached, and save the parameters of the entity relationship joint extraction model with the highest F1 score on the validation set D2.

[0012] Step 5: Input unlabeled text sentences in the carbon finance field into the trained Atom-7B large language model to obtain the feature representation vector of the text sentence. Then, the feature representation vector is passed to the PFN module to output entity relationship triples.

[0013] Furthermore, in step 2, the process of hybrid attention enhancement of the Atom-7B large language model is as follows:

[0014] 2.1 Convert the input text into tokens and map the tokens into a d-dimensional vector H through the embedding layer mapping function Emb(·) 0 As the input of the first-layer Transformer, the semantic modeling of the input sequence is achieved through a multi-layer Transformer architecture:

[0015] H 0 =Emb(t)

[0016]

[0017] Where t={t1,t2,...,t n} represents the tokens of the input text, Stack(Transformer 1 ,...,Transformer L ) represents the encoding module consisting of L stacked Transformer layers, is the last layer representation of the output;

[0018] 2.2 For the lth Transformer layer, its input H l For the representation of the previous layer, the query matrix Q, key matrix K, and value matrix V are generated through linear projection respectively:

[0019]

[0020] in, represents the learnable parameter matrix;

[0021] 2.3 The causal mask M causal , right mask M right and bidirectional mask M bidir Applied to the attention score, and after the Softmax operation, the causal attention weight, the right attention weight, and the bidirectional attention weight are obtained respectively, and multiplied with the value matrix V to generate the causal representation L, the right representation R, and the bidirectional representation B. Process the left, right and bidirectional context information separately:

[0022]

[0023]

[0024] M bidir =0

[0025] Where d is the dimension of the key matrix, and the causal mask M causal To ensure the causality of autoregressive generation, right mask M rightSo that tokenj can only focus on the pre-order tokeni that meets the condition i>j, the bidirectional mask M bidir Allow any token pair (t i ,t j ) Free interaction;

[0026] 2.4 Calculate the weight σ1, σ2 of each token through mapping operation:

[0027]

[0028] in, For the learnable parameter vector, the zero initialization strategy is used to initialize the parameters Initialize to ensure that the attention weights σ1 and σ2 are equal to 0 in the initial stage of fine-tuning, and use tanh(·) as a nonlinear activation function to normalize the original scores to the interval (-1, 1) and satisfy the mathematical property that zero input corresponds to zero output.

[0029] 2.5 Three Representations Based on the weights σ1, σ2, a weighted sum is performed to obtain a mixed representation:

[0030] H l+1 =σ1·R+σ2·B+(1-σ1-σ2)·L

[0031] in, is a mixed representation that will be used as the input to the l+1 layer, Able to effectively control the aggregation ratio of each token context information;

[0032] 2.6 If the last layer of Transformer is not reached, the mixed representation is input to the next layer of Transformer, and steps 2.2 to 2.5 are repeated iteratively, and the representation of the last layer after stacking multiple layers of Transformer is output as the input of the entity relationship joint extraction module.

[0033] Furthermore, in step 3, the process of training the entity relationship joint extraction model is as follows:

[0034] 3.1 Represent the last layer output of step 2 H L Input the PFN model to get the predicted values ​​of entities and relations respectively and in represents the probability that the token pair starting with the i-th token and ending with the j-th token belongs to an entity of type k, Express the probability that the i-th token and the j-th token belong to the subject entity and the object entity starting word of the relation type l, and calculate the overall loss function, that is, add the losses of named entity recognition and relation extraction as the total loss:

[0035]

[0036] L overall =L ner +L re

[0037] Among them, θ is the parameter of the PFN model, and are the labels of entities and relations, S and T are the sets of all entity labels and relation labels, respectively. ner and L re They are the losses of the named entity recognition subtask and the relation extraction subtask, respectively, and BCELoss is the binary cross entropy loss;

[0038] 3.2 Jointly optimize the parameters of the entity relationship joint extraction model: According to the gradient value of the loss function, all parameters of the PFN model θ and the Atom-7B large language model are fine-tuned.

[0039] A system for jointly extracting entity relationships from Chinese texts in the field of carbon finance, the system comprising:

[0040] The carbon finance data collection module is used to obtain data related to carbon finance, carbon emissions, and carbon trading. It obtains unstructured carbon finance Chinese text with entity relationships to be extracted from a specified paragraph, as well as a given ontology constraint set. The ontology constraint set includes the relationship name, head entity type, and tail entity type. The labeled carbon finance Chinese text data is divided into a training set D1 and a validation set D2 according to a preset ratio. The Chinese text data includes the subject, object, relationship, and category label contained in the current sample.

[0041] The hybrid attention module, while maintaining the causal attention mechanism of Atom-7B, introduces the right-side attention mechanism and the bidirectional attention mechanism to build a hybrid attention mechanism to exploit the semantic representation capabilities of the Atom-7B large language model. It inputs sentence text and outputs semantic representation.

[0042] In the joint extraction model training module, in the PFN module, the feature representation vector of the text sentence is input to obtain the corresponding entity and relationship prediction probabilities, and the loss weights of named entity recognition and relationship extraction are calculated. In the joint optimization, all parameters of the Atom-7B and PFN models are fine-tuned, and the model parameters with the best extraction performance on the validation set D2 are saved;

[0043] The entity-relationship triplet output module is used to input unlabeled carbon finance Chinese text into the trained entity-relationship joint extraction model, obtain the feature representation vector of the text sentence through the Atom-7B large model, and then output the entity-relationship triplet through the PFN module.

[0044] The technical concept of the present invention is: in order to better stimulate the feature extraction capability of the Atom-7B large language model for Chinese text sentences, on the basis of maintaining the causal attention mechanism of Atom-7B, the right attention mechanism and the bidirectional attention mechanism are introduced to construct a hybrid attention mechanism. Secondly, different attention modes are dynamically integrated to ensure the effectiveness of the hybrid attention. At the same time, a zero-initialization training strategy is adopted to effectively maintain the consistency of the model with the original pre-training distribution in the early stage of fine-tuning, avoiding potential performance degradation. The PFN module is used to effectively utilize the semantic standards obtained by Atom-7B to generate high-quality relationship triples.

[0045] The beneficial effects of the present invention are as follows: for the Atom-7B large language model, while maintaining the causal attention mechanism, the right-side attention mechanism and the bidirectional attention mechanism are introduced to construct a hybrid attention mechanism; secondly, different attention modes are dynamically integrated to ensure the effectiveness and adaptability of the hybrid attention; finally, a zero-initialization training strategy is adopted to effectively maintain consistency with the original pre-training distribution during the initial stage of model fine-tuning, avoiding potential performance degradation. These measures effectively tap into the Atom-7B large language model's ability to extract features from Chinese text sentences, fully leveraging Atom-7B's advantages in scale and knowledge capacity. At the same time, the hybrid attention mechanism compensates for the inherent defects of traditional decoders in relation extraction tasks, effectively improving the relation extraction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 A block diagram of the method of the present invention. DETAILED DESCRIPTION

[0047] The present invention will be further described below with reference to the accompanying drawings.

[0048] Reference Figure 1 , a joint entity relationship extraction method for Chinese text in the field of carbon finance, including the following steps:

[0049] Step 1: Collection of Chinese carbon finance text data: Batch crawling of raw data related to notices and announcements and related information from the China Carbon Accounting Database, Global Real-Time Carbon Data, Shanghai Energy Exchange, Ministry of Ecology and Environment of the People's Republic of China, Azure Map, World Bank Database, International Institute of Green Finance of Central University of Finance and Economics, etc. Obtain unstructured text of the entity relationships to be extracted from the specified paragraphs, as well as a given set of ontology constraints, including the relationship name, head entity type, and tail entity type. Divide the annotated carbon finance text data into a training set D1 and a validation set D2 according to a preset ratio;

[0050] Step 2: Hybrid Attention Enhancement for the Atom-7B Large Language Model: For the Atom-7B large language model, while maintaining the causal attention mechanism, we introduce the right-side attention mechanism and the bidirectional attention mechanism to construct a hybrid attention mechanism. At the same time, we adopt a zero-initialization training strategy and use the Atom-7B large language model to obtain the semantic representation of the text sentences in the training set D1 as the input for Step 3.

[0051] The process of step 2 is as follows:

[0052] 2.1 Convert the input text into tokens and map the tokens into a d-dimensional vector H through the embedding layer mapping function Emb(·) 0 As the input of the first-layer Transformer, the semantic modeling of the input sequence is achieved through a multi-layer Transformer architecture:

[0053] H 0 =Emb(t)

[0054] H L =Stack(Transformer 1 ,...,Transformer L )(H 0 )

[0055] Where t={t1,t2,...,t n} represents the tokens of the input text, Stack(Transformer 1 ,...,Transformer L ) represents the encoding module consisting of L stacked Transformer layers, is the last layer representation of the output;

[0056] 2.2 For the lth Transformer layer, its input H l For the representation of the previous layer, the query matrix Q, key matrix K, and value matrix V are generated through linear projection respectively:

[0057]

[0058] in, represents the learnable parameter matrix;

[0059] 2.3 The causal mask M causal , right mask M right and bidirectional mask M bidir Applied to the attention score, and after the Softmax operation, the causal attention weight, the right attention weight, and the bidirectional attention weight are obtained respectively, and multiplied with the value matrix V to generate the causal representation L, the right representation R, and the bidirectional representation B. Process the left, right and bidirectional context information separately:

[0060]

[0061]

[0062] M bidir =0

[0063] Where d is the dimension of the key matrix, and the causal mask M causal To ensure the causality of autoregressive generation, right mask M right So that tokenj can only focus on the pre-order tokeni that meets the condition i>j, the bidirectional mask M bidir Allow any token pair (t i ,t j ) Free interaction;

[0064] 2.4 Calculate the weight σ1, σ2 of each token through mapping operation:

[0065]

[0066] in, For the learnable parameter vector, the zero initialization strategy is used to initialize the parameters Initialize to ensure that the attention weights σ1 and σ2 are equal to 0 in the initial stage of fine-tuning, and use tanh(·) as a nonlinear activation function to normalize the original scores to the interval (-1, 1) and satisfy the mathematical property that zero input corresponds to zero output.

[0067] 2.5 Three Representations Based on the weights σ1, σ2, a weighted sum is performed to obtain a mixed representation:

[0068] H l+1 =σ1·R+σ2·B+(1-σ1-σ2)·L

[0069] in, is a mixed representation that will be used as the input to the l+1 layer, Able to effectively control the aggregation ratio of each token context information;

[0070] 2.6 If the last layer of Transformer is not reached, the mixed representation is input to the next layer of Transformer, and steps 2.2 to 2.5 are repeated iteratively, and the representation of the last layer after stacking multiple layers of Transformer is output as the input of the entity relationship joint extraction module.

[0071] Step 3: Train the entity-relationship joint extraction model: Input the semantic representation output by Atom-7B into the entity-relationship joint extraction model PFN, output the entity and relationship prediction values ​​corresponding to the text sentence, and jointly optimize the model parameters based on the loss gradient;

[0072] The process of step three is as follows:

[0073] 3.1 Represent the last layer output of step 2 H L Input the PFN model to get the predicted values ​​of entities and relations respectively and in represents the probability that the token pair starting with the i-th token and ending with the j-th token belongs to an entity of type k, Express the probability that the i-th token and the j-th token belong to the subject entity and the object entity starting word of the relation type l, and calculate the overall loss function, that is, add the losses of named entity recognition and relation extraction as the total loss:

[0074]

[0075] L overall =L ner +L re

[0076] Among them, θ is the parameter of the PFN model, and are the labels of entities and relations, S and T are the sets of all entity labels and relation labels, respectively. ner and L re They are the losses of the named entity recognition subtask and the relation extraction subtask, respectively, and BCELoss is the binary cross entropy loss;

[0077] 3.2 Jointly optimize the parameters of the entity relationship joint extraction model: According to the gradient value of the loss function, all parameters of the PFN model θ and the Atom-7B large language model are fine-tuned.

[0078] Step 4: Repeat steps 2 and 3 until the maximum number of iterations is reached, and save the parameters of the entity relationship joint extraction model with the highest F1 score on the validation set D2;

[0079] Step 5: Input unlabeled text sentences in the carbon finance field into the trained Atom-7B large language model to obtain the feature representation vector of the text sentence. Then, the feature representation vector is passed to the PFN module to output entity relationship triples.

[0080] This embodiment also provides a system for jointly extracting entity relationships from Chinese text in the field of carbon finance, comprising: a carbon finance data collection module, a hybrid attention module, a joint extraction model training module, and an entity relationship triple output module. Each of these modules corresponds to steps 1 to 5 of the method of the present invention, i.e., the system comprises:

[0081] The carbon finance data collection module is used to obtain relevant data such as carbon finance, carbon accounting, and carbon trading. It obtains the unstructured carbon finance Chinese text with the entity relationships to be extracted from the specified paragraph, as well as a given ontology constraint set. The ontology constraint set includes the relationship name, head entity type, and tail entity type. The labeled carbon finance Chinese text data is divided into a training set D1 and a validation set D2 according to a preset ratio. The Chinese text data includes the subject, object, relationship, and category label contained in the current sample;

[0082] The hybrid attention module, while maintaining the causal attention mechanism of Atom-7B, introduces the right-side attention mechanism and the bidirectional attention mechanism to construct a hybrid attention mechanism to explore the semantic representation capabilities of the Atom-7B large language model, input sentence text, and output semantic representation.

[0083] In the joint extraction model training module, in the PFN module, the feature representation vector of the text sentence is input to obtain the corresponding entity and relationship prediction probabilities, and the loss weights of named entity recognition and relationship extraction are calculated. In the joint optimization, all parameters of the Atom-7B and PFN models are fine-tuned, and the model parameters with the best extraction performance on the validation set D2 are saved;

[0084] The entity-relationship triplet output module is used to input unlabeled carbon finance Chinese text into the trained entity-relationship joint extraction model, obtain the feature representation vector of the text sentence through the Atom-7B large model, and then output the entity-relationship triplet through the PFN module.

[0085] As described above, the specific implementation steps of this patent make the present invention clearer. Any modifications and changes made to the present invention within the spirit of the present invention and the scope of protection of the claims fall within the scope of protection of the present invention.

Claims

1. A joint entity relationship extraction method for Chinese text in the field of carbon finance, characterized by: The method comprises the following steps: Step 1: Collect Chinese carbon finance text data: Batch crawl raw data of notices and announcements, obtain unstructured text of the entity relationships to be extracted from the specified paragraphs, and a given set of ontology constraints, which includes the relationship name, head entity type, and tail entity type. The labeled carbon finance text data is divided into a training set D1 and a validation set D2 according to a preset ratio. Step 2: Hybrid Attention Enhancement for the Atom-7B Large Language Model: For the Atom-7B large language model, while maintaining the causal attention mechanism, we introduce the right-side attention mechanism and the bidirectional attention mechanism to construct a hybrid attention mechanism. At the same time, we adopt a zero-initialization training strategy and use the Atom-7B large language model to obtain the semantic representation of the text sentences in the training set D1 as the input for Step 3. Step 3: Train the entity-relationship joint extraction model: Input the semantic representation output by Atom-7B into the entity-relationship joint extraction model PFN, output the entity and relationship prediction values ​​corresponding to the text sentence, and jointly optimize the model parameters based on the loss gradient; Step 4: Repeat steps 2 and 3 until the maximum number of iterations is reached, and save the parameters of the entity relationship joint extraction model with the highest F1 score on the validation set D2. Step 5: Input unlabeled text sentences in the carbon finance field into the trained Atom-7B large language model to obtain the feature representation vector of the text sentence. Then, the feature representation vector is passed to the PFN module to output entity relationship triples.

2. The entity relationship joint extraction method of Chinese text in the field of carbon finance as claimed in claim 1 is characterized in that: In step 2, the process of hybrid attention enhancement of the Atom-7B large language model is as follows: 2.1 Convert the input text into tokens and map the tokens into a d-dimensional vector H through the embedding layer mapping function Emb(·) 0 As the input of the first-layer Transformer, the semantic modeling of the input sequence is achieved through a multi-layer Transformer architecture: H 0 =Emb(t) H L =Stack(Transformer 1 ,...,Transformer L )(H 0 ) Where t={t1,t2,...,t n } represents the tokens of the input text, Stack(Transformer 1 ,...,Transformer L ) represents the encoding module consisting of L stacked Transformer layers, is the last layer representation of the output; 2.2 For the lth Transformer layer, its input H l For the representation of the previous layer, the query matrix Q, key matrix K, and value matrix V are generated through linear projection respectively: in, represents the learnable parameter matrix; 2.3 The causal mask M causal , right mask M right and bidirectional mask M bidir Applied to the attention score, and after the Softmax operation, the causal attention weight, the right attention weight, and the bidirectional attention weight are obtained respectively, and multiplied with the value matrix V to generate the causal representation L, the right representation R, and the bidirectional representation B. Process the left, right and bidirectional context information separately: M bidir =0 Where d is the dimension of the key matrix, and the causal mask M causal To ensure the causality of autoregressive generation, right mask M right So that tokenj can only focus on the pre-order tokeni that meets the condition i>j, the bidirectional mask M bidir Allow any token pair (t i ,t j ) Free interaction; 2.4 Calculate the weight σ1, σ2 of each token through mapping operation: in, For the learnable parameter vector, the zero initialization strategy is used to initialize the parameters Initialize to ensure that the attention weights σ1 and σ2 are equal to 0 in the initial stage of fine-tuning, and use tanh(·) as a nonlinear activation function to normalize the original scores to the interval (-1, 1) and satisfy the mathematical property that zero input corresponds to zero output. 2.5 Three Representations Based on the weights σ1, σ2, a weighted sum is performed to obtain a mixed representation: H l+1 =σ1·R+σ2·B+(1-σ1-σ2)·L in, is a mixed representation that will be used as the input to the l+1 layer, Able to effectively control the aggregation ratio of each token context information; 2.6 If the last layer of Transformer is not reached, the mixed representation is input to the next layer of Transformer, and steps 2.2 to 2.5 are repeated iteratively, and the representation of the last layer after stacking multiple layers of Transformer is output as the input of the entity relationship joint extraction module.

3. The entity relationship joint extraction method of Chinese text in the field of carbon finance as claimed in claim 2 is characterized in that: In step 3, the process of training the entity relationship joint extraction model is as follows: 3.1 Represent the last layer output of step 2 H L Input the PFN model to get the predicted values ​​of entities and relations respectively and in represents the probability that the token pair starting with the i-th token and ending with the j-th token belongs to an entity of type k, Express the probability that the i-th token and the j-th token belong to the subject entity and the object entity starting word of the relation type l, and calculate the overall loss function, that is, add the losses of named entity recognition and relation extraction as the total loss: L overall =L ner +L re Among them, θ is the parameter of the PFN model, and are the labels of entities and relations, S and T are the sets of all entity labels and relation labels, respectively. ner and L re They are the losses of the named entity recognition subtask and the relation extraction subtask, respectively, and BCELoss is the binary cross entropy loss; 3.2 Jointly optimize the parameters of the entity relationship joint extraction model: According to the gradient value of the loss function, all parameters of the PFN model θ and the Atom-7B large language model are fine-tuned.

4. A system for extracting entity relationships from Chinese texts in the field of carbon finance as claimed in claim 1, characterized in that: The system includes a carbon finance data collection module, a hybrid attention module, a joint extraction model training module and an entity relationship triplet output module. The carbon finance data collection module is used to obtain data related to carbon finance, carbon emissions, and carbon trading. It obtains unstructured carbon finance Chinese text with entity relationships to be extracted from a specified paragraph, as well as a given ontology constraint set, where the ontology constraint set includes the relationship name, head entity type, and tail entity type. The labeled carbon finance Chinese text data is divided into a training set D1 and a validation set D2 according to a preset ratio. The Chinese text data includes the subject, object, relationship, and category label contained in the current sample. The hybrid attention module, while maintaining the causal attention mechanism of Atom-7B, introduces the right-side attention mechanism and the bidirectional attention mechanism to construct a hybrid attention mechanism for mining the semantic representation capability of the Atom-7B large language model. It inputs sentence text and outputs semantic representation. In the joint extraction model training module and the PFN module, the feature representation vector of the text sentence is input to obtain the corresponding entity and relationship prediction probabilities, and the loss weights of named entity recognition and relationship extraction are calculated. In the joint optimization, all parameters of the Atom-7B and PFN models are fine-tuned, and the model parameters with the best extraction performance on the validation set D2 are saved. The entity relationship triple output module is used to input unlabeled carbon finance Chinese text into the trained entity relationship joint extraction model, obtain the feature representation vector of the text sentence through the Atom-7B large model, and then output the entity relationship triple through the PFN module.

Citation Information

Patent Citations

  • Entity relationship joint extraction method and system for data discovery

    CN117743475A

  • Text generation method and device and computing equipment

    CN118364877A

  • Entity relationship joint extraction method and system for Chinese text in carbon neutralization field

    CN118585643A

  • MRI image segmentation system and method based on mixed attention supervision U-shaped network

    CN118587442A

  • Entity relationship classification method and related equipment

    CN119721036A