Generative large model-based literature reference content extraction method

By constructing special symbol data content in the generative large language model for fine-tuning and changing its input and output forms, the problem of slow speed and inconsistency in the extraction of cited content in literature is solved, and the speed and effect are improved, while ensuring complete consistency of the content.

CN120046583AActive Publication Date: 2025-05-27TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510525209.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The generative large language model is slow in the process of extracting citation content in literature, and the extracted citation content may have the problem that some of the words and symbols are inconsistent with the original document content.

Method used

By building data content containing special symbols for fine-tuning of the big model, changing the input and output form of its extracted referenced content, reducing the input length of the big model, thereby improving the extraction speed and effect.

Benefits of technology

The speed and effect of extracting cited content of literature has been improved, while ensuring the complete consistency between the extracted cited content and the original document content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046583A_ABST
    Figure CN120046583A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of natural language processing, particularly relates to a method for extracting literature reference content based on a generative large model, and aims to solve the problems that an existing large model is relatively slow in extraction and characters in the content are inconsistent. The method comprises the following steps: constructing model fine tuning data, and performing parameter fine tuning on a generative large model to obtain a reference extraction model; obtaining a to-be-processed document, and further obtaining a text paragraph set; performing formalized special symbol label conversion on each text paragraph in the text paragraph set; and inputting the converted text paragraph set into the large model for reference extraction to obtain a segmentation mark sequence number corresponding to a reference text, and mapping the segmentation mark sequence number to an original text to obtain reference content. According to the method, the reference content is replaced by the special symbols, the input and output forms of the extracted reference content are changed, and the input length of a large model is reduced, so that the extraction speed and the extraction effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] The citation content in the literature reflects the relationships of association, inheritance, and development among different literature contents. The extraction of citation content in the literature plays an important role in fields such as content review and summary, content evaluation, academic tracking, and knowledge graph construction. However, with the continuous development of academic research and the massive growth of the number of literatures, higher requirements are put forward for the accuracy, efficiency, and adaptability of the extraction of citation content in the literature.

[0003] Traditional methods for extracting citation content in the literature mainly rely on pattern matching, indicator word search, and shallow machine learning models. Pattern matching identifies citation content by manually summarizing the footnote and endnote identifiers of citation content (such as "[1]", "【1】", "(Smith, 2020)") and various literature identifier features; indicator word search relies on manually inducing a large number of citation indicator words (such as "a certain author points out", "a certain research shows") for extraction. The above methods rely heavily on the data rule features summarized manually. However, the citation content forms in literatures of different disciplinary fields vary greatly, and there are a large number of citation forms such as indirect paraphrases, making it difficult for these methods based on manually summarized rule features to identify citation types not covered in the manually summarized identifiers and having poor adaptability.

[0004] Shallow machine learning models can achieve automatic learning of the feature of the expression form of citation content through supervised training methods after a certain amount of citation content data annotation, thus effectively improving the recall rate of citation content in the literature. However, this method relies heavily on the scale and quality of citation content data annotation and has poor generalization ability on literatures in different fields and different types of citation data, and cannot effectively cope with complex and changeable citation scenarios.

[0005] Generative large language models represented by ChatGPT, trained based on massive data and ultra-large-scale network parameters, have powerful content understanding capabilities and can understand the complex sentence structures and semantic relationships in the literature. Just through prompt instructions or just giving various examples of citation content extraction, generative large language models can achieve global modeling of the text in the literature, better capture the context semantic information of citation content, and demonstrate powerful generalization ability, and can achieve high-precision and high-recall extraction of citation content of various citation types included in various literatures.

[0006] However, generative large language models also have deficiencies in extracting cited content from literature: on the one hand, there will be cases where words, symbols, or other content in individual sentences of the output cited content are not output as they are, resulting in an inability to ensure that all output cited content results are exactly the same as the original literature sentence content and an inability to ensure complete restoration; on the other hand, since it is necessary to output all recognized cited sentences in full, the speed of cited extraction becomes slower. Usually, in order to enable the generative large model to have better extraction effects, users will also provide complex and detailed instruction descriptions and cited content extraction examples, which also increases the input length of the model and affects the processing speed of cited content extraction. There is an urgent need for new methods to overcome these problems and improve the overall performance of extracting cited content from literature. Summary of the Invention

[0007] To solve the above problems in the prior art, that is, the problems of slow speed in the process of extracting cited content by generative large language models and the possible inconsistency of some words and symbols in the extracted cited content with the original literature content, the first aspect of the present invention proposes a method for extracting cited content from literature based on a generative large model, and the method includes the following steps: S100. Obtain the literature to be processed, and obtain a text paragraph set based on the literature to be processed; S200. For each text paragraph in the text paragraph set, add segmentation marker numbers sentence by sentence, and replace the content containing the cited fragment with a cited identification symbol based on a preset cited identification rule to obtain a set of converted text paragraphs; S300. Input the set of converted text paragraphs into a pre-fine-tuned cited extraction model, identify the sentences containing cited content, and extract the segmentation marker numbers; Wherein, the cited extraction model is obtained by fine-tuning the parameters of the generative large model based on pre-constructed model fine-tuning data, and the cited extraction model adds the cited identification symbol to the cited recognition rule; S400. After mapping the extracted segmentation marker numbers to the sentence position information of the literature to be processed, perform content extraction and output the final cited content.

[0008] In some preferred embodiments, the method for obtaining the set of converted text paragraphs is as follows: S210. Based on a preset cited identification rule, identify and extract the cited fragments with cited identification features in all text paragraphs, form a set of cited fragments, and record the position information of the corresponding text of the cited fragments; S220. Add segmentation marker numbers sentence by sentence in each text paragraph, and record the position information of the sentence corresponding to the segmentation marker numbers; S230. Replace the content containing reference fragments in the text paragraphs with added segmentation marker numbers with reference identification symbols to obtain a set of transformed text paragraphs.

[0009] In some preferred embodiments, the segmentation marker numbers include a sentence - start segmentation marker and a sentence - end segmentation marker.

[0010] In some preferred embodiments, the method of replacing the content containing reference fragments in the text paragraphs with added segmentation marker numbers with reference identification symbols is as follows: If a certain reference fragment and a sentence within a certain segmentation number belong to the same sentence, replace the sentence within the segmentation number with the reference identification symbol; If a certain reference fragment is part of a sentence within a certain segmentation number, directly replace the corresponding part of the content within the segmentation number with the reference identification symbol; If a certain reference fragment spans multiple segmentation numbers, replace the repeated part of the sentences corresponding to the multiple segmentation numbers with the reference identification symbol.

[0011] In some preferred embodiments, the method of extracting the segmentation marker numbers of the sentences containing the reference sentence set content and / or sentences containing reference identification symbols is as follows: Extract the sentence - start segmentation markers of the sentences containing reference identification symbols in the set of transformed text paragraphs; based on a matching algorithm, extract the sentence - start segmentation markers of the sentences similar to the reference sentence set content in the set of transformed text paragraphs; Output all the extracted sentence - start segmentation markers one by one; if no reference identification symbols are recognized and no sentences similar to the reference sentence set content are matched, output as None.

[0012] In some preferred embodiments, the method of fine - tuning the parameters of a generative large - model based on pre - constructed model fine - tuning data is as follows: A. Obtain various types of literature and obtain a set of text paragraphs based on the various types of literature; B. Given an extraction instruction, use the generative large - model to extract all the reference sentences containing reference content in the set of text paragraphs obtained in A to form a reference sentence set; C. For each text paragraph in the set of text paragraphs, obtain a set of transformed text paragraphs by the method of S200; D. Extract the segmentation marker numbers of the sentences containing the reference sentence set content and / or sentences containing reference identification symbols in the set of transformed text paragraphs obtained in C to obtain a set of segmentation marker numbers; E. Use the set of converted text paragraphs obtained by C as the input of the model, and the set of segmentation marker sequence numbers obtained by D as the output of the model to construct fine-tuning data pairs. Based on the fine-tuning data pairs, perform parameter fine-tuning on the generative large model to obtain a citation extraction model.

[0013] In some preferred embodiments, the matching algorithm is specifically as follows: Perform sentence complete character matching on the sentences in the set of converted text paragraphs and the set of citation sentences. If the matching is successful, extract the start segmentation marker of the sentence that matches the sentence. If the matching is not successful, calculate the similarity between the sentences in the set of converted text paragraphs and the set of citation sentences, and screen the similarity calculation results of each sentence to be processed. When the similarity is greater than a preset specific threshold, extract the start segmentation marker of the sentence.

[0014] In some preferred embodiments, when outputting the extracted segmentation marker sequence numbers, only keep one if there are duplicate segmentation marker sequence numbers.

[0015] In the second aspect of the present invention, a device for extracting literature citation content based on a generative large model is proposed. The device includes: A model construction module configured to perform parameter fine-tuning on the generative large model based on a pre-constructed model fine-tuning data pair to obtain a citation extraction model; the citation extraction model adds citation identification symbols to the citation recognition rules. A text acquisition module configured to acquire a literature to be processed and obtain a set of text paragraphs based on the literature to be processed. A format conversion module configured to add segmentation marker sequence numbers to the text paragraphs in the set of text paragraphs, and replace the content containing citation fragments with citation identification symbols based on a preset citation identification rule to obtain a set of converted text paragraphs. A marker extraction unit configured to input the set of converted text paragraphs into the citation extraction model, and use the citation extraction model to identify sentences containing citation content and extract segmentation marker sequence numbers. A citation display module configured to map the output segmentation marker sequence numbers to the sentence position information of the literature to be processed and then perform content extraction to output the final citation content.

[0016] In the third aspect of the present invention, an electronic device is proposed, including: At least one processor; and A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned method for extracting literature citation content based on a generative large model.

[0017] Advantages of the present invention: (1) The method of the present invention constructs data content containing special symbols for fine-tuning the large model, changes the input and output forms of the extracted reference content, reduces the input length of the large model, and thus improves the extraction speed and extraction effect; (2) Only by changing the input and output forms of the large model can the speed and effect of reference content extraction be improved, which is applicable to all types of generative large models; the input and output forms are simple, easy for the model to understand and implement model fine-tuning, and the expected effect can be achieved through the construction of a small amount of fine-tuning data; (3) The time consumption of the added steps is extremely low and can be ignored; after fine-tuning, the model input not only does not need to construct reference content extraction instructions, but also replaces multiple character contents with reference identification features with a single special symbol, greatly saving the input length and output length of the model, and effectively improving the reference extraction speed of the large model; (4) The sentence segmentation marker symbols output by the model can be accurately mapped to the original literature reference content, avoiding the possibility of word and symbol errors in the output of the large model before fine-tuning, and ensuring the correctness of reference content extraction. Description of the Drawings

[0018] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent: Figure 1 is a flowchart of a method for extracting literature reference content based on a generative large model in an embodiment of the present invention; Figure 2 is a flowchart of obtaining a reference extraction model by fine-tuning the parameters of a generative large model in an embodiment of the present invention. Detailed Embodiments

[0019] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0020] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0021] The present invention aims to improve the two deficiencies existing in the process of extracting cited content by generative large language models, namely, the relatively slow speed and the possible partial inconsistency of word symbols between the extracted cited content and the content of the original document. By changing the input and output forms of the large model, the speed and effect of extracting cited content are improved. While achieving the improvement of the speed of extracting cited content, it can also ensure that the extracted cited content is completely consistent with the content in the original document.

[0022] To more clearly illustrate the method for extracting cited content from documents based on a generative large model of the present invention, the following will elaborate on each step in the embodiments of the present invention in conjunction with Figure 1-2 elaborate on each step in the embodiments of the present invention.

[0023] A method for extracting cited content from documents based on a generative large model according to the first embodiment of the present invention includes steps S100 - S500, and each step is described in detail as follows: S100. Fine-tune the parameters of the generative large model based on the pre-constructed model fine-tuning data to obtain a citation extraction model; the citation extraction model adds citation identification symbols to the citation recognition rule and can quickly identify sentences containing cited content, including citation identification symbols.

[0024] Preferably, the method for fine-tuning the parameters of the generative large model based on the pre-constructed model fine-tuning data to obtain a citation extraction model is as follows: 1) Obtain various types of documents and obtain a set A of text paragraphs based on the various types of documents.

[0025] As an option, in this embodiment, the various types of documents are a set of various types of documents, which can be various electronic forms of papers, reports, monographs, patents, etc. The coverage range of the document content is not limited and can include various contents such as natural sciences and social sciences. All the documents in the document set are segmented according to conditions to obtain an academic text paragraph set A.

[0026] Further preferably, segmenting all the documents in the document set according to conditions can improve the data concurrent processing efficiency. The segmentation conditions can be according to natural paragraphs, logical paragraphs, or other methods, and the present invention does not make a limitation. As an option, it is also possible not to segment. Segmenting is beneficial for concurrent execution to speed up the processing speed.

[0027] 2) Given an extraction instruction P, use the generative large model to extract all the citation sentences containing cited content in the set of text paragraphs to form a set B of citation sentences; Among them, the citation content extraction instruction P given to the generative large model can be either a detailed extraction description, or a given set of clear reference examples for extracting citation content, or a combination of the two.

[0028] 3) For the academic text paragraphs A in set Ai Perform formal special symbol label conversion to obtain a set of converted text paragraphs as the input of the model.

[0029] To obtain the set of converted text paragraphs, the specific method is as follows: 301) Identify and extract the reference fragments with reference identification features in the text paragraph to form a set of reference fragments C, and record the position information of the corresponding text of the reference fragments.

[0030] Preferably, the reference identification rules include but are not limited to forms such as footnote endnote identification, quotation mark identification, and various literature identifications. Extract the reference fragments in A i that contain the above reference identification rules to obtain the set of reference fragments C and record the start and end position information of all reference fragments in the original text of A i The reference fragments in C can be a sentence containing the reference identification rules, a part of the content in a sentence containing the reference identification rules, or multiple sentences containing the reference identification rules. For example, taking the "" reference identification rule as an example, "" may span multiple sentences or only a part of the content of one sentence.

[0031] 302) Add segmentation mark serial numbers sentence by sentence in the text paragraph A i and record the position information of the sentence corresponding to the segmentation mark serial number.

[0032] Preferably, the segmentation mark serial numbers include a sentence start segmentation mark and a sentence end segmentation mark. For the purpose of realizing sentence semantic segmentation, it can be strictly divided according to punctuation marks or other more flexible division methods can be adopted.

[0033] As an option, in this embodiment, [Bn] is used as the sentence start part segmentation mark and [En] is used as the sentence end segmentation mark. The segmentation method is not limited here, and n is the sentence serial number included in A i Record the position information of the sentence represented by each segmentation serial number in A i For example: [B2](234:345) means that the start position of the second sentence in A i is 256 and the end position is 345.

[0034] 303) Replace the content containing reference fragments in the paragraph with added segmentation mark serial numbers with reference identification symbols to obtain the set of converted text paragraphs. The specific method is as follows: If a certain reference fragment in the set C and the sentence within a certain segmentation serial number [Bn][En] belong to the same sentence, directly replace the sentence within the segmentation serial number with [CI]; If a reference segment in set C belongs to a part of the sentences within a certain segmentation serial number [Bn][En], directly replace the content of this part within the segmentation serial number with [CI]; If a reference segment in set C spans multiple segmentation serial numbers [Bn] [En], then replace the overlapping parts of the multiple segmentation serial numbers [Bn] [En] with [CI].

[0035] The example is as follows: [B1]...[E1][B2]...[E2][B3][CI][E3][B4][CI][E4]; The two [CI]s may be two reference segments in set C, or may be a reference segment that spans multiple sentences.

[0036] 4) Extract the segmentation marker serial numbers of the sentences containing the reference sentence set content and / or the sentences containing the reference identification symbols in the converted text paragraph set to obtain the segmentation marker serial number set. The method is as follows: Extract the start segmentation marker of the sentences containing the reference identification symbols in the converted text paragraph set; Based on the matching algorithm, extract the start segmentation markers of the sentences in the converted text paragraph set that are similar to the content of the reference sentence set; output all the extracted start segmentation markers one by one; If no reference identification symbol is recognized and no sentence similar to the content of the reference sentence set is matched, the output is None.

[0037] Further preferably, the matching algorithm is specifically as follows: Perform a complete character match of the sentences in the converted text paragraph set and the reference sentence set. If the match is successful, extract the start segmentation marker of the sentence that matches this sentence; If the match is not successful, calculate the similarity of the sentences in the converted text paragraph set and the reference sentence set, screen the similarity calculation results of each sentence to be processed, and when the similarity is greater than a pre-set specific threshold, extract the start segmentation marker of this sentence.

[0038] Specifically, since there may be some characters, words, and symbols in sentence set B that are different from the original sentences, first perform a complete character match of the sentences to extract the sentence segmentation serial numbers. For the sentences that cannot be matched, obtain the sentence segmentation symbols through sentence similarity calculation. The similarity calculation method is not limited, and any method such as the common string length, word and character repetition ratio, vector similarity, etc. can be used, and only obtain the start segmentation serial numbers of the sentences whose similarity is greater than the specific threshold; Preferably, when all the extracted segmentation marker serial numbers are output, there may be duplicates between the content in set B and set C. If there are duplicate segmentation marker serial numbers, only keep one.

[0039] 5) Use the set of converted text paragraphs obtained in step 303) as the input of the model, and the set of segmentation marker sequence numbers obtained in step 4) as the output of the model to construct fine-tuning data pairs. Based on the fine-tuning data pairs, perform parameter fine-tuning on the generative large model to obtain a citation extraction model.

[0040] Preferably, in this embodiment, take a certain paper paragraph as various types of literature, obtain various types of literature, and construct model fine-tuning data based on various types of literature: Obtain a set of text paragraphs A i as follows: {AI can never and should never become an independent subject of rights. Technology should not be allowed to alienate, otherwise it will challenge human rights and dignity. As early as 1950, Asimov clearly put forward the First Law in his Three Laws of Robotics, that is, a robot may not injure a human being or, through inaction, allow a human being to come to harm. It is not difficult to see from the Second Law and the Third Law he set that they are for the functional setting and rights protection of robots. The First Law can be said to be the most basic principle in the development of robotics and is adopted by many researchers. It should also become the basic view of humans on robots or AI. The act of injuring a human being should be extended to the protection of intellectual products with the development of the times and technology [1] If the Second Law is regarded as the basis for positioning the function of AI to serve humans, then the Third Law is, to a certain extent, the basis for striving for rights for AI creations.} Obtain a set of citation sentences B as follows: {"As early as 1950, Asimov clearly put forward the First Law in his Three Laws of Robotics, that is, a robot may not injure a human being or, through inaction, allow a human being to come to harm."; "The First Law can be said to be the most basic principle in the development of robotics and is adopted by many researchers. It should also become the basic view of humans on robots or AI. The act of injuring a human being should be extended to the protection of intellectual products with the development of the times and technology [1] ."} Obtain a set of citation fragments C as follows: {"The First Law can be said to be the most basic principle in the development of robotics and is adopted by many researchers. It should also become the basic view of humans on robots or AI. The act of injuring a human being should be extended to the protection of intellectual products with the development of the times and technology [1] ."} Add segmentation sequence number markers [Bn] [En] to A i as follows: {

B1

E1

B2

E2

B3

E3

B4

CI

E4

B5

E5

B2

B4

B1

E1

B2

E2

B3

E3

B4

CI

E4

B5

E5

B2

B4

[0041] 5) Based on the above fine-tuning data pairs, perform parameter fine-tuning on the generative large model to obtain a reference extraction model.

[0042] Preferably, the model fine-tuning method can be performed in any way such as full parameter fine-tuning, only adjusting some parameters, or adding some small-scale trainable modules. After model fine-tuning, the original output form of the quoted sentence content is changed to the form of outputting the first part of the sentence segmentation number.

[0043] More preferably, the generative large model can be any model with the ability to extract reference content; and add multiple special symbols such as the segmentation symbols

B1

E1

B2

E2

CI

None

[0044] After the trained and fine-tuned citation extraction model adds citation identification symbols to the citation recognition rules, it can understand and recognize the meanings of preset segmentation symbols, special segmentation marker numbers, citation identification symbols, and the entire formalized special symbol label conversion method.

[0045] Step S200: Obtain the document to be processed, and based on the document to be processed, obtain the text paragraph set of the document to be processed; Based on the obtained document to be processed, the content in the document is segmented according to conditions to obtain the academic text paragraph set of the document to be processed. Segmentation can improve the data concurrency processing efficiency, and the segmentation conditions can be in accordance with natural paragraphs, logical paragraphs, or other methods.

[0046] Furthermore, in this embodiment, it is also possible not to perform segmentation.

[0047] Step S300: Perform formalized special symbol label conversion on each text paragraph in the text paragraph set of the document to be processed: Add segmentation marker numbers to the text paragraphs of the document to be processed sentence by sentence, and replace the content containing citation fragments with citation identification symbols based on the preset citation identification rules to obtain the text paragraph set of the converted document to be processed.

[0048] Preferably, the method for obtaining the text paragraph set of the converted document to be processed is as follows: S310: Based on the preset citation identification rules, identify and extract the citation fragments with citation identification features in the text paragraph to form a citation fragment set, and record the position information of the text corresponding to the citation fragments; S320: Add segmentation marker numbers to the text paragraph sentence by sentence, and record the position information of the sentences corresponding to the segmentation marker numbers; S330: Replace the content containing citation fragments in the paragraph with segmentation marker numbers added with citation identification symbols to obtain the converted text paragraph set.

[0049] Preferably, the method for replacing the content containing citation fragments in the text paragraph with segmentation marker numbers added is as follows: If a certain citation fragment and the sentence within a certain segmentation number belong to the same sentence, replace the sentence within the segmentation number with the citation identification symbol; If a certain citation fragment is part of the sentence within a certain segmentation number, directly replace the corresponding part of the content within the segmentation number with the citation identification symbol; If a certain citation fragment spans multiple segmentation numbers, replace the overlapping part of the sentences corresponding to the multiple segmentation numbers with the citation identification symbol.

[0050] Step S400: Input the text paragraph set of the to-be-processed document after conversion into a pre-fine-tuned citation extraction model to identify sentences containing citation content and extract segmentation marker numbers. The citation extraction model identifies all sentences containing citation content in the input text paragraph set. After training, the model incorporates citation identification symbols into the citation recognition rules. The identified sentences containing citation content include sentences with citation identification symbols, and then the segmentation marker numbers of all the identified sentences are extracted.

[0051] Step S500: Map the segmentation marker numbers to the sentence position information of the to-be-processed document and then perform content extraction to output the final citation content.

[0052] When the user inputs the text content to be processed, the text is first formally converted. The citation fragments that can be recognized using the citation identification rules are replaced with citation identification symbols, and corresponding segmentation marker numbers are added to the overall content. When the processed text undergoes all citation content discrimination by the citation extraction model, there is no need to repeatedly identify the content that has already been replaced with citation identification symbols. Additionally, multiple character contents with citation identification features are replaced with a single special symbol, greatly saving the input length of the model and improving the recognition speed. After all citation content is recognized, the corresponding segmentation marker numbers are output, and the original document citation content is obtained through position mapping, avoiding the possibility of word and symbol errors in the output of the large model before fine-tuning and ensuring the correctness of citation content extraction.

[0053] Although the various steps are described in the above sequential order in the above embodiments, those skilled in the art can understand that for the purpose of achieving the effects of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.

[0054] An apparatus for extracting literature citation content based on a generative large model according to the second embodiment of the present invention, the apparatus includes: A model construction module configured to perform parameter fine-tuning on the generative large model based on pre-constructed model fine-tuning data to obtain a citation extraction model; A text acquisition module configured to acquire the to-be-processed document and obtain a text paragraph set based on the to-be-processed document; A format conversion module configured to add segmentation marker numbers to the text paragraphs in the text paragraph set and replace the content containing citation fragments with citation identification symbols based on a preset citation identification rule to obtain a converted text paragraph set; A tagging extraction unit, configured to input the set of converted text paragraphs into a citation extraction model, and use the citation extraction model to identify sentences containing citation content and extract segmentation tag numbers; A citation display module, configured to map the output segmentation tag numbers to the sentence position information of the document to be processed, then perform content extraction, and output the final citation content; Among them, the method for constructing model fine-tuning data is A. Obtain various types of documents, and obtain a set of text paragraphs based on the various types of documents; B. Given an extraction instruction, use a generative large model to extract all citation sentences containing citation content in the set of text paragraphs obtained in A, and form a citation sentence set; C. For each text paragraph in the set of text paragraphs, obtain a set of converted text paragraphs through the method of S200; D. Extract the segmentation tag numbers of the sentences containing the content of the citation sentence set and / or the sentences containing citation identification symbols in the set of converted text paragraphs obtained in C, and obtain a set of segmentation tag numbers; E. Use the set of converted text paragraphs obtained in C as the input of the model, and the set of segmentation tag numbers obtained in D as the output of the model to construct a fine-tuning data pair.

[0055] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process and related explanations of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0056] It should be noted that the above-described device for extracting citation content from documents based on a generative large model only takes the above-mentioned division of each functional module as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing each module or step, and are not regarded as an improper limitation of the present invention.

[0057] An electronic device according to a third embodiment of the present invention includes: At least one processor; and A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned method for extracting citation content from documents based on a generative large model.

[0058] A computer-readable storage medium according to a fourth embodiment of the present invention, wherein the computer-readable storage medium stores computer instructions for being executed by a computer to implement the above-mentioned method for extracting literature citation content based on a generative large model.

[0059] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and related descriptions of the above-described electronic devices and computer-readable storage media can refer to the corresponding processes in the foregoing method embodiments and will not be repeated herein.

[0060] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0061] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0062] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0063] The terms "first", "second", etc. are used to distinguish similar objects and not to describe or indicate a particular order or sequence.

[0064] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to those process, method, article, or apparatus / device.

[0065] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A method for extracting literature citation content based on a generative large model, characterized in that: The following steps are involved: S100, obtaining a document to be processed, and obtaining a text paragraph set based on the document to be processed; S200, for each text paragraph in the text paragraph set, adding a segmentation mark serial number by sentence, and replacing the content containing the reference fragment with a reference identifier based on a preset reference identifier rule, to obtain a converted text paragraph set; S300, inputting the converted text paragraph set into a pre-adjusted citation extraction model, identifying sentences containing citation content and extracting segmentation mark numbers; Wherein, the citation extraction model is obtained by fine-tuning the parameters of the generative large model based on the pre-built model fine-tuning data, and the citation extraction model adds the citation identification symbol to the citation identification rule; S400: After mapping the extracted segmentation mark serial number to the sentence position information of the document to be processed, content extraction is performed to output the final reference content.

2. According to claim 1, a method for extracting reference content from literature based on a generative large model is characterized in that: Get the converted text paragraph set as follows: S210, based on a preset quotation identification rule, identifying and extracting quotation segments with quotation identification features in all text paragraphs, forming a quotation segment set, and recording position information of text corresponding to the quotation segments; S220, adding segmentation mark serial numbers by sentence in each text paragraph, and recording position information of sentences corresponding to the segmentation mark serial numbers; S230: Replace the content containing the quoted fragment in the text paragraph to which the segmentation mark serial number is added with the quoted identifier symbol to obtain a converted text paragraph set.

3. According to claim 2, a method for extracting reference content from literature based on a generative large model is characterized in that: The segmentation mark sequence number includes a sentence start segmentation mark and a sentence end segmentation mark.

4. According to claim 3, a method for extracting reference content from literature based on a generative large model is characterized in that: Replace the content containing the quoted fragment in the text paragraph with the quoted identifier symbol, the method is as follows: If a certain quoted segment and a sentence in a certain segmentation sequence number belong to the same sentence, the sentence in the segmentation sequence number is replaced with the quoted identifier symbol; If a certain quoted segment is part of a sentence in a certain segmentation number, directly replace the corresponding part of the content in the segmentation number with the quoted identifier; If a certain quoted segment spans multiple segmentation numbers, the parts of the sentences corresponding to the multiple segmentation numbers that are repeated with the quoted segment are replaced with the quoted identification symbol.

5. According to claim 2, a method for extracting reference content from literature based on a generative large model is characterized in that: Fine-tune the parameters of the generative large model based on the pre-built model fine-tuning data. The method is as follows: A. obtaining various types of documents, and obtaining a set of text paragraphs based on the various types of documents; B. Given an extraction instruction, use the generative large model to extract all quoted sentences containing quoted content in the text paragraph set obtained in A to form a quoted sentence set; C. For each text paragraph in the text paragraph set, obtaining a converted text paragraph set by the method of S200; D. extracting the segmentation mark serial numbers of the sentences containing the quoted sentence set content and / or the quoted identifier symbol from the converted text paragraph set obtained in C, and obtaining a segmentation mark serial number set; E. Use the converted text paragraph set obtained by C as the input of the model, and the segmentation mark number set obtained by D as the output of the model, construct a fine-tuning data pair, and fine-tune the parameters of the generative large model based on the fine-tuning data pair to obtain a reference extraction model.

6. According to claim 5, a method for extracting reference content from literature based on a generative large model is characterized in that: Extract the segmentation mark sequence number of the sentence containing the quoted sentence set content and\or the sentence containing the quoted identifier symbol, the method is: Extract the sentence start segmentation markers of the sentences containing the reference identifier symbol in the converted text paragraph set; extract the sentence start segmentation markers of the sentences with similar content to the reference sentence set in the converted text paragraph set based on the matching algorithm; Output all the extracted sentence segmentation tags one by one; If the reference identifier is not recognized and no sentence similar to the reference sentence set is matched, the output is None.

7. A method for extracting reference content from literature based on a generative large model according to claim 6, characterized in that: The matching algorithm is specifically: Perform sentence complete character matching on the converted text paragraph set and the sentences in the quoted sentence set. If the match is successful, extract the sentence start segmentation marker that matches the sentence. If the match is not successful, the similarity between the converted text paragraph set and the sentences in the quoted sentence set is calculated, and the similarity calculation results of each sentence to be processed are screened. When the similarity is greater than a preset specific threshold, the sentence beginning segmentation marker of the sentence is extracted.

8. The method for extracting reference content from literature based on a generative large model according to claim 1, characterized in that: When the extracted segmentation marker numbers are output, if there are duplicate segmentation marker numbers, only one is retained.

9. A document citation content extraction device based on a generative large model, according to a document citation content extraction method based on a generative large model according to any one of claims 1 to 8, characterized in that: The device includes: A model building module is configured to fine-tune the parameters of the generative large model based on the pre-built model fine-tuning data to obtain a reference extraction model; the reference extraction model adds the reference identification symbol to the reference identification rule; A text acquisition module, configured to acquire a document to be processed, and obtain a text paragraph set based on the document to be processed; a format conversion module configured to add segmentation mark serial numbers to the text paragraphs in the text paragraph set, and replace the content containing the reference fragment with the reference identification symbol based on a preset reference identification rule to obtain a converted text paragraph set; A tag extraction unit configured to input the converted text paragraph set into a citation extraction model, identify sentences containing citation content using the citation extraction model and extract segmentation tag serial numbers; The citation display module is configured to extract the content after mapping the output segmentation mark serial number to the sentence position information of the document to be processed, and output the final citation content.

10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement a method for extracting document citation content based on a generative large model as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Time word extraction method and device

    CN109657237A

  • Legal provision information recommendation system based on knowledge base and large model

    CN117370539A

  • Document generation method and device based on variable analysis, equipment and storage medium

    CN119129562A

  • Intelligent semantic analysis system and method based on large data model

    CN119443109A

  • Methods and systems for handling a document having content marked using one or more identifiers

    US20210274059A1