Credit labeling method

By extracting the layout information of the text to be marked, dividing the text into subtext, identifying the citation markers and types, and matching it with the citation description set, the problem of poor recognition effect of citation missing is solved, and efficient citation annotation is achieved.

CN120046607AActive Publication Date: 2025-05-27TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510526838.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing citation annotation methods have the problem of poor recognition of citation missing phenomena, and deep learning-based methods lack effective ways to obtain citation context features, resulting in poor automatic annotation effect and low efficiency.

Method used

A citation annotation method is proposed. By extracting the original annotation information based on the layout of the text to be marked, the text is divided into subtext, and the citation annotation symbols and types are identified in the subtext, and the citation description set is matched to identify the phenomenon of missing citations and mark them.

Benefits of technology

This method can effectively identify the phenomenon of missing citations, make up for the shortcomings of independent labeling, improve the efficiency and intelligence of citation labeling, and is suitable for various application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046607A_ABST
    Figure CN120046607A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of text processing, particularly relates to a citation labeling method, and aims to solve the problem of poor citation missing phenomenon recognition effect. The method comprises the following steps: segmenting a to-be-labeled text into a plurality of sub-texts and determining a position identifier of any sub-text; determining a reference type identifier as a first identifier under the condition that a reference tag is identified in any sub-text; under the condition that any sub-text is successfully matched with a reference description set comprising a plurality of reference description structures, the reference type identifier of the sub-text is determined as a second identifier, and the reference description structures are obtained by screening alternative reference description structures; the alternative reference description structure is obtained by performing reference description structure recognition on sub-texts of the to-be-labeled text; and based on the reference type identifier and the position identifier of each sub-text, performing reference labeling in the to-be-labeled text. All quotation information of the text can be labeled, and the quotation labeling efficiency and the intelligent level under various application scenes are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In literary works such as academic papers, citation marking of the research results, data, academic viewpoints, etc. of others can clarify the original source of this part of the content, show respect for previous academic research, and avoid plagiarism.

[0003] Existing citation marking methods are mainly based on authors' self-marking. For example, authors add citation marking symbols and reference content in papers, and there is a common situation of "using without citation" in many papers, that is, the phenomenon of missing citations.

[0004] Existing methods for citation marking based on deep learning lack an effective way to obtain the context features of citations, the automatic marking effect is poor, and the marking speed may be slow and the efficiency is low. Summary of the Invention

[0005] To solve the above problems in the prior art, that is, the problem of poor recognition effect of the missing citation phenomenon, the present application proposes a citation marking method, which includes: Step S10, based on the layout of the text to be marked, extract the original marking information of the text to be marked, where the original marking information includes citation marking symbols; Step S20, divide the text to be marked into multiple sub-texts, and determine the position identifier of any sub-text; Step S30, when a citation marking symbol is recognized in any sub-text, determine that the citation type identifier of any sub-text is the first identifier; Step S40, when any sub-text matches successfully with the citation description set, determine that the citation type identifier of any sub-text is the second identifier, where the citation description set is composed of multiple citation description structures, and the citation description structure is screened from the alternative citation description structures based on the citation type of the reference text and the number of alternative citation description structures. The reference text is the text including the original marking information for reference, and the alternative citation description structures are obtained by identifying the citation description structures of the sub-texts of multiple texts to be marked; Step S50, based on the citation type identifier and position identifier of each sub-text, perform citation marking in the text to be marked.

[0006] In some embodiments, the citation description structure is screened from the alternative citation description structures based on the citation type of the reference text and the number of alternative citation description structures, including: Obtain multiple reference texts, and determine the reference citation type identifier of any reference sub-text in the reference text; Based on the reference citation type identifier, divide the reference sub-text into a first reference sub-text and a second reference sub-text; Screen the target reference sub - text with alternative reference description structures in the first reference sub - text, and determine the hit index of the alternative reference description structure based on the reference type of the target reference sub - text; Count the number of alternative reference description structures included in the second reference sub - text, and determine the exclusion index of the alternative reference description structure based on the number; Based on the hit index and the exclusion index, determine the effective hit index of the alternative reference description structure; Screen the alternative reference description structure based on the effective hit index to obtain the reference description structure.

[0007] In some embodiments, determining the hit index of the alternative reference description structure based on the reference type of the target reference sub - text includes: Determine the weight coefficient corresponding to any reference type; Obtain the product of the weight coefficients of each reference type included in any target reference sub - text as the hit trust value of any target reference sub - text; Obtain the sum of the hit trust values of each target reference sub - text as the hit index.

[0008] In some embodiments, determining the exclusion index of the alternative reference description structure based on the number includes: Determine the exclusion trust value of the second reference sub - text; Obtain the sum of the exclusion trust values of each second reference sub - text as the exclusion index.

[0009] In some embodiments, determining the effective hit index of the alternative reference description structure based on the hit index and the exclusion index includes: Obtain the sum of the exclusion index and a preset value as the target exclusion index; Use the ratio of the hit index to the target exclusion index as the effective hit index.

[0010] In some embodiments, screening the alternative reference description structure based on the effective hit index to obtain the reference description structure includes: When the effective hit index is greater than or equal to the effective hit threshold, determine the alternative reference description structure as the reference description structure.

[0011] In some embodiments, dividing the reference sub - text into the first reference sub - text and the second reference sub - text based on the reference reference type identifier includes: When the reference reference type identifier is not empty, mark the reference sub - text as the first reference sub - text; When the reference reference type identifier is empty, mark the reference sub - text as the second reference sub - text.

[0012] In some embodiments, the alternative reference description structure is obtained by identifying the reference description structure for sub-texts of multiple texts to be annotated, including: Inputting multiple de-tagged sub-texts into a large model, and based on the prompt words for reference description structure identification, obtaining the reference fragment and output reference description structure of any de-tagged sub-text, where the de-tagged sub-text is the sub-text after removing the reference annotation symbol; In the case where the position where the position identifier of the sub-text is located overlaps with the position of the reference fragment, determining the output reference description structure as the alternative reference description structure.

[0013] In some embodiments, determining the position identifier of any sub-text includes: Determining the start character position and end character position of any sub-text; Combining the start character position and end character position into coordinates as the position identifier of any sub-text.

[0014] In some embodiments, based on the reference type identifier and position identifier of each sub-text, performing citation annotation in the text to be annotated, including: Locating the target position of the sub-text based on the position identifier; Performing corresponding-form citation annotation at the target position based on the reference type identifier.

[0015] Advantageous effects of the present application: (1) Extract the original annotation information based on the layout of the text to be annotated, and the citation information independently annotated by the author in the text to be annotated can be obtained; divide the text to be annotated into multiple sub-texts, and the whole text can be divided into multiple small sub-texts as the smallest unit for text processing, and determine the position identifier of any sub-text, and the position information of each sub-text can be determined for text positioning; when a citation annotation symbol is recognized in any sub-text, determine that the citation type identifier of any sub-text is the first identifier, and the citation content and citation type independently annotated by the author can be determined according to the original annotation information; match any sub-text with the citation description set, where the citation description set is composed of multiple citation description structures, and the citation description structure is screened from the alternative citation description structures based on the citation type of the reference text and the number of alternative citation description structures. The reference text is the text including the original annotation information for reference, and the alternative citation description structures are obtained by identifying the citation description structures in the sub-texts of multiple texts to be annotated. By matching any sub-text with the citation description structure, it is determined whether there is a citation type of the citation description structure. If the match is successful, it indicates that there is a citation of the citation description structure in any sub-text, and then determine that the corresponding citation type identifier is the second identifier, which can identify the phenomenon of missing citations and make up for the deficiency of independent annotation; based on the citation type identifiers and position identifiers of each sub-text, perform citation annotation in the text to be annotated, and the citation content independently annotated by the author and the unannotated citation content in the text to be annotated can be identified and annotated, and all citation information of the text can be clearly obtained and annotated, improving the efficiency and intelligent level of citation annotation in various application scenarios.

[0016] (2) This application can detect multiple types of citations at the same time. Citation recognition does not rely solely on the author's independent annotation, and can effectively identify the situation of "using but not citing". The content of citation annotation is comprehensive and the applicable range is wide. Brief Description of the Drawings

[0017] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives and advantages of this application will become more obvious: Figure 1 It is a flowchart of a citation annotation method provided by an embodiment of this application; Figure 2 It is a step flowchart of a citation annotation method provided by an embodiment of this application; Figure 3 It is a system block diagram of a citation annotation system provided by an embodiment of this application. Detailed Embodiments

[0018] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. In addition, it should be noted that for the sake of description, only the parts related to the relevant invention are shown in the drawings.

[0019] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0020] A citation marking method according to a first embodiment of the present application includes steps S10 - S50. As Figure 1 shown, the method includes: Step S10, based on the layout of the text to be marked, extract the original marking information of the text to be marked, where the original marking information includes citation marking symbols.

[0021] Optionally, the layout is the page format, specifically referring to the size of the format, the type area and the surrounding margins, the font, size, typesetting form of the text, the number of words, the arrangement position, and the typesetting methods of the table of contents, headings, notes, tables, figure names, figure captions, punctuation marks, running heads, page numbers, and page decorations, etc.

[0022] In the embodiments of the present application, the layout of the text to be marked at least includes the main text, footnotes, endnotes, references, etc.

[0023] Further, according to the layout of the text to be marked, extract the original marking information marked independently in the text to be marked, where at least the citation marking symbols are included in the original marking information.

[0024] In the embodiments of the present application, the citation marking symbol is a symbol or format used to indicate the citation source, such as parenthetical citation (author, year), citation number, footnote, endnote, abbreviation system (ibid.), etc.

[0025] In other embodiments, the original marking information may further include identification symbols, such as double quotes, for directly quoting the original text of the citation.

[0026] Step S20, divide the text to be marked into multiple sub - texts, and determine the position identifier and citation type identifier of any sub - text.

[0027] Optionally, the splitting method of the text to be marked can determine the delimiter according to the language. For example, Chinese can use full - width symbols such as “。”, “!”, “?” for splitting.

[0028] As an example, split the following fragment: {In that evening dyed orange - red by the setting sun, the streets of the small town were filled with a faint floral fragrance and the warm smell of freshly baked bread. A group of children were chasing and playing, and their laughter was clear and melodious, just like the most beautiful movement on a summer morning! Under the old pagoda tree by the roadside, an old man was leisurely fanning himself, and his eyes revealed deep nostalgia for the past years.} The text is segmented by a delimiter into 3 sub - texts, which are respectively: ① In that evening dyed orange - red by the setting sun, the streets of the small town were filled with a faint floral fragrance and the warm smell of freshly baked bread.

[0029] ② A group of children were chasing and playing, and their laughter was clear and melodious, just like the most beautiful movement on a summer morning! ③ Under the old pagoda tree by the roadside, an old man was leisurely fanning himself, and his eyes revealed deep nostalgia for the past years.

[0030] Optionally, by positioning the text, the position identifier of any sub - text can be obtained.

[0031] As a possible implementation, determine the starting character position B and the ending character position E of any sub - text; form the coordinates (B, E) with the starting character position and the ending character position as the position identifier of any sub - text.

[0032] For the above 3 sub - texts, sub - text ① contains 41 characters, sub - text ② contains 32 characters, and sub - text ③ contains 37 characters. Then the position identifiers (B, E) of the above 3 sub - texts are (0, 40), (41, 72), and (73, 109) respectively.

[0033] It should be noted that after step S20 is executed, steps S30 and S40 are executed in parallel.

[0034] Step S30, when a reference annotation symbol is recognized in any sub - text, determine that the reference type identifier of any sub - text is the first identifier.

[0035] Optionally, if a reference annotation symbol is recognized in any sub - text, the corresponding reference type can be determined and the corresponding reference type identifier can be obtained.

[0036] For example, when a reference annotation symbol of a footnote or endnote is recognized in any sub - text, it indicates that the sub - text includes a reference type with footnote or endnote label information; when a reference annotation symbol of square brackets or parentheses is recognized in any sub - text, it indicates that the sub - text includes a reference type of reference literature.

[0037] As a possible implementation, search for the position of the footnote or endnote label in any sub - text. When it is recognized that there is a footnote or endnote label, determine that the reference type is a footnote or endnote reference.

[0038] As an example, the text content included in any sub - text is as follows: {In that evening dyed orange - red by the setting sun, the streets of the small town were filled with a faint fragrance of flowers and the warm smell of freshly baked bread. A group of children were chasing and playing, and their laughter was clear and pleasant, just like the most beautiful movement in the early summer morning! Under the old pagoda tree by the roadside, an old man was leisurely fanning himself, and his eyes revealed deep nostalgia for the past years. 1 1Author 1: "Example Name 1", Example Publishing House, 2023.} According to the above text, the third sentence is a footnote, and its end contains a footnote symbol mark, so it can be determined that the citation type of the above text is footnote / endnote.

[0039] As a possible implementation, search any sub - text, locate all reference marks using square brackets, obtain the reference mark numbers, and retrieve them in the "References" of the text to be marked. If it exists, determine that the citation type is reference citation.

[0040] As another possible implementation, search any sub - text, locate all reference marks using parentheses, obtain the author name + year inside the parentheses, and retrieve them in the "References" of the text to be marked. If there is a reference with both the author name and year matching, determine that the citation type is reference citation.

[0041] As another possible implementation, obtain the full text of each reference in the text to be marked, compare any sub - text with each reference for similarity. If there is overlapping text, determine that the citation type is reference citation, and the overlapping position with each reference is the citation position.

[0042] As an example, the text content included in any sub - text is as follows: {In that evening dyed orange - red by the setting sun, the streets of the small town were filled with a faint fragrance of flowers and the warm smell of freshly baked bread [2]. A group of children were chasing and playing, and their laughter was clear and pleasant, just like the most beautiful movement in the early summer morning (Author 2. 2023). Under the old pagoda tree by the roadside, an old man was leisurely fanning himself, and his eyes revealed deep nostalgia for the past years.

[0043] References [1]Author 1. Example Name 1[J]. Example Publishing House, 2023, (12): 129 - 130. [2]Author 2. Example Name 2[J]. Example Publishing House, 2023, (12): 129 - 130. [3]Author 3. Example Name 3[J]. Example Publishing House, 2023, (12): 129 - 130.} It is understandable that both the first sentence and the second sentence contain established format references, namely square bracket references and parentheses references.

[0044] Furthermore, extract the full text of the third reference and compare it with the above original text. If the third sentence is highly similar to a certain fragment in the original text, it is determined as a reference citation.

[0045] In other embodiments, a knowledge base can also be established to retrieve any sub - text in the knowledge base. If there is similar data, the citation type is determined as a database citation.

[0046] Among them, the knowledge base includes common - sense texts such as axioms, concepts, policies, legal provisions, ancient poems, and masterpieces.

[0047] As a possible implementation, retrieve any sub - text in the knowledge base. If any sub - text overlaps in the knowledge base, determine the citation type of any sub - text as a knowledge - base citation.

[0048] In other embodiments, the citation type can also be determined by identifying identification symbols.

[0049] As a possible implementation, traverse each sub - text. If there are double quotes, determine the citation type as a marked citation.

[0050] As an example, the text content included in any sub - text is: {Economist XX pointed out in his research: "The rapid development and wide application of automation technology, although it has improved production efficiency to a certain extent, may also lead to the disappearance of some traditional occupations, thus triggering changes in the employment structure." This sentence deeply reveals the complex relationship between technological progress and the employment market.} By traversing the above text and identifying double quotes, determine that the citation type of the above text is a marked citation. Optionally, the citation type TYPE is distinguished by a citation - type identifier.

[0051] In the embodiments of the present application, the first identifier includes the citation - type identifiers for footnote and endnote citations, and the citation - type identifiers for reference citations.

[0052] As a possible implementation, query the citation - type identifier table to determine the citation - type identifier of any sub - text, where any citation type in the citation - type identifier table corresponds to a uniquely determined citation - type identifier.

[0053] As an example, the citation - type identifier for a footnote and endnote citation is a, and the citation - type identifier for a reference citation is b.

[0054] Further, in the embodiments of the present application, quadruples can be set for each sub - text to describe the citation annotation situation of each sub - text. Among them, the quadruple is expressed as (ID, (B, E), TYPE, MESSAGE), where ID is the unique flag bit of the sub - text in the text to be annotated, (B, E) is the position of the sub - text in the text to be annotated, TYPE is the identifier of the citation type existing in the sub - text, and MESSAGE is the supporting information of the citation type identifier TYPE.

[0055] As an example, search for the positions of footnote and endnote labels in the text to be annotated, locate the sub - text where the footnote and endnote are located, add the citation type identifier a to the TYPE of the quadruple corresponding to the sub - text, and add the footnote and endnote information to the MESSAGE of the quadruple corresponding to the sub - text.

[0056] For example, in the above footnote citation example, the third sentence is a footnote, with the ID recorded as 3, the position identifier as (73, 109), and the citation type identifier as a. The corresponding quadruple obtained is (3, (73, 109), [a], [Author 1: "Example Name 1", Example Publishing House, 2023.]).

[0057] As another example, search the text to be annotated, locate all the reference markers using square brackets, obtain the reference marker numbers, and retrieve them in the "References" of the text to be annotated. If they exist, add the citation type identifier b to the TYPE of the quadruple corresponding to the sub - text, and add the corresponding reference entry to the MESSAGE of the quadruple corresponding to the sub - text.

[0058] As another example, search the text to be annotated, locate all the reference markers using parentheses, obtain the author's name + year inside the parentheses, and retrieve them in the "References" of the text to be annotated. If there is a reference where both the author's name and year match, add the citation type identifier b to the TYPE of the quadruple corresponding to the sub - text, and add the corresponding reference entry to the MESSAGE of the quadruple corresponding to the sub - text.

[0059] As another example, obtain the full text of each reference in the text to be annotated, compare the similarity between the sub - text and each reference. If there is overlapping text, determine the overlapping position between the sub - text and each reference as the citation position, add the citation type identifier b to the TYPE of the quadruple corresponding to the sub - text, and add the corresponding reference entry to the MESSAGE of the quadruple corresponding to the sub - text.

[0060] For example, in the above reference citation example, the position identifiers (B, E) of the 3 sub-texts are (0, 40), (41, 72), and (73, 109) respectively, the corresponding IDs are 1, 2, and 3 respectively, and the citation type identifier is b. The corresponding quadruples obtained are (1, (0, 40), [b], [Author 1. Example Name 1 [J]. Example Publishing House, 2023, (12): 129 - 130.]), (2, (41, 72), [b], [Author 2. Example Name 2 [J]. Example Publishing House, 2023, (12): 129 - 130.]), (3, (73, 109), [b], [Author 3. Example Name 3 [J]. Example Publishing House, 2023, (12): 129 - 130.]).

[0061] In other embodiments, the citation type identifier of the database citation and the citation type identifier of the identification citation can also be recorded as the first identifier.

[0062] As an example, the citation type identifier for the knowledge base citation is c, and the citation type identifier for the identification citation is d.

[0063] As a possible implementation, a knowledge base is established, and any sub-text is retrieved in the knowledge base. If similar data exists, the citation type identifier c is added to the TYPE of the quadruple corresponding to the sub-text, and the hit knowledge base data information is added to the MESSAGE of the quadruple corresponding to the sub-text.

[0064] As another possible implementation, traverse the text to be annotated, obtain the position intervals of each pair of double quotes, compare the text content within the position intervals of each pair of double quotes with any sub-text. If there are overlapping positions, the citation type identifier d is added to the TYPE of the quadruple corresponding to the sub-text, and the position information of the overlapping positions relative to each sub-text is calculated and added to the MESSAGE of the sub-text where it is located.

[0065] For example, in the above annotation citation example, this fragment is divided into 2 sub-texts, and the division positions (B, E) are [0, 75] and [76, 100] respectively. The position interval of the double quotes is [(14, 75)]. Compare the two sub-texts with the double quote positions crosswise, and it is found that the first sub-text has an intersection with the double quote position interval. Then there is an identification citation for this sub-text, and the quadruple is updated to (1, (0, 75), [d], [(14, 75)]).

[0066] Step S40. When any sub - text matches successfully with the reference description set, determine that the reference type identifier of any sub - text is the second identifier. Herein, the reference description set is composed of multiple reference description structures, and the reference description structure is filtered from the alternative reference description structures based on the reference type of the reference text and the number of alternative reference description structures. The reference text is the text including the original annotation information for reference, and the alternative reference description structures are obtained by identifying the reference description structures for the sub - texts of multiple texts to be annotated.

[0067] Optionally, match any sub - text with the reference description set. If the match is successful, it can be determined that the reference type of the sub - text is reference description structure reference, and then determine that the reference type identifier of any sub - text is the second identifier.

[0068] In the embodiments of the present application, the second identifier includes the reference type identifier of the reference description structure reference, and the reference type identifier for describing the reference type of the text is e.

[0069] Among them, the alternative reference description structures are obtained through the following method: Input multiple de - tagged sub - texts into the large model, and based on the prompt words for reference description structure recognition, obtain the reference fragment and the output reference description structure of any de - tagged sub - text. Herein, the de - tagged sub - text is the sub - text after removing the reference annotation symbol. When the position where the position identifier of the sub - text is located overlaps with the position of the reference fragment, determine the output reference description structure as the alternative reference description structure.

[0070] As a possible implementation manner, obtain multiple texts to be annotated. Through annotation recognition and text segmentation, obtain multiple sub - texts and corresponding quadruples. Remove the reference annotation symbols of the sub - texts to obtain the unmarked texts, and re - obtain the position identifiers to establish the mapped quadruple (ID, (B1, E1), TYPE, MESSAGE).

[0071] As an example, a certain sub - text is: {On the one hand, digital technology has broken the limitations of traditional education in terms of time and space. For example, online education platforms enable students in remote areas to access high - quality educational resources, which to a certain extent narrows the educational gap between urban and rural areas and between regions. As pointed out by UNESCO in a report, "Online education provides new ways of learning for those who cannot obtain traditional educational opportunities" [2]. Through Internet technology, students can participate in course learning anytime and anywhere, interact with teachers and classmates around the world, and greatly broaden the horizons and channels of learning.} The quadruple corresponding to this sub - text can be represented by a quadruple table, as shown in Table 1.

[0072] Table 1

[0073] The de-tagged sub-text after removing the reference markers is: On the one hand, digital technology has broken the limitations of traditional education in terms of time and space. For example, online education platforms enable students in remote areas to access high-quality educational resources, which to a certain extent narrows the educational gap between urban and rural areas and between regions. As pointed out by UNESCO in a report, online education provides new ways of learning for those who cannot obtain traditional educational opportunities. Through Internet technology, students can participate in course learning anytime and anywhere, interact with teachers and classmates around the world, and greatly broaden the horizons and channels of learning.

[0074] Since the two explicit reference markers, double quotes and square brackets, are removed, the character positions change, generating new quadruples, which can be represented by mapping the quadruple table, as shown in Table 2.

[0075] Table 2

[0076] Furthermore, apply the prompt words and use the large model text generation technology to obtain the reference fragments within the de-tagged sub-text.

[0077] As an example, the prompt words used are: {Determine whether there is a reference in the following text. If there is, output the full text of the reference fragment and give the characteristics of the original text that determine it as a reference On the one hand, digital technology has broken the limitations of traditional education in terms of time and space. For example, online education platforms enable students in remote areas to access high-quality educational resources, which to a certain extent narrows the educational gap between urban and rural areas and between regions. As pointed out by UNESCO in a report, online education provides new ways of learning for those who cannot obtain traditional educational opportunities. Through Internet technology, students can participate in course learning anytime and anywhere, interact with teachers and classmates around the world, and greatly broaden the horizons and channels of learning.} Among them, is used to divide the instruction and the content to strengthen the understanding of the large model prompt words. In the above prompt words, divides the instruction and the text that the large model needs to judge.

[0078] The result returned by the large model is: There is a reference in this text passage, and the full text of the reference fragment is: "As pointed out by UNESCO in a report, online education provides new ways of learning for those who cannot obtain traditional educational opportunities." The characteristics of the original text that determine it as a reference are: "...... pointed out in......".

[0079] It should be noted that there are no obvious reference identifiers marked by users in the de-tagged sub-text, such as double quotes and square brackets. The reference recognition result is obtained by the large model through pre-training and fine-tuning with a large amount of data. This method has a slow judgment speed, is not applicable to high-concurrency, real-time, and massive call processing, and has extremely high economic costs and cannot be implemented. Therefore, the large model cannot be directly used for reference annotation.

[0080] On the other hand, if the reference markers in the sub-text are not removed, such as double quotes "" and square brackets [2], the large model will give these two explicit reference features based on established rules and ignore the hidden unknown description features. Therefore, in order to ensure the comprehensiveness of reference recognition, the input to the large model is the de-tagged sub-text.

[0081] Furthermore, determine the position of the reference segment generated by the large model in the de-tagged sub-text and compare it with each new quadruple. When the position identified by the position identifier in the new quadruple overlaps with the position of the reference segment, determine that the output reference description structure is the alternative reference description structure.

[0082] Taking the recognition result of the above large model as an example, the reference segment in the de-tagged sub-text is: "As the United Nations... new approach." Its position in the tagged sub-text is (84, 132), and the position identified by the position identifier in the new quadruple is (84, 132). Since the two positions coincide, the reference feature of this segment, "…… pointed out in ……", is used as the alternative reference description structure.

[0083] It can be understood that in this way, the implicit reference description features are mined. At the same time, since this sentence is indeed a reference, the credibility probability that this structure is a reference expression is enhanced. However, this is only an individual case, and it cannot be determined as a reference description structure based on a single judgment of the large model. Therefore, it is used as an alternative reference description structure and then screened and verified through subsequent methods.

[0084] In the embodiments of the present application, by obtaining multiple texts including original annotation information as reference texts, and based on the reference type and the number of alternative reference description structures of the reference texts, the alternative reference description structures are screened to obtain reference structure descriptions, which form a reference description set.

[0085] Please refer to Figure 2 , the specific steps for obtaining the reference structure description include: Step S401, obtain multiple reference texts and determine the reference reference type identifier of any reference sub-text in the reference texts.

[0086] Optionally, obtain a large collection of papers. Since the papers all include original annotation information, any paper can be used as a reference text. By performing text segmentation on the reference text, reference sub-texts are generated, and quadruples are constructed for each reference sub-text to determine the reference citation type identifier of any reference sub-text.

[0087] Step S402: Based on the reference citation type identifier, divide the reference sub-texts into first reference sub-texts and second reference sub-texts.

[0088] Optionally, when the reference citation type identifier is not empty, mark the reference sub-text as a first reference sub-text; when the reference citation type identifier is empty, mark the reference sub-text as a second reference sub-text.

[0089] It can be understood that the first reference sub-text has obvious citation identifiers, that is, the reference type is one or more of a, b, c, and d; there are no obvious citation identifiers in the second reference sub-text, and it is initially identified that there is no citation.

[0090] Step S403: Screen out target reference sub-texts with alternative citation description structures in the first reference sub-texts, and determine the hit index of the alternative citation description structures based on the citation types of the target reference sub-texts.

[0091] Optionally, to determine whether there is a citation description structure in the first reference sub-text, multiple methods can be used. For example, regular expression matching: converting "……pointed out in……" into a regular expression: ".*?pointed out in.*?".

[0092] It should be noted that by using regular expression matching to screen target reference sub-texts in the first reference sub-texts, the recognition by the large model (the single recognition time-consuming is above seconds) is converted into a simple regular expression (below ms level).

[0093] Furthermore, determine the weight coefficient corresponding to any citation type; obtain the product of the weight coefficients of each citation type included in any target reference sub-text as the hit confidence value of any target reference sub-text; obtain the sum of the hit confidence values of each target reference sub-text as the hit index.

[0094] Optionally, assign corresponding weight coefficients to each citation type. The value of the weight coefficient is greater than 0 and is set according to prior knowledge. For example, in terms of experience, when using double quotes for citation, the possibility of a fixed sentence pattern of citation description appearing in front is greater, so a high weight is assigned to this citation type identifier.

[0095] In the embodiments of the present application, the weight coefficient corresponding to footnote and endnote references (i.e., the reference type identifier is a) is set to 2, the weight coefficient corresponding to reference literature references (i.e., the reference type identifier is b) is set to 2, the weight coefficient corresponding to knowledge base references (i.e., the reference type identifier is c) is set to 4, and the weight coefficient corresponding to identifier references (i.e., the reference type identifier is d) is set to 6.

[0096] It should be noted that the setting of the above weight coefficients is only an example. In other embodiments, it can also be set to other values, and can also be optimized to values that can be dynamically adjusted, which is not limited in the present application.

[0097] Optionally, the hit trust value v can be determined by the following formula:

[0098] where represents the weight coefficient with the reference type identifier t.

[0099] It can be understood that the hit trust value corresponding to the first reference sub-text without an alternative reference description structure is 0.

[0100] As an example, a certain target reference sub-text includes two reference types, b and d, and the corresponding hit trust values .

[0101] It can be understood that in the case of coexistence of multiple reference types, the credibility of the existing reference description structure in the sentence is increased, and the credibility is rewarded by calculating the product.

[0102] Further, calculate the sum of the hit trust values of each target reference sub-text as the hit index of the alternative reference description structure.

[0103] It can be understood that the first reference sub-text is the sub-text where a reference is determined to exist. The more the alternative reference description structure hits in the first reference sub-text, the higher the hit index, indicating that the alternative reference description structure is more credible as a reference description structure and is more likely to be a reference description structure.

[0104] Step S404, count the number of alternative reference description structures included in the second reference sub-text, and determine the exclusion index of the alternative reference description structure based on the number.

[0105] Optionally, determine the exclusion trust value of the second reference sub-text; obtain the sum of the exclusion trust values of each second reference sub-text as the exclusion index.

[0106] In an embodiment of the present application, the exclusion trust value of any reference sub - text is set to 1, and the sum of the exclusion trust values of each second reference sub - text, that is, the number of second reference sub - texts with an alternative reference description structure, is obtained as the exclusion index.

[0107] In other embodiments, the exclusion trust value can also be set to other values, which are not limited herein.

[0108] It can be understood that there is no obvious reference identifier in the second reference sub - text, which is a sub - text without a reference. If an alternative reference description structure is recognized in the second reference sub - text, the credibility of the alternative reference description structure will be reduced. The more second reference sub - texts in which an alternative reference description structure is recognized, the greater the degree of reduction in the corresponding credibility. Therefore, the exclusion index is determined by counting the number of alternative reference description structures included in the second reference sub - text.

[0109] Step S405: Based on the hit index and the exclusion index, determine the effective hit index of the alternative reference description structure.

[0110] Optionally, obtain the sum of the exclusion index and a preset value as the target exclusion index, and use the ratio of the hit index to the target exclusion index as the effective hit index.

[0111] As an example, the effective hit index R can be determined by the following formula:

[0112] Wherein, represents the hit index, represents the exclusion index, 1 is the preset value set in the embodiment of the present application, represents the target exclusion index.

[0113] It can be understood that by adding the preset value, the target exclusion index is obtained, which can avoid the risk of the denominator being 0.

[0114] By calculating the ratio of the hit index to the target exclusion index, the credibility of the alternative reference description index is quantified. The more alternative reference description structures are recognized in the first reference sub - text with reference content, the higher the credibility of the alternative reference description structure for identifying the reference content. On the contrary, the fewer alternative reference description structures are recognized in the second reference sub - text without reference content, the higher the credibility of the alternative reference description structure for identifying the reference content.

[0115] Step S406: Based on the effective hit index, screen the alternative reference description structure to obtain the reference description structure.

[0116] Optionally, when the effective hit index is greater than or equal to the effective hit threshold, determine that the alternative reference description structure is a reference description structure.

[0117] As an example, the effective hit threshold is set to 5. In other embodiments, the effective hit threshold can also be set according to usage, which is not limited herein.

[0118] It can be understood that after screening by the effective hit threshold, each reference description structure has passed the statistical verification based on the large model and massive data, greatly improving the credibility of the description structure.

[0119] Optionally, after determining the reference description structure, a reference description set including all reference description structures is constructed.

[0120] It can be understood that by constructing the reference description set, any sub - text is matched with the reference description structures in the reference description set, enabling reference recognition for sub - texts without obvious markings and making the application scenarios of reference recognition more comprehensive.

[0121] Step S50: Based on the reference type identifier and position identifier of each sub - text, perform citation annotation in the text to be annotated.

[0122] Optionally, by determining the reference type identifier of each sub - text, the reference type of each sub - text can be symbolized. After determining the position identifier of each sub - text, the position of each sub - text in the text to be annotated can be confirmed. Further, based on the reference type and position information of the sub - text, citation annotation can be performed on each sub - text.

[0123] As a possible implementation, locate the target position of the sub - text based on the position identifier; perform corresponding - form citation annotation at the target position based on the reference type identifier.

[0124] In the embodiments of the present application, the citation annotation forms for different reference types are different. For example, the sub - text with any reference type identifier is highlighted at the target position with a uniquely corresponding color.

[0125] It should be noted that in other embodiments, the citation annotation forms for different reference types can be the same. For example, the identified citations are all displayed in a unified annotation format at the target position.

[0126] Although the above - mentioned steps are described in the above - mentioned sequential order in the above - mentioned embodiments, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present application.

[0127] The embodiment of the present application extracts the original annotation information based on the layout of the text to be annotated, and can obtain the citation information independently annotated by the author in the text to be annotated; divides the text to be annotated into multiple sub-texts, can divide the whole text into multiple small sub-texts as the smallest unit for text processing, and determines the position identifier of any sub-text, can determine the position information of each sub-text for text positioning; when a citation annotation symbol is recognized in any sub-text, determines that the citation type identifier of any sub-text is the first identifier, can determine the citation content and citation type independently annotated by the author according to the original annotation information; matches any sub-text with the citation description set, where the citation description set is composed of multiple citation description structures, and the citation description structure is screened from the alternative citation description structures based on the citation type of the reference text and the number of alternative citation description structures. The reference text is the text including the original annotation information for reference, and the alternative citation description structures are obtained by identifying the citation description structures in the sub-texts of multiple texts to be annotated. By matching any sub-text with the citation description structure, it is determined whether there is a citation type of the citation description structure. If the match is successful, it indicates that there is a citation of the citation description structure in any sub-text, and then determines that the corresponding citation type identifier is the second identifier, can identify the phenomenon of missing citations, and makes up for the deficiency of independent annotation; based on the citation type identifiers and position identifiers of each sub-text, performs citation annotation in the text to be annotated, can identify and annotate both the citation content independently annotated by the author and the unannotated citation content in the text to be annotated, can clearly obtain all the citation information of the text and annotate it, and improves the efficiency and intelligent level of citation annotation in various application scenarios.

[0128] The citation annotation system according to the second embodiment of the present application, as Figure 3 shown, includes an extraction module 100, a segmentation module 200, a matching module 300, a generation module 400, and a citation annotation module 500.

[0129] The extraction module 100 is used to extract the original annotation information of the text to be annotated based on the layout of the text to be annotated, where the original annotation information includes citation annotation symbols; The segmentation module 200 is used to divide the text to be annotated into multiple sub-texts and determine the position identifier of any sub-text; The first identifier determination module 300 is used to determine that the citation type identifier of any sub-text is the first identifier when a citation annotation symbol is recognized in any sub-text; The second identifier determination module 400 is configured to determine that the reference type identifier of any sub - text is the second identifier when any sub - text successfully matches the reference description set. The reference description set is composed of multiple reference description structures, and the reference description structures are filtered from the alternative reference description structures based on the reference type of the reference text and the number of alternative reference description structures. The reference text is the text including the original annotation information for reference, and the alternative reference description structures are obtained by identifying the reference description structures for the sub - texts of multiple texts to be annotated; The citation annotation module 500 is configured to perform citation annotation in the text to be annotated based on the reference type identifier and the position identifier of each sub - text.

[0130] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process and related descriptions of the above - described system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0131] It should be noted that the citation annotation system provided in the above embodiments is only illustrated by the division of the above - mentioned functional modules. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the modules or steps in the embodiments of the present application can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub - modules to complete all or part of the functions described above. For the names of the modules and steps involved in the embodiments of the present application, they are only used to distinguish each module or step, and are not regarded as an improper limitation of the present application.

[0132] An electronic device according to the third embodiment of the present application includes: At least one processor; and A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above - mentioned citation annotation method.

[0133] A computer - readable storage medium according to the fourth embodiment of the present application, the computer - readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above - mentioned citation annotation method.

[0134] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process and related descriptions of the above - described electronic device and computer - readable storage medium can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0135] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0136] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order from that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0138] The terms "first", "second", etc. are used to distinguish similar objects and are not used to describe or indicate a specific order or sequence.

[0139] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to those processes, methods, articles, or apparatus / devices.

[0140] So far, the technical solution of the present application has been described in connection with the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the protection scope of the present application is obviously not limited to these specific embodiments. Without departing from the principle of the present application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present application.

Claims

1. A citation annotation method, characterized in that: include: Step S10, extracting original annotation information of the text to be annotated based on the layout of the text to be annotated, wherein the original annotation information includes a reference annotation symbol; Step S20, dividing the text to be annotated into a plurality of sub-texts, and determining a position mark of any sub-text; Step S30, when the reference marker is identified in any of the sub-texts, determining the reference type identifier of the any of the sub-texts as a first identifier; Step S40, when any of the sub-texts successfully matches the reference description set, determining the reference type identifier of the any of the sub-texts as a second identifier, wherein the reference description set is composed of a plurality of reference description structures, the reference description structure is obtained by screening the candidate reference description structures based on the reference type of the reference text and the number of the candidate reference description structures, the reference text is a text including original annotation information for reference, and the candidate reference description structure is obtained by performing reference description structure recognition on the sub-texts of the plurality of texts to be annotated; Step S50: marking citations in the text to be marked based on the reference type identifier and the position identifier of each sub-text.

2. A citation marking method according to claim 1, characterized in that: The reference description structure is obtained by screening the candidate reference description structures based on the reference type of the reference text and the number of the candidate reference description structures, and includes: Acquire multiple reference texts, and determine a reference type identifier of any reference subtext in the reference texts; Based on the reference type identifier, dividing the reference subtext into a first reference subtext and a second reference subtext; Screening the first reference subtext for a target reference subtext containing the candidate reference description structure, and determining a hit index of the candidate reference description structure based on the reference type of the target reference subtext; Counting the number of the candidate reference description structures included in the second reference subtext, and determining an exclusion index of the candidate reference description structure based on the number; Determining an effective hit index of the candidate reference description structure based on the hit index and the exclusion index; The candidate reference description structure is screened based on the effective hit index to obtain the reference description structure.

3. A citation marking method according to claim 2, characterized in that: The step of determining a hit index of the candidate reference description structure based on the reference type of the target reference subtext includes: Determine the weight coefficient corresponding to any reference type; Obtaining the product of weight coefficients of each reference type included in any target reference subtext as a hit confidence value of the any target reference subtext; The sum of the hit confidence values ​​of each target reference subtext is obtained as the hit index.

4. A citation marking method according to claim 2, characterized in that: The determining of the exclusion index of the candidate reference description structure based on the quantity includes: determining an exclusion confidence value of the second reference subtext; The sum of the exclusion confidence values ​​of the second reference subtexts is obtained as the exclusion index.

5. A citation marking method according to claim 2, characterized in that: The determining, based on the hit index and the exclusion index, the effective hit index of the candidate reference description structure comprises: Obtaining the sum of the exclusion index and a preset value as a target exclusion index; The ratio of the hit index to the target exclusion index is used as the effective hit index.

6. A citation marking method according to claim 2, characterized in that: The step of screening the candidate reference description structure based on the effective hit index to obtain the reference description structure includes: In a case where the valid hit index is greater than or equal to a valid hit threshold, the candidate reference description structure is determined to be the reference description structure.

7. A citation marking method according to claim 2, characterized in that: The dividing the reference subtext into a first reference subtext and a second reference subtext based on the reference type identifier comprises: In the case where the reference type identifier is not empty, marking the reference subtext as the first reference subtext; When the reference type identifier is empty, the reference subtext is marked as a second reference subtext.

8. A citation marking method according to claim 1, characterized in that: The candidate reference description structure is obtained by identifying the reference description structure of subtexts of multiple texts to be annotated, including: Inputting a plurality of unmarked subtexts into the large model, obtaining a reference segment of any unmarked subtext and outputting a reference description structure based on the prompt words recognized by the reference description structure, wherein the unmarked subtext is a subtext without the reference marker; In the case where the position located by the position identifier of the subtext overlaps with the position of the reference segment, the output reference description structure is determined to be an alternative reference description structure.

9. A citation marking method according to claim 1, characterized in that: The step of determining the position mark of any subtext includes: Determine the first character position and the last character position of any subtext; The first character position and the last character position are combined into coordinates as a position identifier of any subtext.

10. A citation marking method according to claim 1, characterized in that: The step of marking citations in the text to be marked based on the reference type identifier and the position identifier of each sub-text includes: Locating a target position of the subtext based on the position identifier; Based on the reference type identifier, a corresponding form of citation annotation is performed at the target location.

Citation Information

Patent Citations

  • Automatic quotation extraction method and device with semantic integrity kept

    CN104050158A

  • Information entropy-based quotation recommendation method, device and terminal

    CN117076658A

  • Periodical metering method and device, electronic equipment and storage medium

    CN117972119A

  • Document generation method and device, storage medium and electronic equipment

    CN118485065A

  • Comprehensive measurement method for reusability of data publication

    CN118897986A