Deep learning-based Chinese quotation description item automatic indexing method, system and equipment
Through the combination of deep learning models and preset rules, the accuracy and efficiency of automatic citation of Chinese citation entries is solved, and seamless conversion from original citations to structured entries is achieved, improving the citation recall and accuracy, and reducing dependence on the manual rule base.
Patent Information
- Application Number
- CN202510528505.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, automatic citation of Chinese citations entries relies on manual rules, resulting in high maintenance costs, poor generalization, low accuracy, and lack of effective deep learning research methods.
A deep learning-based method is adopted to count the proportion of each entry type, extract the training corpus from the manual index data, convert it into a single-character sequence format, and use the context semantic encoding module and the label sequence optimization module to build a deep learning model, and combine preset rules for indexing, including splitting the data balance strategy of combining most corpus and minority corpus and label correction of specific strings.
It realizes automatic quotation of efficient Chinese citations without relying on manual rules, improves the quotation recall and accuracy, adapts to diversified citation formats, reduces dependence on the manual rule base, and enhances the migration and application capabilities in cross-domain scenarios.
Smart Images

Figure CN120471020A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information indexing, and specifically relates to a method, system and device for automatic indexing of Chinese citation entries based on deep learning. Background Art
[0002] Automatic indexing of Chinese citation entries is to mark each entry in Chinese citations (author, title, source, edition description, year, volume number, issue number, page number, place of publication, publisher, parent document source, thesis awarding institution, editor-in-chief, to be published, etc.) to improve the accuracy and completeness of citation links and enrich the knowledge network node citation network.
[0003] In existing technologies, most Chinese citation entry indexing is based on manual rule matching. These methods extract fields based on regular expressions or fixed templates (e.g., "Author: {content}," "Publication Year: {number}"). However, Chinese citation formats are complex and varied (e.g., abbreviations, mixed punctuation, and multiple languages), requiring a large number of customized rules for different scenarios. This leads to high maintenance costs and poor generalization, resulting in low indexing efficiency, poor scalability, and limited accuracy.
[0004] Currently, no research is available on the automated indexing of Chinese citation entries. This task falls under the umbrella of natural language processing, where common strategies include rules, statistics, and deep learning. Currently, deep learning algorithms have demonstrated excellent performance in a variety of natural language processing applications. Their advantage lies in achieving excellent results without the need to summarize literal rules.
[0005] Based on this, the present invention proposes a method, system and device for automatic indexing of Chinese citation entries based on deep learning. Summary of the Invention
[0006] In order to solve the above-mentioned problems in the prior art, namely, the lack of effective research methods and mature technical solutions for the automatic indexing of Chinese citation entries, especially the problem of how to use natural language processing technologies such as deep learning to efficiently identify and annotate complex and diverse Chinese citation entry information without relying on manual rules, the present invention provides a method, system and device for automatic indexing of Chinese citation entries based on deep learning.
[0007] In a first aspect of the present invention, a method for automatic indexing of Chinese citation entries based on deep learning is proposed, the method comprising:
[0008] According to the Chinese citation item indexing standards, the proportion of each item type is counted, and training corpus is extracted from the manually indexed data according to the said proportion;
[0009] Converting the training corpus into a single-character sequence format and adding a label to each character as input data; the label includes a combination of a start and end mark and a record type;
[0010] Input the input data into a pre-built and trained deep learning model for training, wherein the deep learning model is constructed based on a combination of a contextual semantic encoding module and a label sequence optimization module;
[0011] The citation data to be indexed is converted into a new format and then fed into a trained deep learning model for prediction. This generates a sequence of entries and merges them into complete entries based on the labels.
[0012] The merged complete entry is modified according to preset rules, wherein the preset rules include: determining the entry type according to the citation context characteristics, and modifying the label of the specific character string based on the label mapping rule table.
[0013] Furthermore, after extracting the training corpus according to the proportion, the following is also included:
[0014] The training dataset is balanced by splitting the majority class corpus and merging it with the minority class corpus.
[0015] Furthermore, the majority class corpus is split and merged with the minority class corpus, specifically including:
[0016] Calculate the ratio of the sample size of the minority class to the sample size of the majority class in all categories of bibliographic items. When the ratio is less than a preset percentage, split the training corpus based on preset rules.
[0017] The preset rule is: set the number of splits k according to the business volume, and split the majority class samples into k subsets according to the set sampling method;
[0018] Each majority class subset is merged with all minority class samples to generate k balanced sub-datasets, and the deep learning model is trained based on the k balanced sub-datasets.
[0019] Furthermore, the contextual semantic encoding module is used to extract bidirectional context features from the character sequence;
[0020] The label sequence optimization module is used to learn the transition probability between labels based on context features and generate a globally optimal label sequence.
[0021] Furthermore, the start and end marks include B to mark the start of the entry, I to mark the middle position of the entry, E to mark the end of the entry, and S to mark the entry as a single word.
[0022] Furthermore, the entry types include at least author, title, source, year, edition description, volume number, issue number, page number, place of publication, publisher, parent document source, thesis awarding institution, editor-in-chief and to be published.
[0023] Furthermore, the label of a specific string is modified based on the entry type and the preset label mapping rules, specifically:
[0024] Modify tags based on preset string features, including keyword recognition, legal document determination, and adjacent tag association rules;
[0025] Based on the document type identification results, labels are modified according to the specific format features of conference papers, patents, standards, and journals;
[0026] For strings containing year-number structures, differentiated corrections are implemented based on the number arrangement format and value range;
[0027] Perform tag conflict detection across document types and perform tag correction based on institution suffixes, document identifiers, and special symbols;
[0028] Among them, the document types include at least conference papers, patents, standards, journals, books, dissertations, reports, archives, newspapers, case judgments, laws and regulations, and yearbooks.
[0029] Another aspect of the present invention provides a deep learning-based automatic indexing system for Chinese citation entries, based on a deep learning-based automatic indexing method for Chinese citation entries. The system includes:
[0030] A corpus generation module is configured to calculate the proportion of each type of item according to the Chinese citation item indexing standard, and extract training corpus according to the proportion from the manually indexed data;
[0031] a format conversion module configured to convert the training corpus into a single-character sequence format and add a label to each character as input data; the label includes a combination of a start and end mark and a record type;
[0032] A training module configured to input the input data into a pre-built and trained deep learning model for training, wherein the deep learning model is constructed based on a contextual semantic encoding module and a label sequence optimization module;
[0033] A prediction module is configured to convert the format of the citation data to be indexed and then input it into a trained deep learning model for prediction, thereby obtaining a sequence of entries and merging them into complete entries according to the labels;
[0034] A correction module is configured to correct the merged complete entry according to preset rules, wherein the preset rules include: determining the entry type based on the citation context characteristics, and correcting the label of the specific string based on the entry type and preset label mapping rules.
[0035] A third aspect of the present invention provides an electronic device, comprising:
[0036] at least one processor; and
[0037] a memory communicatively connected to at least one of the processors; wherein,
[0038] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned method for automatic indexing of Chinese citation entries based on deep learning.
[0039] In a fourth aspect of the present invention, a computer-readable storage medium is proposed, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned method for automatic indexing of Chinese citation entries based on deep learning.
[0040] Beneficial effects of the present invention:
[0041] (1) By deeply encoding the contextual semantics of Chinese citations through a deep learning model, the problem that traditional rule-based methods are difficult to cover complex text patterns is effectively solved, the adaptability to diverse citation formats is significantly improved, and the indexing recall rate has achieved a fundamental breakthrough.
[0042] (2) The corpus extraction strategy based on statistical proportions ensures that the natural distribution characteristics of various record types in the training data are fully preserved, avoiding the model recognition bias problem caused by data imbalance.
[0043] (3) The system innovatively combines character-level sequence annotation with a label sequence optimization module. By capturing the contextual dependencies between characters and the label transfer constraints, it achieves precise positioning of entry boundaries, showing stronger robustness when dealing with nested, continuous, or incomplete entries.
[0044] (4) Post-processing rules based on a dual correction mechanism are introduced, which not only use citation context features for type verification, but also correct special character strings through label mapping rules, effectively solving the label conflicts and logical contradictions that may arise from pure end-to-end models.
[0045] (5) The single-character serialization processing method breaks through the limitations of traditional word segmentation technology on citation annotation tasks and can completely retain the fine-grained features of the original text. It is particularly suitable for processing citation entries containing unconventional separators or special punctuation.
[0046] (6) Through end-to-end automated processing, seamless conversion from original citations to structured entries is achieved, which greatly reduces the dependence on manual rule base maintenance and significantly improves the system's migration and application capabilities in cross-domain scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0048] Figure 1 This is a flowchart of a method for automatic indexing of Chinese citation entries based on deep learning according to the present invention;
[0049] Figure 2 1. It is a schematic diagram of the training process of a deep learning model in a method for automatic indexing of Chinese citation entries based on deep learning of the present invention;
[0050] Figure 3 This is a schematic diagram of the prediction process of the deep learning model in the automatic indexing method of Chinese citation entries based on deep learning of the present invention. DETAILED DESCRIPTION
[0051] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.
[0052] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0053] A first embodiment of the present invention provides a method for automatically indexing Chinese citation entries based on deep learning, the method comprising:
[0054] Step S10: According to the Chinese citation item indexing standard, the proportion of each item type is counted, and training corpus is extracted from the manually indexed data according to the proportion;
[0055] Step S20, converting the training corpus into a single-character sequence format and adding a label to each character as input data; the label includes a combination of a start and end mark and a record type;
[0056] Step S30: inputting the input data into a pre-built and trained deep learning model for training, wherein the deep learning model is constructed based on a combination of a contextual semantic encoding module and a label sequence optimization module;
[0057] Step S40: After the format of the citation data to be indexed is converted, it is input into the trained deep learning model for prediction to obtain a sequence of entries, and then merged into a complete entry according to the label;
[0058] Step S50 , amending the merged complete entry according to preset rules, wherein the preset rules include: determining the entry type according to the citation context features, and amending the label of the specific character string based on the label mapping rule table.
[0059] In order to more clearly explain the automatic indexing method of Chinese citation entries based on deep learning of the present invention, the following is combined with Figure 1-Figure 3 Each step in the embodiment of the present invention is described in detail, including step S10 to step S50, and each step is described in detail as follows:
[0060] Step S10: According to the Chinese citation item indexing standard, the proportion of each item type is counted, and training corpus is extracted from the manually indexed data according to the proportion;
[0061] After extracting the training corpus according to the above proportions, it also includes:
[0062] The training dataset is balanced by splitting the majority class corpus and merging it with the minority class corpus. Specifically:
[0063] Calculate the ratio of the sample size of the minority class to the sample size of the majority class in all categories of bibliographic items. When the ratio is less than a preset percentage, split the training corpus based on preset rules.
[0064] The preset rule is: set the number of splits k according to the business volume, and split the majority class samples into k subsets according to the set sampling method;
[0065] Each majority class subset is merged with all minority class samples to generate k balanced sub-datasets, and the deep learning model is trained based on the k balanced sub-datasets.
[0066] In this embodiment, the preset percentage is set to 50%. The sample sizes of all record categories in the training corpus are counted. If the ratio of the sample size of the minority class (the smallest sample class) to the sample size of the majority class (the largest sample class) in a certain category is less than 50% (i.e., the minority class is less than half of the majority class), the splitting process is triggered.
[0067] Based on business needs, in this embodiment, the maximum number of splits is set to 5. Calculate the number of samples of the majority class as a multiple of the minority class (e.g., 3.2 times), and take an integer multiple (e.g., 3 times) as the splitting benchmark. In this case, k = 3. If the multiple exceeds 5 (e.g., 6 times), then k = 5.
[0068] If the majority class sample size is not enough to split into k parts (for example, the sample size is too small), the actual number of splits is used (for example, when the sample size is 4 times that of the minority class, k = 4).
[0069] In this embodiment, the majority class samples are split into k subsets according to a set sampling method, specifically:
[0070] The majority class samples are numbered sequentially, and samples are extracted at fixed intervals to generate k subsets. For example:
[0071] If the sample size of the majority class is 320 (the minority class is 100, with a ratio of 32%), k=3, then the sampling interval is 3, and the way to generate subsets is: the first subset takes samples No. 1, 4, 7..., the second subset takes samples No. 2, 5, 8..., and the third subset takes samples No. 3, 6, 9...
[0072] Taking Chinese journal citation indexing as an example, assuming that the "author" category is the majority category (1,200 entries) and the "place of publication" category is the minority category (300 entries), the ratio is 25% (<50%), triggering a split.
[0073] Calculate k=4 (1200 / 300=4 times), and generate 4 majority class subsets (300 items in each subset, extracted by interval sampling).
[0074] Each subset was merged with 300 entries of the “place of publication” category to form four balanced data sets (600 entries each).
[0075] like Figure 2 As shown, step S20, converting the training corpus into a single-character sequence format and adding a label to each character as input data; the label includes a combination of a start and end mark and a record type;
[0076] Among them, each character is labeled and converted into a vector as input data.
[0077] Specifically, the acquired training corpus is processed and the data format is converted into the input format required by the model; the data is segmented by characters, and consecutive letters or numbers are regarded as one character. A tag is added to each character in the entry, taking a single entry as a unit;
[0078] The types of entries shall at least include author, title, source, edition description, year, volume number, issue number, page number, place of publication, publisher, source of parent document, institution awarding the thesis, editor-in-chief and to be published.
[0079] In this embodiment, the manually indexed original citation data is cleaned to remove redundant spaces, garbled characters and non-standard symbols (such as "@" and "#", etc.), while retaining Chinese characters, letters, numbers and standard symbols.
[0080] Treat consecutive letters / numbers as a whole (e.g., "ISBN" and "978704052123" in "ISBN978-7-04-052123" are each treated as a single character);
[0081] Split text by characters (Chinese characters and combined letters / numbers are all considered independent characters);
[0082] In this embodiment, a mark is added to each character in the entry based on preset symbols. The preset symbols are: @@ represents the author, ## represents the title, $$ represents the source, 『『 represents the year, 』』 represents the volume number, 〖〖 represents the issue number, 〖〖 represents the page number, ! ! represents the place of publication, :: represents the publisher,
【
represents the edition description,
[0083] Take the citation "Zhang San. Deep Learning [M]. Beijing: Science and Technology Press, 2021." as an example:
[0084] The character sequence after segmentation is: ["Zhang","three","."","deep","degree","learn","learn","[","M","]","."","Beijing","Beijing",":","science","technology","publishing","press",",","2021","."];
[0085] Corresponding label sequence:
[0086] B-@@,E-@@,SO,B-$$,I-$$,I-$$,I-$$,I-$$,I-$$,E-$$,SO,B-! ! ,E-! ! ,SO,B-::,I-::,I-::,I-::,E-::,SO,S-『『,SO.
[0087] In this embodiment, the cited article is a book, the title of the book is the document source, and the document carrier type "[M]" is marked in the document source together with the title of the book.
[0088] like Figure 3As shown, in step S30, the input data is input into a pre-built and trained deep learning model for training, and the deep learning model is constructed based on a combination of a context semantic encoding module and a label sequence optimization module; after training, it is optimized through a test set, which is part of the training corpus.
[0089] The context semantic encoding module is used to extract bidirectional context features from the character sequence;
[0090] The label sequence optimization module is used to learn the transition probability between labels based on context features and generate a globally optimal label sequence.
[0091] In this embodiment, the contextual semantic encoding module uses the BiLSTM model, and the label sequence optimization module uses the CRF model. The training corpus is vectorized and BiLSTM+CRF is used for automatic indexing modeling. The model parameters are:
[0092] optimizer = adadelta, dropout = 0.5, hidden_dim = 200, batch_size = 20, learningrate = 0.1. Iterate training and save model parameters.
[0093] This paper uses a BiLSTM module to capture the dependencies between characters in a quote, processing sequences both forward and backward (in forward order) and backward and forward (in reverse order), enhancing understanding of contextual semantics (for example, in "Zhang San. Journal of Computer Science," the contextual dependencies of "Zhang" include the "three" in the following text and the "." in the preceding text). This solves the problem of long-range dependencies and avoids the vanishing / exploding gradients of traditional RNNs.
[0094] As a label sequence optimization module, the CRF model globally constrains the label probability output by BiLSTM to ensure the legitimacy of the label sequence.
[0095] optimizer=adadelta indicates that the optimizer selects the Adadelta algorithm, which can adaptively adjust the learning rate without manually setting the initial learning rate. It is suitable for processing sparse data (such as uneven character distribution in text).
[0096] Dropout = 0.5 means that the random deactivation rate is 0.5, that is, during the training process, neurons are randomly "turned off" with a probability of 50% to prevent the model from overfitting the training data.
[0097] hidden_dim=200, indicating that the hidden layer dimension is 200.
[0098] batch_size=20, indicating that the batch size is 20, and 20 samples are input for parameter update each time training.
[0099] learning_rate=0.1 means the initial learning rate is 0.1.
[0100] Step S40: After the format of the citation data to be indexed is converted, it is input into the trained deep learning model for prediction to obtain a sequence of entries, and then merged into a complete entry according to the label;
[0101] In this embodiment, among the entries with the same characters in the label, those starting with B and ending with E are merged, and the entry label is saved.
[0102] Step S50 , amending the merged complete entry according to preset rules, wherein the preset rules include: determining the entry type according to the citation context features, and amending the label of the specific character string based on the label mapping rule table.
[0103] Modify the label of a specific string based on the entry type and the preset label mapping rules, specifically:
[0104] Modify tags based on preset string features, including keyword recognition, legal document determination, and adjacent tag association rules;
[0105] Based on the document type identification results, labels are modified according to the specific format features of conference papers, patents, standards, and journals;
[0106] For strings containing year-number structures, differentiated corrections are implemented based on the number arrangement format and value range;
[0107] Perform tag conflict detection across document types and perform tag correction based on institution suffixes, document identifiers, and special symbols;
[0108] Among them, the document types include at least conference papers, patents, standards, journals, books, dissertations, reports, archives, newspapers, case judgments, laws and regulations, and yearbooks.
[0109] For specific label mapping rules, see Table 1 and Table 2.
[0110] Table 1:
[0111] Table 2:
[0112] Assume that the citation information is: Wang Fang, Chen Li. Research on indexing model based on BiLSTM[J]. Journal of the China Society for Scientific and Technical Information, 2024, 61(4): 400-408.
[0113] The model's initial prediction result is the label "61(4)": "61" is mistakenly labeled as "issue number" and "4" is mistakenly labeled as "volume number".
[0114] “400-408” tag: not recognized as “page number”, mistakenly marked as “source”.
[0115] Triggering journal-specific correction rules:
[0116] In the format of "year, number 1 (number 2)", "number 1" is the volume number (61), and the label should be 』』; "number 2" is the period number (4), and the label should be 〖〖.
[0117] Correct "61"→ (volume number), "4"→〖〖(issue number).
[0118] Triggering universal format matching: "400-408" contains "-" and is a consecutive number, which conforms to the "page number" format (rule: the number range containing "-" is marked as page number), and the correction label is "page number".
[0119] Although the various steps in the above embodiment are described in the above-mentioned order, those skilled in the art will understand that in order to achieve the effect of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple changes are within the scope of protection of the present invention.
[0120] A second embodiment of the present invention provides a deep learning-based automatic indexing system for Chinese citation entries, based on a deep learning-based automatic indexing method for Chinese citation entries. The system includes:
[0121] A corpus generation module is configured to calculate the proportion of each type of item according to the Chinese citation item indexing standard, and extract training corpus according to the proportion from the manually indexed data;
[0122] a format conversion module configured to convert the training corpus into a single-character sequence format and add a label to each character as input data; the label includes a combination of a start and end mark and a record type;
[0123] A training module configured to input the input data into a pre-built and trained deep learning model for training, wherein the deep learning model is constructed based on a contextual semantic encoding module and a label sequence optimization module;
[0124] A prediction module is configured to convert the format of the citation data to be indexed and then input it into a trained deep learning model for prediction, thereby obtaining a sequence of entries and merging them into complete entries according to the labels;
[0125] A correction module is configured to correct the merged complete entry according to preset rules, wherein the preset rules include: determining the entry type based on the citation context characteristics, and correcting the label of the specific string based on the entry type and preset label mapping rules.
[0126] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process and related instructions of the system described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0127] It should be noted that the above embodiment provides a deep learning-based automatic indexing system for Chinese citation entries, which is illustrated only by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are merely for distinguishing the modules or steps and are not to be regarded as improper limitations on the present invention.
[0128] An electronic device according to a third embodiment of the present invention includes:
[0129] at least one processor; and
[0130] a memory communicatively connected to at least one of the processors; wherein,
[0131] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned method for automatic indexing of Chinese citation entries based on deep learning.
[0132] A fourth embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned method for automatic indexing of Chinese citation entries based on deep learning.
[0133] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes and related instructions of the storage device and processing device described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0134] Those skilled in the art should be able to appreciate that, in conjunction with the modules and method steps of each example described in the embodiments disclosed herein, it is possible to implement them with electronic hardware, computer software, or a combination of the two, and the programs corresponding to the software modules and method steps can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. In order to clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0135] The terms "first", "second", etc. are used to distinguish similar objects, rather than to describe or indicate a particular order or sequence.
[0136] The term "comprise" or any other similar term is intended to cover non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0137] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. A method for automatic indexing of Chinese citation entries based on deep learning, characterized in that: The method includes: According to the Chinese citation item indexing standards, the proportion of each item type is counted, and training corpus is extracted from the manually indexed data according to the said proportion; Converting the training corpus into a single-character sequence format and adding a label to each character as input data; the label includes a combination of a start and end mark and a record type; Input the input data into a pre-built and trained deep learning model for training, wherein the deep learning model is constructed based on a combination of a contextual semantic encoding module and a label sequence optimization module; The citation data to be indexed is converted into a new format and then fed into a trained deep learning model for prediction. This generates a sequence of entries and merges them into complete entries based on the labels. The merged complete entry is modified according to preset rules, wherein the preset rules include: determining the entry type according to the citation context characteristics, and modifying the label of the specific string based on the entry type and preset label mapping rules.
2. The method for automatic indexing of Chinese citation entries based on deep learning according to claim 1, characterized in that: After extracting the training corpus according to the above proportions, it also includes: The training dataset is balanced by splitting the majority class corpus and merging it with the minority class corpus.
3. The method for automatic indexing of Chinese citation entries based on deep learning according to claim 2, characterized in that: Split the majority class corpus and merge it with the minority class corpus, specifically including: Calculate the ratio of the sample size of the minority class to the sample size of the majority class in all categories of bibliographic items. When the ratio is less than a preset percentage, split the training corpus based on preset rules. The preset rule is: set the number of splits k according to the business volume, and split the majority class samples into k subsets according to the set sampling method; Each majority class subset is merged with all minority class samples to generate k balanced sub-datasets, and the deep learning model is trained based on the k balanced sub-datasets.
4. The method for automatic indexing of Chinese citation entries based on deep learning according to claim 1, characterized in that: The context semantic encoding module is used to extract bidirectional context features from the character sequence; The label sequence optimization module is used to learn the transition probability between labels based on context features and generate a globally optimal label sequence.
5. The method for automatic indexing of Chinese citation entries based on deep learning according to claim 1, characterized in that: The start and end marks include B to mark the start of the entry, I to mark the middle position of the entry, E to mark the end of the entry, and S to mark the entry as a single word.
6. The method for automatic indexing of Chinese citation entries based on deep learning according to claim 1, characterized in that: The types of entries shall at least include author, title, source, edition description, year, volume number, issue number, page number, place of publication, publisher, source of parent document, institution awarding the thesis, editor-in-chief and to be published.
7. The method for automatic indexing of Chinese citation entries based on deep learning according to claim 1, characterized in that: Modify the label of a specific string based on the entry type and the preset label mapping rules, specifically: Modify tags based on preset string features, including keyword recognition, legal document determination, and adjacent tag association rules; Based on the document type identification results, labels are modified according to the specific format features of conference papers, patents, standards, and journals; For strings containing year-number structures, differentiated corrections are implemented based on the number arrangement format and value range; Perform tag conflict detection across document types and perform tag correction based on institution suffixes, document identifiers, and special symbols; Among them, the document types include at least conference papers, patents, standards, journals, books, dissertations, reports, archives, newspapers, case judgments, laws and regulations, and yearbooks.
8. A deep learning-based automatic indexing system for Chinese citation entries, based on the deep learning-based automatic indexing method for Chinese citation entries according to any one of claims 1 to 7, characterized in that: The system includes: A corpus generation module is configured to calculate the proportion of each type of item according to the Chinese citation item indexing standard, and extract training corpus according to the proportion from the manually indexed data; a format conversion module configured to convert the training corpus into a single-character sequence format and add a label to each character as input data; the label includes a combination of a start and end mark and a record type; A training module configured to input the input data into a pre-built and trained deep learning model for training, wherein the deep learning model is constructed based on a contextual semantic encoding module and a label sequence optimization module; A prediction module is configured to convert the format of the citation data to be indexed and then input it into a trained deep learning model for prediction, thereby obtaining a sequence of entries and merging them into complete entries according to the labels; A correction module is configured to correct the merged complete entry according to preset rules, wherein the preset rules include: determining the entry type based on the citation context characteristics, and correcting the label of the specific string based on the entry type and preset label mapping rules.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the automatic indexing method of Chinese citation entries based on deep learning as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to be executed by the computer to implement the automatic indexing method for Chinese citation entries based on deep learning as described in any one of claims 1 to 7.