Reference review method and system, electronic equipment and storage medium
By building a rule feature library and a regular expression rule set, combined with citation feature data from academic databases, automated reference review is achieved, solving the problem of low efficiency of manual review and improving the accuracy and reliability of review.
Patent Information
- Application Number
- CN202511156598.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-08-19
AI Technical Summary
In the existing technology, the review of references relies on manual work, which is inefficient and easily affected by human factors. It is difficult to adapt to various changing factors, resulting in inaccurate review results.
By building a rule feature library, using matching tests of regular expression rule sets and sample sets, dynamically adjusting rules to adapt to format specification changes of publishing institutions, and combining citation feature data from academic databases, automated reference review is achieved, including multi-level verification and rule update mechanisms.
It improves the accuracy and reliability of reference review, reduces the uncertainty of manual review, and improves review efficiency.
Smart Images

Figure CN120780833A_ABST
Abstract
Description
TECHNICAL FIELD
[0002] The present application belongs to the field of document proofreading, and particularly relates to a reference document proofreading method and system, an electronic device and a storage medium. BACKGROUND
[0003] In today's academic research field, the speed of knowledge innovation and dissemination is increasing, and academic achievements in the form of papers, monographs and other forms are emerging. As an important part of academic research, reference documents not only respect and cite the research achievements of predecessors, but also are an important link for academic exchange and knowledge inheritance. It provides readers with clues for further research, and helps to promote the in-depth development of academic research.
[0004] With the rapid development of Internet technology, academic information is more convenient to obtain, and academic databases have included a large amount of academic documents to provide rich resources for researchers. However, this has also brought about the problem of a sharp increase in the number of reference documents, which has put forward higher requirements for the proofreading of reference documents.
[0005] In the process of academic publishing, different periodical societies and publishing houses have different requirements for the format and recording of reference documents. These requirements usually include the arrangement order of reference documents, the format of recording items (such as the writing method of author's name, the abbreviation form of periodical name, etc.), the use of punctuation marks, etc. Accurate recording of reference documents not only reflects the professionalism of the author, but also relates to the standardization and credibility of academic achievements.
[0006] In the prior art, the proofreading of reference documents is mainly completed manually. Editors or scholars need to spend a lot of time and effort to check the format and content of reference documents according to the requirements of periodical societies. This manual proofreading method is low in efficiency and is easily affected by human factors such as fatigue and negligence, resulting in inaccurate proofreading results. In addition, with the updating of academic standards and the changes in the requirements of periodical societies, manual proofreading is difficult to adapt to these changes in real time, further increasing the difficulty of proofreading.
[0007] Therefore, the present application provides a reference document proofreading method to solve the above technical problems. SUMMARY
[0008] The purpose of the present application is to provide a reference document proofreading method, system, electronic device and storage medium to solve the technical problem that manual review in the prior art is difficult to adapt to the influence caused by various changing factors, resulting in low review efficiency.
[0009] In order to solve the above technical problems, the present application provides a reference document proofreading method, comprising: In response to the captured publishing agency format specification change event and the extracted citation feature dataset from the academic database, the format specification change event and the citation feature dataset are fused and analyzed to dynamically build a rule feature library, wherein the rule feature library includes the distribution rule of the entry format; In response to an operation instruction of a user interface, the operation instruction is converted into a regular expression rule set based on the distribution rule of the entry format in the rule feature library, and a test sample set is generated according to a sample feature in the rule feature library, a matching test of the rule set and the sample set is performed, error type distribution of unmatching samples is analyzed, and parameters of the regular expression are optimized; Based on the optimized regular expression rule set and the sample set, logical contradictions of rules in the rule set and coverage defects of samples in the sample set are verified and calculated, a comprehensive evaluation value is generated, and the rule is activated when the comprehensive evaluation value is higher than a preset quality threshold; The activated rule is converted into a structured storage format according to a feature dimension, a version control index is recorded to record a rule generation time and a change history, and a change tracking mechanism is synchronously set to respond to and update the rule feature library; In response to a target reference, the rule in the rule feature library is called to perform regular matching, and authenticity multi-level verification is performed on the regular matching result to generate a proofreading report including the verification result.
[0010] In some embodiments, in response to the captured publishing agency format specification change event and the extracted citation feature dataset from the academic database, the format specification change event and the citation feature dataset are fused and analyzed to dynamically build a rule feature library, wherein the rule feature library includes the distribution rule of the entry format, and further includes: Periodically scanning a reference format page of a publishing agency official website to identify the format specification change event through content hash comparison; Analyzing the appearance frequency and distribution rule of the author name and title structure in the citation feature dataset in the academic database; Fusing the captured format specification change event and the analyzed appearance frequency and distribution rule of the main fields to generate a rule feature library with a time-stamped version identifier; When the format specification change event is identified to be changed, a new and old version difference comparison is performed to generate a rule feature library update instruction.
[0011] In some embodiments, in response to operation instructions of the user interface, based on the distribution rule of the catalog item format in the rule feature library, the operation instructions are converted into a regular expression rule set, a test sample set is generated according to the sample features in the rule feature library, a matching test of the rule set and the sample set is performed, the error type distribution of the unmatched samples is analyzed, and the parameters of the regular expression are optimized, further comprising: parsing the distribution rule of the catalog item format in the rule feature library to generate a regular expression template including a named capture group; based on the regular expression template, the operation instructions of the user in the interface are mapped into regular expression parameter configuration in real time; based on the sample features recorded in the rule feature library, a test sample set covering multiple scenarios is constructed; performing batch matching test of the rule set and the sample set, counting error type distribution data of the unmatched samples, and adjusting character range definition and quantifier parameters of the named capture group according to the error type distribution data of the unmatched samples.
[0012] In some embodiments, based on the optimized regular expression rule set and the sample set, the logical contradiction of the rules in the rule set and the coverage defects of the samples in the sample set are verified and calculated, a comprehensive evaluation value is generated, and when the comprehensive evaluation value is higher than a preset quality threshold, the rule is activated, further comprising: inputting the regular expression rule set and the associated sample set into a pre-trained verification model; verifying the logical contradiction of the rules in the rule set, identifying sample coverage blind area and boundary condition missing; calculating the comprehensive evaluation value of the logical defect rate of the rules in the rule set and the sample coverage rate, and when the comprehensive evaluation value is higher than a preset quality threshold, activating the rule and marking; when the comprehensive evaluation value is lower than a preset quality threshold, extracting error type distribution data of the corresponding rule, reconstructing the process and optimizing the parameters of the regular expression accordingly.
[0013] In some embodiments, in response to the target reference, the corresponding structured rule in the rule feature library is called to perform regular matching, multi-level verification of the authenticity of the regular matching result is performed, an editing report including the verification result is generated, and further comprising: according to the journal type of the target reference selected by the user, the corresponding active rule in the rule feature library is called to perform regular expression matching; extracting the catalog item metadata of the matching success, and initiating the authenticity verification of the title, author and journal name fields to the academic database; performing semantic completion on the failed matching field to generate candidate metadata and reinitiating the authenticity verification; generating a proofreading report based on the authenticity verification result, including integrating format errors and authenticity deviation data, and generating a structured proofreading report containing error location, error type and correction scheme.
[0014] In some embodiments, when the format specification change event change is identified, a new and old version difference comparison is performed and a rule feature library update instruction is generated, further including: locating the specific change point of the new and old format specification change event through a text difference algorithm; analyzing the influence range of the change point on the format distribution rule of the bibliographic item in the rule feature library; reconstructing the format distribution model of the rule feature library according to the influence range and generating a new version identifier; writing the version change record into the history track of the rule feature library and triggering an associated rule update notification.
[0015] In some embodiments, the authenticity multi-level verification further includes: basic field verification, including accurately comparing bibliographic item metadata that successfully matches regular expressions with corresponding data in an academic database to detect title spelling and author name order differences; semantic completion verification, including completing the bibliographic item metadata through semantic analysis for fields that fail to match regular expressions, and performing cross-database joint query verification; deformation compatibility verification, including performing multi-language spelling deformation compatibility verification on author name fields to handle abbreviations, aliases and cultural difference variants.
[0016] Based on the same concept, the present application also provides a reference literature proofreading system, comprising: a rule feature library construction module configured to respond to captured publishing agency format specification change events and extracted citation feature data sets from an academic database, to fuse and analyze the format specification change events and the citation feature data sets, and to dynamically construct a rule feature library, wherein the rule feature library includes bibliographic item format distribution rules; a matching test module configured to respond to user interface operation instructions, to convert the operation instructions into a regular expression rule set based on the bibliographic item format distribution rules in the rule feature library, to generate a test sample set according to sample features in the rule feature library, to perform matching test of the rule set and the sample set, to analyze error type distribution of unmatched samples and to optimize parameters of the regular expression; The comprehensive evaluation module is configured to verify and calculate logical contradictions of rules in the rule set and coverage defects of samples in the sample set based on the optimized regular expression rule set and the sample set, and generate a comprehensive evaluation value, and activate the rules when the comprehensive evaluation value is higher than a preset quality threshold; The rule updating module is configured to convert the activated rules into a structured storage format according to feature dimensions, record the rule generation time and change history by establishing a version control index, and set a change tracking mechanism to respond to and update the rule feature library; The document proofreading module is configured to execute regular matching by calling corresponding structured rules in the rule feature library in response to target reference literature, perform multi-level verification on the regular matching results, and generate a proofreading report including the verification results.
[0017] Based on the same concept, the present application further provides an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus; the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the reference literature proofreading method.
[0018] Based on the same concept, the present application further provides a computer readable storage medium, which stores a computer program executable by an electronic device, and when the computer program runs on the electronic device, the electronic device executes the steps of the reference literature proofreading method.
[0019] Compared with the prior art, the present application has the beneficial effects that: The present application discloses a reference literature proofreading method and system, an electronic device and a storage medium, which can improve the accuracy and reliability of literature proofreading, avoid the uncertainty of manual proofreading, and improve the proofreading efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0020] Other features, objects and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings: Figure 1 is a flowchart of a reference literature proofreading method in some specific embodiments of the present application; Figure 2 is one of the flowcharts of a reference literature proofreading method in another specific embodiment of the present application; Figure 3 is the second flowchart of a reference literature proofreading method in another specific embodiment of the present application; Figure 4is a flowchart of a reference proofreading method in another embodiment of the present application; Figure 5 is a flowchart of a reference proofreading method in another embodiment of the present application; Figure 6 is a structural diagram of a reference proofreading system in some embodiments of the present application; Figure 7 is a structural diagram of an electronic device in some embodiments of the present application; In the figure, 710 processor; 720 memory; 730 input device; 740 output device. DETAILED DESCRIPTION
[0021] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0022] The terms used in the embodiments of the present application are only for the purpose of describing the specific embodiments, and are not intended to limit the present application. The singular forms "a", "an" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Multiple" generally includes at least two.
[0023] It should be understood that the term "and / or" used herein only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper generally represents that the front and rear associated objects are a "or" relationship.
[0024] It should be understood that although the terms first, second, third, etc. can be used in the embodiments of the present application to describe, these descriptions should not be limited to these terms. These terms are only used to distinguish the description. For example, without departing from the scope of the embodiments of the present application, the first can also be called the second, and similarly, the second can also be called the first.
[0025] Depending on the context, the word "if' can be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," generally or as used in this document. Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]," generally or as used in this document.
[0026] It is also important to note that the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0027] It is particularly important to note that the symbols and / or numbers present in the specification, if not marked in the description of the figures, are not figure references.
[0028] With reference to Figure 1 A reference proofreading method comprises: S101, in response to a captured publishing agency format specification change event and a citation feature data set extracted from an academic database, performing fusion analysis on the format specification change event and the citation feature data set to dynamically construct a rule feature library, wherein the rule feature library comprises a catalog item format distribution rule; S102, in response to an operation instruction of a user interface, converting the operation instruction into a regular expression rule set based on the catalog item format distribution rule in the rule feature library, simultaneously generating a test sample set according to a sample feature in the rule feature library, performing a matching test of the rule set and the sample set, analyzing an error type distribution of unmatched samples and optimizing parameters of the regular expression; S103, based on the optimized regular expression rule set and the sample set, verifying and calculating a logical contradiction of a rule in the rule set and a coverage defect of a sample in the sample set, generating a comprehensive evaluation value, and activating the rule when the comprehensive evaluation value is higher than a preset quality threshold; S104, converting the activated rule into a structured storage format according to a feature dimension, recording a rule generation time and a change history by establishing a version control index, and synchronously setting a change tracking mechanism to respond to and update the rule feature library; S105, in response to the target reference, calling the corresponding structured stored rules in the rule feature library to perform regular matching, performing authenticity multi-level verification on the regular matching result, and generating a proofreading report including the verification result.
[0029] Specifically, in the embodiments of the present application, in response to a reference format specification change event published by a publishing agency website (for example, the journal "Medical Research" changes the author name format from "surname, comma, space, name abbreviation" to "surname, space, name abbreviation"), the change event is captured through an automated program; a citation feature data set is extracted from an academic database (for example, 100,000 CNKI papers are analyzed to find that the proportion of "Li, Y." in the author field is 68%, the proportion of "Li Y" is 30%, and the proportion of "Y. Li" is 2%), and the change event and the feature data set are fused and analyzed: first, the key difference points in the change event (such as the deletion of the comma) are extracted, and the format distribution rules in the citation feature data set (such as the frequency of use of the comma in the author field) are combined to dynamically build a rule feature library with version identifier V2.1, which records the cataloging item format distribution rules, including the new distribution rule for the author field: {surname} {space} {name abbreviation} accounts for 100%, and the title field capitalization accounts for 92% (for example, the title "Deep learning" matches the capitalization rule, while "deep learning" is considered abnormal).
[0030] The user checks the "no comma in author name" option through the visual interface and submits the operation instruction, which is converted into a regular expression rule set based on the cataloging item format distribution rules in the rule feature library: Original rule (?<last_name>[^,]+),\s(?<first_initial>[A-Z]\.); Optimized to (?<last_name>[^\s]+)\s(?<first_initial>[A-Z]\.); At the same time, a test sample set containing 500 samples is generated according to the sample features in the rule feature library (including compliant samples such as "Zhang S." and "Wang H." and error samples such as "Zhang, S." and "Wang H"); after performing matching test of the rule set and the sample set, the error type distribution of the unmatched samples is analyzed: the proportion of samples with commas that are not matched is 30%, and the proportion of name abbreviation samples without punctuation that are not matched is 25%; accordingly, the regular expression parameters are optimized: the name field capture group is expanded to (?<first_initial>[A-Z]\.?) to be compatible with the no-punctuation case.
[0031] Based on the optimized regular expression rule set and sample set, verify and calculate the logical contradiction of the rules in the rule set and the coverage defects of the samples in the sample set: 2 places of logical contradiction of the rules are detected (such as the rule requires "space between surname and name" but the sample contains "ZhangS." without space), 1 type of sample coverage defects is identified (the test sample set does not contain "full name" variants such as "Zhang San"); generate a comprehensive evaluation value: 2 places of logical contradiction, sample coverage rate 87%, comprehensive score 89 points; activate the rule when the comprehensive score is higher than the preset quality threshold 85 points. Convert the activated rule into a structured storage format (for example, stored as a JSON structure {"rule_id":"AUTH_V2.1","regex":"..."} ) according to the feature dimension coding, establish version control index (record version number: V2.1, parent version: V2.0, change summary: delete comma), and synchronize the change tracking mechanism: when the new specification requires "author's surname in capital letters", automatically trigger the rule feature library update to V2.2.
[0032] In response to the target reference "Zhang S. et al. Medical AI. 2023", the AUTH_V2.1 rule stored in the rule feature library is called to perform regular matching: capture the author field "Zhang S."; Perform multi-level verification on the matching result: first, perform basic field verification, compare with academic database metadata (CDB author is "Zhang, S."), and detect "comma missing" deviation; Perform semantic completion verification on the uncommon abbreviation "et al.", and verify after completing "et alii"; Finally, perform deformation compatibility verification, compatible with the edit distance difference between "ZhangS" and "Zhang, S." (the calculated character difference value is 2, which is lower than the threshold 3); generate a proofreading report containing the verification result: {error location: author field, error type: format deviation, suggested correction: "Zhang, S."}.
[0033] In some applications, in response to the captured publishing agency format specification change event and the extracted citation feature data set from the academic database, the format specification change event and the citation feature data set are fused and analyzed to dynamically build a rule feature library, wherein the rule feature library includes bibliographic item format distribution rules, including periodically scanning the reference format page of the publishing agency website, and identifying the format specification change event through content hash comparison; analyzing the author name, title structure, and main field appearance frequency and distribution rules in the citation feature data set in the academic database; fusing the captured format specification change event with the analyzed main field appearance frequency and distribution rules to generate a rule feature library with a timestamp version identifier; when the format specification change event is identified to be changed, performing new and old version difference comparison and generating a rule feature library update instruction.
[0034] It is understandable that in this application, in response to captured publishing institution format specification change events and academic database citation feature datasets, change detection is achieved by periodically scanning the reference format pages on the publishing institution's official website: an automated program is configured to scan the target journal's official website (such as "Natural Science Reviews") daily, calculate the hash value of the page text for comparison, and trigger a change event when the hash value does not match. For example, the scan discovered that the journal "Cell Research" updated its reference format. The old specification was "Author.Title[J].Journal Name,Year,Volume(Issue):Page Number", and the new specification was changed to "Author.Title[J].Journal Name,Year of Publication,Volume(Issue):Start Page-End Page". The key change points identified are "Year→Year of Publication", "Volume(Issue)→Volume(Issue)", and "Page Number→Start Page-End Page". At the same time, we extracted citation feature datasets from academic databases for analysis: A statistical analysis of the references of 500,000 documents revealed that the "surname + comma + first name abbreviation" format (e.g., "Zhang, S.") accounted for 58% of the author field, and the "surname + space + first name abbreviation" format (e.g., "Zhang S.") accounted for 37%; the "number bracket" format (e.g., "15(2)") accounted for 62% of the volume field, and the "text prefix" format (e.g., "Vol.15, No.2") accounted for 35%. We then integrated the captured change events with the analyzed distribution patterns: for changes in the author field, the "surname + comma + first name abbreviation" format accounted for 58% of the data in the distribution pattern; for changes in the volume field, the "text prefix" format accounted for only 35%, but this is a mandatory requirement of the new standard. Generate a rule feature library with a timestamp and version identifier (e.g., version number ref_rules_v20250801). The author field is compatible with two formats (comma format 58% of the time, space format 37%), and the volume / issue field is required to use the "volume (issue)" format (covering 35% of the data). When a change event is identified, a difference comparison is performed. For example, the journal "Medical Frontiers" has added a new requirement for "author last names to be capitalized." This is compared with the feature library's case-insensitive author last names. An update instruction is generated, including the field change type, the new rule description, and the compatibility coefficient with existing data (in this case, only 8% of the citation data conforms to the uppercase format). This ultimately triggers an update to the rule feature library to the new version ref_rules_v20250810, marking the new constraint on the author field.
[0035] Furthermore, in this application, when a change in the format specification change event is identified, a comparison of the difference between the new and old versions is performed and a rule feature library update instruction is generated, including locating the specific change point of the new and old format specification change events through a text difference algorithm; analyzing the impact range of the change point on the format distribution pattern of the entry in the rule feature library; reconstructing the format distribution model of the rule feature library according to the impact range, and generating a new version identifier; writing the version change record into the historical track of the rule feature library and triggering an associated rule update notification.
[0036] It can be understood that in this application, when a publishing agency format specification change event is identified, a new-old version difference comparison is performed and a rule feature library update instruction is generated, first, the new-old format specification text is parsed line by line through a text difference algorithm, the specific change point is located and the change type and location are marked; for example, the Environmental Science journal changes the reference title format from “《Title》” to “Title” (delete the book name), the difference algorithm identifies that the deletion operation occurs at the 5th-15th character of the 3rd line, and the change type is “format simplification”. Then analyze the influence range of this change point on the format distribution law of the rule feature library: based on the citation feature data set statistics, 82% of the literature uses the original book name format, after the new specification forces to delete the book name, the title field distribution law needs to be recalculated, the influence range covers all the rule items related to the title format in the rule feature library (such as 12 rule items such as title verification rule, title capture rule, etc.). According to the influence range, the format distribution model of the rule feature library is reconstructed: delete the original “with book name” distribution item (original proportion 82%), add a new “without book name” forced distribution item (initial coverage rate 0%, need to be dynamically adapted), generate a new version identifier “ref_rules_v20250901_EM1” (where “EM” represents emergency update, “1” represents the first iteration). Finally, write the version change record into the rule feature library history track: record the change time, change journal, number of affected rules, data compatibility coefficient (in this case, the compatibility of new and old data is only 18%), and trigger the associated rule update notification - send a warning email to the subscription user (content includes change summary and affected rule list), and at the same time, push the rule update instruction (instruction code “RU_UPDATE_TITLE_FORMAT”) to the proofreading system through API.
[0037] In some of the applications, in response to the operation instruction of the user interface, based on the format distribution law of the entry in the rule feature library, the operation instruction is converted into a regular expression rule set, at the same time, a test sample set is generated according to the sample features in the rule feature library, the matching test of the rule set and the sample set is performed, the error type distribution of the unmatched samples is analyzed and the parameters of the regular expression are optimized, including parsing the format distribution law of the entry in the rule feature library, generating a regular expression template including a named capture group; based on the regular expression template, the operation instruction of the user in the interface is mapped into a regular expression parameter configuration in real time; based on the sample features recorded in the rule feature library, a test sample set covering multiple scenarios is constructed; the batch matching test of the rule set and the sample set is performed, the error type distribution data of the unmatched samples is counted, and the character range definition and quantifier parameters of the named capture group are adjusted according to the error type distribution data of the unmatched samples.
[0038] It can be understood that in this application, in response to the operation instruction of the user interface, the rule set construction and optimization are performed based on the distribution rule of the catalog item format in the rule feature library. First, the distribution rule of the catalog item format recorded in the rule feature library is parsed to generate a regular expression template containing a naming capture group. For example, the author field distribution rule in the rule feature library is "full name in uppercase + space + name abbreviation" (such as "ZHANG S." accounts for 80%). The basic template is generated as (?<last_name>[A-Z]+)\s(?<first_initial>[A-Z]\.). The last_name capture group corresponds to the name field, and the first_initial capture group corresponds to the name abbreviation field.
[0039] Based on this template, the operation instruction of the user in the interface is mapped into a regular expression parameter configuration in real time: when the user drags the "surname length limit" slider to "3-8 characters" and checks the "name abbreviation optional punctuation" option in the visual interface, the operation instruction is converted into a parameter constraint: modify the last_name capture group to [A-Z]{3,8} to limit the character length, and extend the first_initial capture group to [A-Z]\.? to be compatible with the no-punctuation format (such as "S" and "S." can be matched).
[0040] At the same time, test sample set is constructed according to the sample features recorded in the rule feature library: Positive sample (consistent with new rules): "ZHANG S. Medical AI, 2025" (5 characters of surname length + name abbreviation with punctuation); "WANG X Physics Review, 2024" (4 characters of surname length + name abbreviation without punctuation); Negative sample (covering common errors): "zhang S...." (surname not in uppercase); "LIU YJ..." (name abbreviation too long); Finally, a test set containing 200 samples is generated, covering 92% of the format variants in the feature library.
[0041] Perform batch matching test of rule set and sample set: Run regular rule set to match all samples; Statistical error type distribution of unmatched samples: Error type A: surname not in uppercase (accounting for 40%, such as sample "chen H."); Error type B: name abbreviation without punctuation but rule requires punctuation (accounting for 35%, such as "LEE T"); Error type C: Last name exceeds the limit (25% of the time, such as "CHRISTOPHER K.") Optimize parameters based on error distribution data: For error type A: Expand the character range of the last_name capture group from [AZ] to [A-Za-z] to be lowercase compatible, but add a case conversion marker (generate an intermediate rule (?<last_name> [A-Za-z]+)\s→automatically converted to uppercase when output); For error type B: Modify the quantifier parameters and optimize the first_initial group from [AZ]\. to [AZ]\.? (the question mark indicates matching 0 or 1 punctuation marks). For error type C: Add a length check sub-rule to trigger a warning when the number of characters in the captured name is greater than 8; Comparison of rules before and after optimization: Original rules: (?<last_name> [AZ]{3,8})\s(?<first_initial> [AZ]\.); Optimized rules: (?<last_name> [A-Za-z]{3,8})\s(?<first_initial> [AZ]\.?); Additional check: if len(last_name)>8 then flag("Last name is too long"); Test verification: Original unmatched sample "chen H." → successfully matched after optimization (last name automatically capitalized to "CHEN"); The original unmatched sample "LEE T" → successfully matched after optimization (compatible with "T" without punctuation); Sample "CHRISTOPHER K." → captures the last name "CHRISTOPHER" (length 11) and triggers a warning.
[0042] In some applications, based on the optimized regular expression rule set and the sample set, logical contradictions of rules in the rule set and coverage defects of samples in the sample set are verified and calculated, a comprehensive evaluation value is generated, and when the comprehensive evaluation value is higher than a preset quality threshold, the rule is activated, including inputting the regular expression rule set and the associated sample set into a pre-trained verification model; verifying the logical contradiction points of the rules in the rule set, identifying sample coverage blind areas and boundary condition missing; calculating the comprehensive evaluation value of the logical defect rate of the rules in the rule set and the sample coverage rate, and when the comprehensive evaluation value is higher than a preset quality threshold, the rule is activated and marked; when the comprehensive evaluation value is lower than a preset quality threshold, the error type distribution data of the corresponding rule is extracted, the process is reconstructed and the parameters of the regular expression are optimized accordingly.
[0043] It can be understood that in this application, based on the optimized regular expression rule set and the sample set, rule verification and activation are performed, the regular expression rule set and its associated test sample set are input into a pre-trained verification model (a natural language processing model based on a deep learning architecture), and the model verifies the self-consistency of the rules through syntax tree analysis and logical reasoning. For example, input the author field rule set (?<last_name>[A-Za-z]{3,8})\s(?<first_initial>[A-Z]\.?) and 200 test samples (including positive examples "ZHANGS." and negative examples "chen H."). The verification model performs the following operations: Logical contradiction point verification: Detect rule self-conflict: If there are two constraints in the rule set: constraint A: name abbreviation must contain punctuation (first_initial group needs to match [A-Z]\.) constraint B: punctuation is optional (first_initial group is [A-Z]\.?) The model marks the conflict type as "quantifier conflict" and outputs a conflict point report: conflict position: first_initial capture group; conflict type: punctuation necessity conflict; conflict rule ID: Rule_Author_2025-001; Sample coverage defect identification: Boundary condition missing detection: The test sample set does not contain the following boundary scenarios: double-surname authors (such as "O'Brien T.") and names containing hyphens (such as "Smith-John A."). The model identifies the coverage blind area and outputs a defect report: missing sample type: compound surname; missing proportion: 12% of feature library distribution; boundary condition: unhandled apostrophe / hyphen; Comprehensive evaluation value calculation: a weighted algorithm is used to generate a comprehensive score: Logical defect rate (weight 60%): defect rate = number of contradictory rules / total number of rules, for example: 2 contradictory rules are detected (total 50) → defect rate 4% → score 96 points (weight value 57.6); Sample coverage rate (weight 40%): coverage rate = number of matched samples / total number of samples, for example: 184 of 200 samples are matched → coverage rate 92% → score 92 points (weight value 36.8); Comprehensive score: 57.6 + 36.8 = 94.4 points (preset quality threshold 85 points); Activation decision: 94.4>85→activate the rule and mark the effective state; Substandard rule processing: when the comprehensive score is 72 points (lower than 85 points): Extract error type distribution: main errors: surname length exceeds limit (25%) → sample "CHRISTOPHER K."; name abbreviation format error (30%) → sample "LEE T" Reconstruction optimization process: for length exceeding: add dynamic length check sub-rule: (?<last_name>[A-Za-z]{3,8})(?(?<=.{8})[^.]*|)# more than 8 characters trigger exception; for name abbreviation error: expand character set compatible with number: -[A-Z]\.?;+[A-Z0-9]\.?# compatible with "ZHANG R1.".
[0044] In some applications, in response to the target reference, the corresponding structured stored rules in the rule feature library are called to perform regular matching, the authenticity of the regular matching result is verified, an editing report including the verification result is generated, including calling the corresponding active rules in the rule feature library according to the journal type of the target reference selected by the user to perform regular expression matching; extracting the metadata of the bibliographic item that matches successfully, initiating authenticity verification of fields such as title, author and journal name to an academic database; performing semantic completion on the fields that fail to match, generating candidate metadata and reinitiating authenticity verification; generating an editing report based on the authenticity verification result, including integrating format errors and authenticity deviation data, generating a structured editing report containing error position, error type and correction scheme.
[0045] It can be understood that in this application, in response to the target reference "Application of Multimodal Deep Learning in Medical Imaging" (Author: Zhang S. et al, Journal: Medical AI Review, 2025, Vol.12, No.3:45-58), the activation rules stored in the structured rule feature library are called to perform regular matching: according to the journal type "Medical AI Review" selected by the user, the rule library is indexed, and the volume field rule (? <volume>Vol \. \d+),\s(?-improve <issue>No\.\d+) performs a match, successfully capturing the volume number "Vol. 12" and the issue number "No. 3", but the author field rule (?<last_name>[A-Z][a-z]+)\s(?<first_initial>[A-Z]\. does not match the "etal" abbreviation part. The successfully matched bibliographic metadata (title, volume, issue, page) is extracted and a verification of authenticity is initiated towards the academic database: through a distributed query interface, the CNKI inclusion data is retrieved, and it is found that the title record is "Multimodal Deep Learning in Medical Imaging" (abbreviation not expanded), a title abbreviation deviation is detected; the volume and issue information is consistent with the database metadata. Semantic completion is performed on the failed author field "et al": through a pre-trained model, it is restored to the complete bibliographic entry "Zhang S., Li M., Wang H.", the author authenticity verification is re-initiated, and it is confirmed that the three authors exist in the database signature list. Based on the verification results, an editing report is generated: the volume and issue field format correctness flag, the title field semantic deviation (detected as "uncommon abbreviation", recommended to be corrected to the full name), and the author field completion verification results are integrated, and a structured report is outputted, including the error location (title field), error type (authenticity deviation), correction scheme (expand the title abbreviation), and author field completion details (original input "et al" is completed to a three-person signature list), while the volume and issue field verification is marked as passed.
[0046] Further, in this application, the authenticity multi-level verification includes basic field verification, including accurate comparison of bibliographic metadata matched by regular expression with corresponding data in the academic database, detecting title spelling, author name order differences; semantic completion verification, including completing the complete bibliographic metadata through semantic analysis for the fields that fail to match the regular expression, and performing cross-database joint query verification; deformation compatible verification, including performing multi-language spelling deformation compatibility verification on the author name field, processing abbreviations, aliases, and cultural difference variants.
[0047] It can be understood that in this application, the authenticity multi-level verification is implemented through a three-level verification mechanism: Basic field verification: accurate comparison of bibliographic metadata matched by regular expression with authoritative records in a distributed academic database cluster at the field level, such as title field verification detecting spelling deviation ("Learing" vs "Learning") between the target literature title "Deep Learing Applications" and the database record "Deep Learning Applications", and author field verification finding name order difference (input "Zhang, S." corresponds to database "San Zhang" with inverted order); Semantic completion verification: For fields that fail regular matching, complete the complete metadata through a pre-trained semantic model, for example, the original input "Wang et al" in the author field is completed to the complete bibliographic entry "Wang H., Zhang L., Chen T.", then perform cross-library joint query (search databases such as CNKI and Web of Science), verify the authenticity of the signature of the three authors in the target literature "Nanometer material synthesis"; Deformation compatibility verification: Perform multi-language spelling deformation compatibility verification on the author name field, for example, handle the deformation of the English author "Smith, J." as follows: Abbreviation compatibility: Verify that "J. Smith" and "Smith J." are the same author; Alias processing: Identify the alias association of "Robert Smith" and "Bob Smith"; Cultural differences: Chinese author "Zhang Wei" compatible spelling variants "Zhang Wei" and "Wei Zhang" By calculating the string edit distance (set threshold ≤2) and cultural naming convention analysis, confirm the equivalence of deformation.
[0048] The following will be combined Figures 2 to 5 Another specific embodiment of the reference proofreading method is described below: As Figure 2 And Figure 3 shown, Use the requests library and BeautifulSoup library of Python to write a web crawler program, regularly access the official website of each journal society, and find the page related to the reference proofreading requirements. Parse the contents of the crawled page, extract the key information of the proofreading requirements such as the reference format template, bibliographic entry requirements, etc., and store them in the database.
[0049] Collect a large amount of citation data included in CNKI, including title, author, journal name, year, volume, page number, etc.
[0050] Use data analysis tools (such as the pandas library of Python) to analyze the citation data, count the frequency and format characteristics of different bibliographic entries, and extract common reference bibliographic entries and format patterns.
[0051] Journal or publisher reviewers set rules through the visual interface of the interface: Based on the publication of big data mining analysis of citation data to build rules (regular expression writing) According to the requirements and analysis results obtained by proofreading, regular expressions are written for different types of references, such as the following journal article citation, the expression is as follows, where author is the author entry, title is the journal article title, etc. The matching entry can not only be used for format checking, but also for subsequent error identification.
[0052] (? <author> [^\.]+)(? <symb1> \s*\.)(? <title>[^\.]+)(?< / title> <symb2> \[)(? <type> J)(? <symb3> \])(? <symb4> \.)(? <pubunit> [^,]+)(? <symb5> ,)(? <pubyear> [^,]+)((((? <symb6> ,)(? <volume> [^\(^\s]+))?(? <symb7> \()(? <issue> [^\)]+)(? <symb8> \)))|((? <symb11> ,)(? <volume1> [^\(^\s]+)((? <symb12> \()(? <issue1> [^\)]+)(? <symb13> \)))?))(? <symb9> :)(? <page> [^\.]+)(? <symb10>. Customized construction of visualization interface rules: Support for publishers such as editors or reviewers to set personalized rules, build through the visualization interface, and generate regular expressions through the background program. Use front-end development technologies (HTML, CSS, JavaScript, etc.) and front-end frameworks (Vue.js or React) to build the visualization interface. Design interface elements such as rule display area, rule editing area, reference input area, and proofreading result display area, and implement corresponding interactive functions.
[0053] Rule matching test and optimization: Collect a batch of reference samples and use the written regular expressions to perform matching tests. Record the matching results and analyze the reasons for unsuccessful matching.
[0054] According to the test results, adjust and optimize the regular expressions, such as modifying character ranges, adding quantifiers, etc., to improve the matching accuracy and coverage of the rules.
[0055] Proofreading rule verification and storage: Large model verification: Input the constructed proofreading rules into the pre-trained large model and ask the model questions such as "According to the rule, does the following reference meet the requirements? [Reference example]".
[0056] Analyze the model's answer to determine the reasonableness and correctness of the rule. If the model's answer does not match the expectation, further analyze and modify the rule. Calculate the accuracy of the rule, and only when the accuracy is above 90% can the rule be preliminarily confirmed as meeting the requirements.
[0057] Rule storage: Use a database management system (MySQL or the domestically developed KBASE database) to create a database table for storing proofreading rules. The rule table can include fields such as rule number, rule description, applicable scope, and regular expression.
[0058] As shown in Figure 4 , insert the verified proofreading rules into the database table for subsequent calling and management As shown in Figure 5 , according to the rules, the references are proofread, and the bibliographic items matched by regular expressions are verified and corrected in the database: Format proofreading; The rule calling function is provided in the visualization interface, and the user can select the corresponding rule for reference literature proofreading according to the needs. When the user inputs the reference literature to be proofread, the corresponding rule is called from the database, and the regular expression is used to match and check the reference literature. The proofreading result is displayed in the proofreading result display area, and for the bibliographic item that does not meet the rule, the corresponding prompt and modification suggestion are given. According to the rule selected by the user, the format and bibliographic item of the reference literature are checked one by one. For example, whether the format of the author's name meets the rule, whether the journal name is correct, etc. For the bibliographic item that does not meet the rule, it is marked out in red font or other eye-catching way in the proofreading result display area, and modification suggestions are provided.
[0059] Bibliographic item inspection proofreading: According to the regular matching bibliographic item, the related information is retrieved from the academic database (CiteNet bibliographic record). For example, according to the author's name and the literature title, whether the literature exists and whether the journal name is correct are checked, and according to the literature name and the year period, whether the author is correct is judged according to the journal name. If the regular expression does not match the bibliographic item, the publication big model based on the Huazhi big model is used to identify the bibliographic item, especially to identify the author, and the effect of compatible English author is the most significant. After identifying the bibliographic item, the correctness of the bibliographic item is checked. If the database retrieval result is inconsistent with the information in the reference literature, the corresponding prompt is given in the proofreading result display area, reminding the user to further verify and modify.
[0060] For the method steps disclosed in the above embodiments, the method steps are described as a series of action combinations for the purpose of simple description, but those skilled in the art should know that the embodiments of the present application are not limited by the order of the described actions, because according to the embodiments of the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily necessary for the embodiments of the present application.
[0061] As Figure 6 shown, the present application also provides a reference literature proofreading system, comprising: A rule feature library construction module 201 configured to respond to the captured publishing agency format specification change event and the extracted citation feature data set from the academic database, fuse and analyze the format specification change event and the citation feature data set, and dynamically construct a rule feature library, wherein the rule feature library comprises a bibliographic item format distribution rule; The matching test module 202 is configured to, in response to an operation instruction of a user interface, convert the operation instruction into a regular expression rule set based on the distribution rule of the catalog item format in the rule feature library, generate a test sample set according to sample features in the rule feature library, perform matching test of the rule set and the sample set, analyze error type distribution of unmatched samples and optimize parameters of the regular expression; The comprehensive evaluation module 203 is configured to verify and calculate logical contradiction of rules in the rule set and coverage defects of samples in the sample set based on the optimized regular expression rule set and the sample set, generate a comprehensive evaluation value, and activate the rule when the comprehensive evaluation value is higher than a preset quality threshold; The rule updating module 204 is configured to convert the activated rule into a structured storage format according to feature dimensions, record rule generation time and change history by establishing a version control index, and set a change tracking mechanism to respond to and update the rule feature library; The literature proofreading module 205 is configured to, in response to a target reference, call the corresponding structured rule in the rule feature library to perform regular matching, perform multi-level verification of the regular matching result, and generate a proofreading report including the verification result.
[0062] It is worth noting that, although only some basic function modules are disclosed in the embodiments of the present application, it does not mean that the composition of the system is limited to the above basic function modules. On the contrary, the meaning expressed in the embodiments is that one or more function modules can be added to the above basic function modules by those skilled in the art in combination with existing technology to form infinite embodiments or technical solutions. That is, the system is open rather than closed, and the protection scope of the present application claim cannot be limited to the disclosed basic function modules. At the same time, for the convenience of description, the above device is described as various units and modules. Of course, the functions of the units and modules can be realized in the same software and / or hardware in the implementation of the present application.
[0063] As shown in Figure 7 The present application also provides an electronic device, comprising: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus; the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the reference literature proofreading method.
[0064] Figure 7 is a structural schematic diagram of an electronic device provided by the embodiments of the present application. As Figure 7 As shown in the structure, the electronic device provided in the embodiment of the present application comprises one or more processors 710 and a memory 720; the processor 710 in the electronic device can be one or more, Figure 7 The memory 720 is used for storing one or more programs; the one or more programs are executed by the one or more processors 710, so that the one or more processors 710 implement a reference proofreading method according to any one of the embodiments of the present application.
[0065] The electronic device can further comprise an input device 730 and an output device 740.
[0066] The processor 710, the memory 720, the input device 730 and the output device 740 in the electronic device can be connected through a bus or other means, Figure 7 For example, the connection through the bus is taken as an example.
[0067] The memory 720 in the electronic device is a kind of computer readable storage medium, which can be used to store one or more programs, and the program can be a software program, a computer executable program and a module, such as program instructions / modules of a reference proofreading method provided in the embodiment of the present application. The processor 710 executes the software program, instruction and module stored in the memory 720, thereby performing various function applications and data processing of the electronic device, i.e. implementing the reference proofreading method in the above method embodiment.
[0068] The memory 720 can include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required by a function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 720 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some examples, the memory 720 can further include a memory remotely arranged with respect to the processor 710, which can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0069] The input device 730 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the electronic device. The output device 740 can include a display device such as a display screen.
[0070] The present application also provides a computer readable storage medium storing a computer program executable by an electronic device, which causes the electronic device to perform the steps of a reference proofreading method when the computer program runs on the electronic device.
[0071] In particular, a computer storage medium of embodiments of the present application can adopt any combination of one or more computer readable medium. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.
[0072] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application. < / page> < / symb9> < / symb13> < / issue1> < / symb12> < / volume1> < / symb11> < / symb8> < / issue> < / symb7> < / volume> < / symb6> < / pubyear> < / symb5> < / pubunit> < / symb4> < / symb3> < / type> < / symb2> < / symb1> < / author> < / issue> < / volume>
Claims
1. A reference review method, characterized in that: include: In response to captured publishing institution format specification change events and citation feature datasets extracted from academic databases, the format specification change events and the citation feature datasets are fused and analyzed to dynamically construct a rule feature library, wherein the rule feature library includes the format distribution rules of the entry; In response to an operation instruction from a user interface, based on the format distribution pattern of the entry in the rule feature library, the operation instruction is converted into a regular expression rule set, and a test sample set is generated according to the sample features in the rule feature library, a matching test is performed between the regular expression rule set and the sample set, the error type distribution of the unmatched samples is analyzed, and the parameters in the regular expression rule set are optimized; Based on the optimized regular expression rule set and the sample set, verify and calculate the logical contradictions of the rules in the regular expression rule set and the coverage defects of the samples in the sample set, generate a comprehensive evaluation value, and activate the rule when the comprehensive evaluation value is higher than a preset quality threshold; Convert the activated rules into a structured storage format according to the feature dimension encoding, establish a version control index to record the rule generation time and change history, and simultaneously set up a change tracking mechanism to respond to and update the rule feature library; In response to the target reference document, the corresponding structured stored rules in the rule feature library are called to perform regular matching, a multi-level authenticity verification is performed on the regular matching result, and a review report including the verification result is generated.
2. A reference review method according to claim 1, characterized in that: In response to captured publishing institution format specification change events and citation feature datasets extracted from academic databases, the format specification change events and the citation feature datasets are fused and analyzed to dynamically construct a rule feature library, wherein the rule feature library includes the format distribution rules of the entry and further includes: Periodically scan the reference format pages on the publisher's official website and identify format specification changes by comparing the content hashes. Analyze the frequency and distribution of author names and main fields of title structure in the citation feature dataset of the academic database; The captured format specification change event is combined with the analyzed occurrence frequency and distribution pattern of the main fields to generate a rule feature library with a timestamp version identifier; When the format specification change event is identified, a difference comparison between the new and old versions is performed and a rule feature library update instruction is generated.
3. A reference review method according to claim 1, characterized in that: In response to an operation instruction from a user interface, based on the format distribution pattern of the entry in the rule feature library, the operation instruction is converted into a regular expression rule set, and a test sample set is generated according to the sample features in the rule feature library, a matching test is performed between the regular expression rule set and the sample set, the error type distribution of the unmatched samples is analyzed, and the parameters of the regular expression are optimized, further comprising: parsing the format distribution pattern of the entry in the rule feature library to generate a regular expression template including a named capture group; Based on the regular expression template, the user's drag and drop and check operation instructions in the interface are mapped to regular expression parameter configurations in real time; Constructing a test sample set covering multiple scenarios based on the sample features recorded in the rule feature library; A batch matching test is performed between the regular expression rule set and the sample set, error type distribution data of unmatched samples is counted, and character range definition and quantifier parameters of the named capture group are adjusted according to the error type distribution data of the unmatched samples.
4. A reference review method according to claim 1, characterized in that: Based on the optimized regular expression rule set and the sample set, verifying and calculating logical contradictions of rules in the regular expression rule set and coverage defects of samples in the sample set, generating a comprehensive evaluation value, and activating the rule when the comprehensive evaluation value is higher than a preset quality threshold, further comprising: Inputting the regular expression rule set and the associated sample set into a pre-trained verification model; Verify the logical contradictions of the rules in the regular expression rule set, and identify sample coverage blind spots and missing boundary conditions; Calculating a comprehensive evaluation value of the logic defect rate and sample coverage of the rules in the regular expression rule set, and activating and marking the rules when the comprehensive evaluation value is higher than a preset quality threshold; When the comprehensive evaluation value is lower than a preset quality threshold, error type distribution data is extracted for the corresponding rule, the process is rebuilt and the parameters of the regular expression are optimized accordingly.
5. A reference review method according to claim 1, characterized in that: In response to the target reference document, calling the corresponding structured stored rule in the rule feature library to perform regular matching, performing multi-level authenticity verification on the regular matching result, and generating a review report including the verification result, further comprising: Calling the corresponding activation rule in the rule feature library to perform regular expression matching according to the journal type of the target reference selected by the user; Extract metadata of successfully matched entries and initiate authenticity verification of fields such as title, author, and journal name in academic databases; Perform semantic completion on the fields that failed to match, generate candidate metadata, and then re-initiate authenticity verification; Generate a review report based on the authenticity verification results, including integrating format errors and authenticity deviation data, and generating a structured review report that includes error location, error type, and correction plan.
6. A reference review method according to claim 2, characterized in that: When the format specification change event is identified, performing a comparison between the new and old versions and generating a rule feature library update instruction further includes: Use text difference algorithm to locate the specific change points of the new and old format specification changes; Analyzing the impact of the change point on the format distribution pattern of the entry in the rule feature library; Reconstructing the format distribution model of the rule feature library according to the impact scope to generate a new version identifier; The version change record is written into the historical track of the rule feature library and an associated rule update notification is triggered.
7. A reference review method according to claim 5, characterized in that: The multi-level authenticity verification further includes: Basic field verification, including accurate field comparison of metadata of successfully matched entries with corresponding data in academic databases, and detection of differences in title spelling and author name order; Semantic completion verification, including completing the complete metadata of the entry through semantic analysis for fields that fail regular expression matching, and performing cross-database joint query verification; Variant compatibility checking, including multilingual spelling variant compatibility checking of the author name field, handling abbreviations, aliases, and cultural variations.
8. A reference review system, characterized in that: include: a rule feature library construction module configured to respond to captured publishing institution format specification change events and citation feature datasets extracted from academic databases, fuse and analyze the format specification change events and the citation feature datasets, and dynamically construct a rule feature library, wherein the rule feature library includes a distribution pattern of entry formats; a matching test module configured to respond to an operation instruction from a user interface, convert the operation instruction into a regular expression rule set based on the format distribution pattern of the entry in the rule feature library, generate a test sample set based on the sample features in the rule feature library, perform a matching test between the regular expression rule set and the sample set, analyze the error type distribution of unmatched samples, and optimize the parameters of the regular expression; a comprehensive evaluation module configured to verify and calculate, based on the optimized regular expression rule set and the sample set, logical contradictions of rules in the regular expression rule set and coverage defects of samples in the sample set, generate a comprehensive evaluation value, and activate the rule when the comprehensive evaluation value is higher than a preset quality threshold; A rule update module is configured to convert the activated rules into a structured storage format according to the feature dimension encoding, establish a version control index to record the rule generation time and change history, and simultaneously set up a change tracking mechanism to respond to and update the rule feature library; The document review module is configured to respond to the target reference document, call the corresponding structured stored rules in the rule feature library to perform regular matching, perform multi-level authenticity verification on the regular matching results, and generate a review report including the verification results.
9. An electronic device, characterized in that: include: A processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that It stores a computer program that can be executed by an electronic device. When the computer program runs on the electronic device, the electronic device executes the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Reference literature check method
CN106326199A
Reference reference vacancy checking method and device, equipment and storage medium
CN113505570A
Reference review method and device, medium and equipment
CN118261115A
Method and system for appraising the extent to which a publication has been reviewed by means of a peer-review process
US20110270847A1
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1