Data processing method, system and equipment for realizing full self-service of paper collection and introduction proof
Through the process of data cleaning, splitting, matching and verification of the paper citation service system, the problem that the existing system cannot support the fully self-service citation proof service is solved, and the full self-service citation proof of papers is achieved and the high accuracy and reliability of the data is achieved.
Patent Information
- Application Number
- CN202411951332.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-23
AI Technical Summary
The existing paper inspection and citation service system cannot support the fully self-service intake and certification service process, and there are problems such as low matching accuracy of institutions and authors, inconsistent data processing process, and lack of data verification process.
Through the processes of cleaning and integration of data from different institutions, data splitting, data matching and data verification and data correction, a full self-service service for citation certificates is realized. The specific steps include selecting unique identifiers or combined identifiers for data merging, splitting the paper structure, performing field matching, using ambiguous matrix analysis method for data verification, and improving the accuracy of data matching through multiple rounds of correction algorithms.
It realizes full self-service service for essay citation certificates, ensures the integrity, authenticity and reliability of data, supports dynamic data updates, and meets the needs of fully self-service certification issuance, accurate and effective data, real-time update of dynamic information, reporting anti-counterfeiting and authenticity verification, and reporting secure storage.
Smart Images

Figure CN120030276A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and more specifically, to a data processing method, system and device for realizing full self-service of paper citation proof. Background Art
[0002] Paper retrieval and citation service is a knowledge service provided by information service institutions such as libraries to users. It is a service that provides users with search certification reports during scientific research processes such as project application, award application, talent plan application, graduation defense and professional title evaluation. The content of retrieval and citation service is relatively stable. The retrieval and citation service content of domestic universities mainly includes SCI, SSCI, A&HCI, EI, CPCI-S, CPCI-SSH, CSSCI, CSCD, SCOPUS, Chinese core journals, PUBMED, etc., impact factors and classifications, and highly cited papers. Some researchers have thought about, designed or practiced the construction of retrieval and citation service systems. The previous research mainly focused on the design and practice of retrieval and citation management systems, improving the efficiency of retrieval and citation services with the help of existing technologies, and the design and implementation of retrieval and citation service systems based on existing systems.
[0003] In terms of the design of the receiving and citation service and management system, the system generally includes module functions such as bibliographic information submission, target database adaptation, data retrieval and acquisition, and the export of inclusion and citation proof reports. Some systems also include functions such as user management, workload statistics, and cost calculation. The emergence of related systems has provided convenience for users' entrustment and librarians' work, but librarians are still required to handle tasks manually and interact with users in multiple links. In terms of using existing technologies to improve the efficiency of receiving and citation services, it mainly includes tools or technologies such as VBA, database technology, Python, etc. to issue receiving and citation reports and improve the efficiency of receiving and citation services. However, existing technologies have problems such as low accuracy in matching institutions and authors, insufficient consistency in data processing, and missing data verification processes, and cannot support the design and implementation of a fully self-service receipt and citation service process.
[0004] Therefore, developing a data processing method and system that can achieve full self-service for paper citation proof is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In view of the problems existing in the prior art, the present invention proposes a data processing method, system and equipment for realizing full self-service of paper citation proof, including the processes of cleaning and integrating data from different institutions, data splitting, data matching and data verification, and data correction, so that the integrity, authenticity and reliability of the data are guaranteed, and the data is dynamically updated to support the full self-service of paper citation proof.
[0006] To achieve the above objectives, in a first aspect, the present invention provides a data processing method for realizing a full self-service service for paper citation proof, comprising the following steps:
[0007] Cleaning and integration of data from different institutions, by selecting unique identifiers or combined identifiers to merge data from different institutions;
[0008] Data splitting: split the papers into different fields according to the obtained paper structure;
[0009] Data matching: the split fields are matched individually or in combination with the unit standard name, personnel name, and journal type to match the paper to a specific unit or individual, and classify the paper according to different journal types;
[0010] Data verification, including data verification of included papers, data verification of merged papers, data verification of author splitting, and data verification of author matching; data verification is carried out using the confusion matrix analysis method, and the first full data verification and subsequent additional data verification are carried out;
[0011] Data correction, based on the results of the data verification indicators, multiple rounds of corrections are made to the data matching algorithm.
[0012] The data processing method ensures the intelligence and accuracy of data processing through more effective data integration, splitting, matching and verification algorithms, meeting the needs of fully self-service certificate issuance, accurate and effective certificate data, real-time updating of dynamic information, verifiable report anti-counterfeiting, and secure report storage.
[0013] Furthermore, the combined identifier is a plurality of data including but not limited to title, journal, and issue; and several pieces of paper data with the same combined identifier are merged.
[0014] Furthermore, the fields include but are not limited to institution, author, address, journal, and subject.
[0015] Furthermore, in the data matching, the authors are matched by matching institutions and authors at the same time; institutions are matched by institution names in different languages; journal attributes are matched by journal names. There are certain difficulties in matching authors with the same name. The system matches institutions and authors at the same time, which can effectively improve the accuracy.
[0016] Furthermore, the authors are matched by matching the authors and addresses of the papers with the names and addresses of the units in the personnel list. In order to better identify the authors of the papers, all addresses in the papers are standardized to be consistent with the units in the personnel list. Using the method of paper author + paper address and name + unit in the personnel list for identification can greatly improve the recognition accuracy and reduce recognition redundancy, but it will also result in some scholars whose papers are not signed in a standardized manner being unable to match the person themselves.
[0017] Furthermore, the confusion matrix analysis method marks the data verification results of each or each indicator dimension as true positive (TP), true negative (TN), false positive (FP) and false negative (FN); the indicators of data verification include: accuracy , Negative element accuracy , overall accuracy , overall error rate , coverage , false positive rate , F value ,in is the weight of CR, which can be 0.5 or 1.
[0018] In a second aspect, the present invention provides a data processing system for realizing a full self-service of paper citation proof, which is arranged at the back end of a data collection system for a full self-service of paper citation proof, and comprises:
[0019] Data deduplication module, which combines data from different institutions by selecting unique identifiers or combined identifiers;
[0020] The data splitting module splits the paper into different fields according to the obtained paper structure;
[0021] Data matching module, the split fields are matched individually or in combination with the unit standard name, personnel name, and journal type, the papers are matched to specific units and individuals, and the papers are classified according to different journal types;
[0022] The data verification module includes data verification of included papers, data verification of merged papers, data verification of author splitting, and data verification of author matching. The data verification is carried out by using the confusion matrix analysis method, and the first full data verification and subsequent additional data verification are adopted;
[0023] The data correction module performs multiple rounds of corrections on the data matching algorithm based on the results of the data verification indicators.
[0024] Finally, the present invention provides a device for implementing a fully self-service paper citation certificate, including a certificate issuing system, wherein the data management of the certificate issuing system executes the data processing method for implementing a fully self-service paper citation certificate as described above.
[0025] The most important thing in building a self-service system for receipt and citation proof is to ensure data quality, intelligent data processing and verification, and accuracy. Therefore, data processing is the core process for realizing self-service receipt and citation. Compared with the prior art, the present invention has the following technical effects:
[0026] (1) The present invention realizes the full self-service issuance of certificates through the data processing process of cleaning and integrating data from different institutions, data splitting, data matching, data verification, and data correction, without waiting for manual review;
[0027] (2) The data matching of the present invention takes into account factors that make the data inaccurate, such as duplicate data and duplicate authors. Data verification is carried out from four angles: paper inclusion data verification, paper merging data verification, author splitting data verification, and author matching data verification, which can ensure the intelligence and accuracy of data processing and verification;
[0028] (3) The data processing method of the present invention feeds back the data verification results to the data matching module, and corrects the algorithm multiple times to achieve the purpose of data accuracy;
[0029] (4) The data processing method of the present invention supports dynamic data updating and data quality checking, so that the data in the database has integrity, authenticity and reliability, thereby ensuring data quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 The figure is a flow chart of a data processing method in one embodiment of the present invention.
[0031] Figure 2 This is a system architecture diagram of a fully self-service device for paper citation proof in one embodiment of the present invention. DETAILED DESCRIPTION
[0032] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0033] In the description of the present invention, the execution order of the actions, steps, etc. in the devices and methods shown in the claims, specifications and drawings can be implemented in any order as long as there is no special explicit limitation on the order and as long as the output of the previous processing is not used in the subsequent processing.
[0034] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the modules / units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0035] Example 1
[0036] See also Figure 1 This embodiment provides a data processing method and system for realizing a full self-service service for paper citation proof, wherein the data processing method includes the following steps:
[0037] Cleaning and integration of data from different institutions, by selecting unique identifiers or combined identifiers to merge data from different institutions;
[0038] Data splitting: split the papers into different fields according to the obtained paper structure;
[0039] Data matching: the split fields are matched individually or in combination with the unit standard name, personnel name, and journal type to match the paper to a specific unit or individual, and classify the paper according to different journal types;
[0040] Data verification, including data verification of included papers, data verification of merged papers, data verification of author splitting, and data verification of author matching; data verification is carried out using the confusion matrix analysis method, and the first full data verification and subsequent additional data verification are carried out;
[0041] Data correction, based on the results of the data verification indicators, multiple rounds of corrections are made to the data matching algorithm.
[0042] The data processed in this embodiment meets the following characteristics: ① Data collection integrity. Data collection completion is the most basic requirement for self-service data collection and citation. The types of certificates commonly issued include SCI, SSCI, A&HCI, CPCI-S, CPCI-SSH, CSCD, EI and CSSCI, etc.; ② Data acquisition stability. Manual collection and citation services can realize the database opening as soon as it is received, which is a challenge for self-service collection and citation services. The system needs to have stable data acquisition capabilities and can quickly process new papers. Therefore, the continuity and stability of data acquisition are also necessary prerequisites for self-service collection and citation services; ③ Intelligent data analysis. Data analysis includes data splitting and merging. As the number of papers is always increasing, the workload of data analysis is getting bigger and bigger. The data analysis process needs to be intelligent to meet user needs more quickly and keep the data formats of various sources consistent. ④ Real-time data update. Data real-time mainly includes two parts. One is the real-time data related to the characteristics of the paper itself, such as whether the paper changes from priority publication to formal publication, or whether it is withdrawn, etc. The second is the real-time data related to the external characteristics of the paper, such as the number of citations, the number of uses and the journal impact factor. The real-time data of both parts is very important. The four processes of data collection, data acquisition, data analysis and data update are interconnected and interdependent. The data quality of each link is crucial to the realization of self-service retrieval and citation.
[0043] The front end of data processing is data collection. As an example, here, data collection refers to the process of collecting paper data from a database to the local. Once the data comes from multiple databases, the first step of data processing involves merging different data sources by selecting unique identifiers such as DOI, or combined identifiers such as title, journal, issue, etc. to merge the data. Secondly, data splitting: According to the obtained paper structure, the paper is split according to institutions, authors, addresses, journals, disciplines, etc., so as to better match the papers in different dimensions. Next, data matching is performed: the split institutional data, author data, journal data, etc. are matched with the standardized names of units (such as secondary units such as colleges of universities), names of institutional personnel, journal types, etc., so as to match the papers to specific units and individuals, and classify the papers according to different types. Author matching has certain challenges, and matching authors with the same name has certain difficulties. The system matches institutions and authors at the same time, which can effectively improve the accuracy. Institutional matching is relatively accurate, as long as the institution names in other languages are translated and merged. Journal attributes are usually matched, such as impact factors and partitions. The journal matching accuracy is relatively high, but there are a small number of name changes or abbreviations, which require manual intervention during the data processing process.
[0044] In order to better identify the authors of the papers, all addresses in the papers were standardized to be consistent with the units in the personnel list. Using the method of paper author + paper address and name + unit in the personnel list for identification can greatly improve the recognition accuracy and reduce recognition redundancy, but it will also result in some scholars with irregular signatures being unable to match their own papers.
[0045] Data verification process:
[0046] In order to enable users to issue certificates by themselves after logging into the system without waiting for manual review, the data review process is set as the data processing flow. The staff needs to verify the data collection, splitting, field indexing, etc. The data verification workload is large, and the first full data verification and subsequent additional data verification are adopted.
[0047] (1) Data verification method
[0048] The system mainly uses confusion matrix analysis for data verification. The data verification results of each item or each indicator dimension are marked as true positive (TP), true negative (TN), false positive (FP) and false negative (FN). Taking the SCI paper author matching as an example, the papers that are actually teacher-student papers and accurately matched by the system developer are true positives, the papers that are not teacher-student papers and are not matched by the system developer are true negatives, the papers that are not teacher-student papers but are matched by the system developer are false positives, and the papers that are actually teacher-student papers but are not matched by the system developer are false negatives. Due to the large number of papers that are not from Tongji, the true negative data of some dimensions are no longer obtained (or can be understood as ∞). Except for the author matching data verification, the TN indicator is not calculated for other data.
[0049] (2) Data verification results
[0050] The data quality is verified from four perspectives: paper inclusion data verification, paper merging data verification, author splitting data verification, and author matching data verification. Data verification is a process of correcting the matching algorithm and discovering problems. Data verification usually requires multiple rounds. Each time, the results of the previous round of data verification are analyzed and correction suggestions are made. Through continuous iteration, the integrity and accuracy of the data are improved. The main verification indicators are shown in Table 1.
[0051] Table 1 Data verification related indicators
[0052]
[0053] Taking the data of 2019 as an example, the accurate data and the data obtained by the data processing system were compared. The data of 2019 are accurate data confirmed by the author. Table 2 is a data verification sample. The position marked with "-" in the table means that the indicator does not need to be analyzed. As can be seen from Table 2, the system can generally handle the data collection, paper merging data and data splitting well. In terms of author matching, the author matching results are not ideal mainly because English authors have full names, abbreviations, and different characters. Based on the verification results, the Scopu, EI and other databases were further analyzed, the various deformations of the author's name were supplemented, and the paper address was translated. The method of matching the author and the address was adopted to improve the integrity and accuracy of the data. Compared with manual verification, the system's verification of data is more about the verification of existing algorithms. By verifying the data, the algorithms of each link of data processing are corrected to improve the accuracy of data processing. For exceptions that the algorithm cannot handle, the manual processing or modification function is retained through system design.
[0054] Table 2 Comparison of real data of the unit and data processing system data
[0055]
[0056] From the above description, it can be seen that the data processing method described in the present invention ensures the intelligence and accuracy of data processing through more effective data integration, splitting, matching and verification algorithms, ensures the accuracy and integrity of data, realizes automatic processing and dynamic updating of data, and meets the needs of fully self-service issuance of certificates, accurate and valid certificate data, real-time updating of dynamic information, verifiable anti-counterfeiting of reports, and safe storage of reports.
[0057] Example 2
[0058] See also Figure 2 This embodiment provides a device for implementing a fully self-service paper citation certificate, including a certificate issuance system, and the data management of the certificate issuance system executes the data processing method for implementing a fully self-service paper citation certificate as described in Example 1.
[0059] The process of the certificate issuance system includes:
[0060] 1. New construction certificate;
[0061] 2. Determine whether the paper has been claimed; if not, enter the database, execute the data processing flow, update the database, and claim the paper;
[0062] 3. Determine whether it is the school's achievement; if yes, proceed to the next step; if no, exit;
[0063] 4. Select a template.
[0064] 5. Select a paper;
[0065] 6. Select fields;
[0066] 7. Generate proof.
[0067] The essence of this technical solution or the part that contributes to the prior art or the part of this technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic system (which may be a personal computer, a server, or a network system, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0068] The data processing method for realizing the full self-service of paper citation proof can be embodied in the form of a computer program product or a software functional unit. If the data processing method for realizing the full self-service of paper citation proof is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0069] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0070] Those skilled in the art should understand that those skilled in the art can implement variations by combining the prior art and the above embodiments, which will not be described in detail here. Such variations do not affect the essential content of the present invention, and will not be described in detail here.
[0071] The above describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above-mentioned specific embodiments, and the systems and structures that are not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, or modify them into equivalent embodiments of equivalent changes, which does not affect the essential content of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention are still within the scope of protection of the technical solutions of the present invention.
Claims
1. A data processing method for realizing full self-service of paper citation proof, characterized in that: The following steps are involved: Cleaning and integration of data from different institutions, by selecting unique identifiers or combined identifiers to merge data from different institutions; Data splitting: split the papers into different fields according to the obtained paper structure; Data matching: the split fields are matched individually or in combination with the unit standard name, personnel name, and journal type to match the paper to a specific unit or individual, and classify the paper according to different journal types; Data verification, including paper inclusion data verification, paper merging data verification, author splitting data verification, and author matching data verification; The data verification was conducted by using confusion matrix analysis method, and the verification was conducted by first verifying the entire data and then verifying the additional data; Data correction, based on the results of the data verification indicators, multiple rounds of corrections are made to the data matching algorithm.
2. The data processing method for realizing full self-service of paper citation proof according to claim 1 is characterized in that: The combined identifier is a plurality of data including but not limited to title, journal, and issue; and several paper data with the same combined identifier are merged.
3. The data processing method for realizing full self-service of paper citation proof according to claim 1 is characterized in that: The fields include but are not limited to institution, author, address, journal, and subject.
4. The data processing method for realizing full self-service of paper citation proof according to claim 1 is characterized in that: In the data matching, institutions and authors are matched simultaneously to match authors; institutions are matched through institution names in different languages; and journal attributes are matched through journal names.
5. The data processing method for realizing full self-service of paper citation proof according to claim 4 is characterized in that: Authors are matched by simultaneously matching the paper authors and paper addresses with the names and unit addresses in the personnel list.
6. The data processing method for realizing full self-service of paper citation proof according to claim 1 is characterized in that: The confusion matrix analysis method marks the data verification results of each or each indicator dimension as true positive (TP), true negative (TN), false positive (FP) and false negative (FN); the indicators of data verification include: accuracy , Negative element accuracy , overall accuracy , overall error rate , coverage , false positive rate , F value ,in is the weight of CR, which can be 0.5 or 1.
7. A data processing system for realizing full self-service of paper citation proof is provided at the back end of the data collection system for full self-service of paper citation proof, and is characterized by: include: Data deduplication module, which combines data from different institutions by selecting unique identifiers or combined identifiers; The data splitting module splits the paper into different fields according to the obtained paper structure; Data matching module, the split fields are matched individually or in combination with the unit standard name, personnel name, and journal type, the papers are matched to specific units and individuals, and the papers are classified according to different journal types; The data verification module includes data verification of included papers, data verification of merged papers, data verification of author splitting, and data verification of author matching. The data verification is carried out by using the confusion matrix analysis method, and the first full data verification and subsequent additional data verification are adopted; The data correction module performs multiple rounds of corrections on the data matching algorithm based on the results of the data verification indicators.
8. A device for realizing full self-service of paper citation proof, including a proof issuing system, characterized in that: The data management of the certificate issuance system executes the data processing method for realizing full self-service of paper citation certificate as described in any one of claims 1 to 6.