Literature Misquotation Identification and Dissemination Tracking System Based on Retracted Papers
By constructing a citation map subnet for retracted papers, identifying misguided paths and calculating the constraint value of the propagation structure, and dynamically tracking the propagation paths of retracted papers, the problem of low propagation efficiency of retracted papers in the existing technology is solved, and efficient dissemination of retracted papers and optimized resource allocation is achieved.
Patent Information
- Application Number
- CN202510654171.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing technology cannot effectively identify and track the misinformation of retracted papers in academic literature, resulting in the inefficient dissemination of retracted papers, the inability to optimize resource allocation, and the failure to pass important statements to key institutions in a timely manner.
By constructing a citation map subnet with the withdrawal paper as the starting point, combining the withdrawal time and effective time intervals of citation behavior, identifying misguided paths, calculating the constraint value of the propagation structure and the changes in the node citation frequency, dynamically tracking the propagation path, and optimizing the push batch and institutional coverage of the withdrawal statement.
Accurate and rapid dissemination of revoked statements has been achieved, the efficiency of responsiveness of revoked statements has been improved, and important statements have been given priority to the delivery of important statements to key institutions.
Smart Images

Figure CN120181078B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of citation network analysis, and particularly to a system for identifying and tracking the misquotation and dissemination of documents based on retracted papers. Background Art
[0002] The technical field of citation network analysis includes the modeling and analysis of citation relationships between academic documents, involving the extraction of citation structures between documents, the tracking of citation paths, data processing and information mining based on citation maps. The core content of the technical field is to dynamically analyze the citation connection relationships between nodes by constructing a directed graph structure between documents, identify knowledge dissemination paths, mine citation rules, analyze research hotspots or misquotation behaviors. The citation network analysis technology represents documents and citation relationships through the nodes and edges of the graph structure, constructs a citation network relying on the literature database and citation information, and uses graph algorithms to achieve a structured understanding and dynamic tracking of document citation behaviors, realizing multiple applications such as misquotation identification, knowledge diffusion modeling, and information dissemination path reconstruction.
[0003] Among them, the system for identifying and tracking the misquotation and dissemination of documents based on retracted papers refers to a technical solution for structurally identifying and tracking the information dissemination chain formed by the continued citation of retracted academic papers in subsequent documents. The patent theme covers the identification, classification, and tracking of the residual citations of retracted papers in the citation network. By constructing a sub-network of the citation map starting from the retracted papers, extracting the dissemination path nodes based on the retraction status mark and citation time node, identifying the misquoted documents and marking the misquotation occurrence nodes and time paths. The system uses citation frequency sorting and node impact factor calculation to automatically identify high-frequency dissemination nodes and continuously spreading institutions, and constructs a list of target institutions for preferentially pushing retraction statements based on the dissemination chain structure to support the directional dissemination of retraction statements. The misquotation detection relies on the time-annotated citation comparison method to identify the citation entries after the time node based on the retraction time. The dissemination tracking relies on the graph traversal algorithm and node path backtracking to construct the citation path chain in chronological order to achieve the identification and annotation of the misquotation diffusion trajectory.
[0004] The misquotation identification and dissemination tracking technology of traditional literature relies on the static citation network constructed based on literature databases and citation information, lacking dynamic monitoring and real-time misquotation tracking for retracted papers, resulting in the ineffective handling of the misquoted dissemination of retraction information in academic literature. Relying on the static representation of nodes and edges in the graph structure, it can only provide a rough analysis of citation relationships and cannot achieve dynamic tracking of the citation paths of retracted papers. There are timeliness issues in misquotation identification and dissemination path reconstruction, and it is impossible to effectively identify misquotation behaviors and prevent the spread of misquotations according to actual time nodes. After retraction, some papers are still cited by newly published papers, and traditional technologies are difficult to identify and eliminate incorrect citations in a timely manner, leading to the spread of misquotation behaviors. When pushing retraction statements, it is impossible to dynamically adjust the push batches according to the actual influence and coverage of the papers, resulting in a low dissemination efficiency of retraction statements, failing to optimize resource allocation, and causing some important retraction statements not to be preferentially transmitted to key institutions or fields, affecting the timeliness of handling retraction statements. Summary of the Invention
[0005] An object of the present invention is to solve the deficiencies existing in the prior art and propose a literature misquotation identification and dissemination tracking system based on papers after retraction.
[0006] To achieve the above object, the present invention adopts the following technical solutions: The literature misquotation identification and dissemination tracking system based on papers after retraction includes:
[0007] A misquotation identification module, which is used to obtain the publication and retraction times of the literature, set an effective citation detection duration interval, extract the citation time of each citation behavior in the citation chain, compare the citation time with the effective citation detection duration interval, identify the misquotation path, and obtain the misquotation section identification value;
[0008] A path truncation module, which is used to remove the misquotation path according to the misquotation section identification value, extract the number of nodes and citation edges in each layer of the remaining path, calculate the node density value and hierarchical span value of each layer, and obtain the propagation structure constraint value;
[0009] A node extraction module, which is used to extract the citation frequency change value of each node in the path within a continuous time period according to the propagation structure constraint value, identify the transition nodes and analyze the proportion of transition nodes in each layer, calculate the diffusion degree according to the number of transition nodes and propagation levels of the retracted paper in each path, and obtain the paper influence level;
[0010] A density sorting module, which is used to extract the set of propagation paths corresponding to each paper retraction entity according to the paper influence level, calculate the occurrence frequency and proportion of each node in the path, and sort the set of paths of each paper retraction entity in combination with the diffusion degree to obtain the retraction entity list.
[0011] As a further solution of the present invention, the mis-citation section identification value includes the time overrun citation number, the total legal citation volume, and the mis-citation path marking matrix. The propagation structure constraint degree value includes the citation sparsity coefficient, the structural distribution span, and the node hierarchy composition ratio. The specific paper influence level is the diffusion index, the path transition frequency, and the cross-layer node activity rate. The retraction entity list includes the entity number, the sorting position, and the path coincidence degree weight.
[0012] As a further solution of the present invention, the mis-citation identification module includes:
[0013] A time interval setting sub-module for obtaining the publication and retraction times of the literature, obtaining the valid citation time, and setting the valid citation detection duration interval;
[0014] A time matrix construction sub-module for extracting the citation time data of each citation behavior in the citation chain according to the valid citation detection duration interval, using the time range of the cited literature as the column vector and the citation behavior time as the row vector to construct a time matrix, and obtaining the citation behavior time matrix;
[0015] A mis-citation path identification sub-module for comparing the citation time with the detection duration interval according to the citation behavior time matrix, identifying the mis-citation path, and extracting the corresponding node numbers and path numbers to obtain the mis-citation section identification value.
[0016] As a further solution of the present invention, the path truncation module includes:
[0017] A path elimination sub-module for eliminating the mis-citation path according to the mis-citation section identification value, extracting the number of nodes and the number of citation edges at each level in the remaining paths, and obtaining the structure data set after path elimination;
[0018] A structure parameter extraction sub-module for calculating the node density value within each level based on the structure data set after path elimination according to the number of nodes and the number of citation edges at each level in the remaining paths, extracting the level span value of each level, and obtaining the path level density and span value;
[0019] A structure constraint construction sub-module for calculating the propagation structure constraint degree value according to the path level density and span value, and according to the node density value and the level span value of each level.
[0020] As a further solution of the present invention, the node extraction module includes:
[0021] A frequency change extraction sub-module for extracting the citation times of each node in the path within a continuous time period according to the propagation structure constraint degree value, constructing a citation frequency time series corresponding to each node, and calculating the citation frequency increment in adjacent time periods to obtain the node citation frequency change sequence;
[0022] A transition node identification sub-module, which is used to analyze according to the node reference frequency change sequence and the frequency increment of each node in a continuous time period, identify transition nodes, and obtain the hierarchical transition node proportion value according to the number of transition nodes and the number of nodes in each path level;
[0023] A node extraction sub-module, which is used to calculate the diffusion degree according to the hierarchical transition node proportion value, the number of transition nodes and the propagation level of the retracted paper in each path, and obtain the paper influence level.
[0024] As a further solution of the present invention, the density sorting module includes:
[0025] A path set extraction sub-module, which is used to extract the propagation path set corresponding to each retracted paper entity according to the paper influence level, identify the node numbers in each path, and obtain path node data;
[0026] A node frequency statistics sub-module, which is used to calculate the occurrence frequency of the same node in each path according to the path node data, obtain the ratio between the number of repeated nodes and the number of path nodes, and generate node frequency statistics data;
[0027] A priority sorting sub-module, which is used to calculate the priority score of the retracted entity according to the node frequency statistics data and in combination with the diffusion degree index, sort and obtain a list of retracted entities.
[0028] As a further solution of the present invention, the system further includes:
[0029] A statement push module, which is used to extract the path coverage quantity and the number of node attribution institutions corresponding to each retracted entity according to the list of retracted entities, calculate the coverage range of each retracted entity in the propagation map, compare with a preset path coverage level threshold, divide the retracted entities into multiple push batches, and monitor the acceptance and processing status of each entity for the statement to obtain a statement push grouping result;
[0030] The statement push grouping result specifically refers to the push batch number, the institutional response status, and the statement receiving node proportion.
[0031] As a further solution of the present invention, the statement push module includes:
[0032] A coverage quantity extraction sub-module, which is used to extract the propagation path set corresponding to each retracted entity according to the list of retracted entities, identify the path coverage quantity of each path, and generate path coverage data;
[0033] The affiliated institution number calculation sub-module is used to obtain the number of affiliated institutions of each node of each retracted entity in the propagation graph according to the path coverage data by extracting the number of affiliated institutions of the nodes associated with each path, and generate the data of the number of affiliated institutions of the nodes.
[0034] The push batch division sub-module is used to calculate the coverage range of each retracted entity in the propagation graph according to the data of the number of affiliated institutions of the nodes, combine the preset path coverage level threshold, divide the retracted entities into multiple push batches, and monitor the acceptance and processing status of each entity's statement, and generate the statement push grouping result.
[0035] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0036] In the present invention, by combining the publication time and retraction time of the retracted paper and comparing with the effective time interval of the citation behavior, the effective identification of academic mis-citation is realized. Using the propagation structure constraint value and the data of the change in the node citation frequency, the propagation path of the retracted paper is dynamically tracked, the high-frequency propagation nodes and the continuously spreading institutions are timely identified, and combined with the path coverage quantity and the data of the number of affiliated institutions of the nodes, the push batches are efficiently divided to ensure that the retraction statement can be preferentially transmitted to important institutions, optimize the push efficiency, and by calculating the influence degree of each path, the retraction statement is accurately and quickly spread, significantly improving the response efficiency of the retraction statement. Description of the Drawings
[0037] Figure 1 is the system flow chart of the present invention;
[0038] Figure 2 is the mis-citation identification module flow chart of the present invention;
[0039] Figure 3 is the path truncation module flow chart of the present invention;
[0040] Figure 4 is the node extraction module flow chart of the present invention;
[0041] Figure 5 is the density sorting module flow chart of the present invention;
[0042] Figure 6 is the statement push module flow chart of the present invention. Detailed Embodiments
[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0045] See also Figure 1 , the literature miscitation identification and dissemination tracking system based on retracted papers includes:
[0046] The miscitation identification module is used to obtain the publication and withdrawal time of the document, set the effective citation detection time interval, extract the citation time of each citation behavior in the citation chain, use the time range of the cited document as the column vector, and the citation behavior time as the row vector, construct a time matrix, compare the citation time and the effective citation detection time interval, identify the miscitation path, and obtain the miscitation segment identification value;
[0047] Effective citation detection time interval: indicates the time window in which the cited document is reasonably cited, that is, the legal citation interval from the publication time to the withdrawal time;
[0048] The path truncation module is used to remove the misquoted path according to the misquoted segment identification value, extract the number of nodes and the number of reference edges at each level in the remaining path, record the node density value and level span value of each level, build the structural constraint parameters according to the reference density and level span, and obtain the propagation structure constraint value;
[0049] Node density value: the ratio of the number of node reference edges to the number of nodes in each layer;
[0050] Layer span value: indicates the difference between the deepest and shallowest layers in the path, reflecting the extension range of the path structure;
[0051] The node extraction module is used to extract the citation frequency change value of each node in the path within a continuous time period according to the propagation structure constraint value, identify the transition node by comparing it with the preset citation frequency increase threshold, and record the corresponding path level. By analyzing the proportion of transition nodes in each level, according to the number of transition nodes and propagation levels involved in each path of the retracted paper, the diffusion degree index is calculated to obtain the paper impact level;
[0052] A density sorting module, which is used to extract the set of propagation paths corresponding to each paper retraction entity according to the paper influence level, identify the node numbers in each path, calculate the occurrence frequency of the same node in each path, obtain the ratio between the number of repeated nodes and the total number of path nodes, and combine the diffusion degree index of the paper to perform priority sorting on the path sets of each paper retraction entity to obtain a list of retraction entities;
[0053] A statement push module, which is used to extract the path coverage quantity and the number of node attribution institutions corresponding to each retraction entity according to the list of retraction entities, calculate the coverage range of each retraction entity in the propagation graph, divide the retraction entities into multiple push batches by comparing with a preset path coverage level threshold, and monitor the acceptance and processing status of each entity for the statement to obtain the statement push grouping result;
[0054] Path coverage level threshold: Set a critical value for the combination of the number of propagation paths corresponding to a unit retraction statement and the number of institutional nodes involved, which is used to hierarchically judge the push priority of the statement.
[0055] The misquotation section identification values include the number of time-limit exceeded citations, the total number of legal citations, and the misquotation path marking matrix. The propagation structure constraint values include the citation sparsity coefficient, the structure distribution span, and the node hierarchy composition ratio. The paper influence level specifically includes the diffusion index, the path transition frequency, and the cross-layer node activity rate. The list of retraction entities includes the entity number, the sorting position, and the path coincidence degree weight. The statement push grouping result specifically refers to the push batch number, the institutional response status, and the statement receiving node ratio.
[0056] Please refer to Figure 2 , the misquotation identification module includes:
[0057] A time interval setting sub-module, which is used to obtain the publication and retraction times of the literature, obtain the valid citation time, and set the duration interval for valid citation detection;
[0058] First, collect the publication time and retraction time corresponding to the literature number registered in the literature database, and associate them with the citation relationship of the literature in the citation network through number indexing. For example, if the publication time of literature A001 is 2017 and the retraction time is 2021, then its effective citation interval is from 2017 to 2021. Further, construct an interval vector with consecutive years for this time interval, that is, the vector interval is {2017, 2018, 2019, 2020, 2021}. This vector is used for subsequent comparison and judgment with the actual citation time. At this time, if the citation time is 2022, it can be determined that it exceeds the retraction period. According to the definition of citation nature, this citation is a mis-citation item. In actual implementation, to construct a more explicit definition standard for this interval, a citation interval offset tolerance needs to be set. The tolerance value δ is set to 1 year to absorb the time difference in database record updates or the time error in delayed release of retraction announcements. Therefore, the effective detection time interval is set as: starting from the publication year, to the retraction year plus the tolerance year, that is, the vector interval is extended to {2017, 2018, 2019, 2020, 2021, 2022}. This interval is used as the basis for subsequent citation time comparison. All records whose citation time is not in this vector are determined to be invalid citation path data. In this setting, the δ value is obtained through statistics of the update lag period of typical literature, and its empirical value is set to 1 year. As shown in Table 1, the citation time of literature A001 is 2022, and after tolerance correction, it is within the effective interval, so it does not belong to mis-citation. However, without tolerance correction, it needs to be excluded, and finally the effective citation time interval value is generated.
[0059] Table 1 Information Table of Literature Citation Time
[0060] ;
[0061] As shown in Table 1, the time information of some literatures is listed, which has been used to detect the construction of the effective citation time interval.
[0062] The time matrix construction sub-module is used to extract the citation time data of each citation behavior in the citation chain according to the effective citation detection duration interval, take the time range of the cited literature as the column vector, take the citation behavior time as the row vector, construct a time matrix, and obtain the citation behavior time matrix;
[0063] First, read the time fields in the citation records of each node in sequence according to the path structure in the citation network. This time field is derived from the official publication time of the cited literature, and this time will be filled into the corresponding row position of the matrix. Each row represents the timestamp of a citation behavior, while each column represents the time range of the cited literature. The column vector is generated in sequence from the starting year of the literature to the withdrawal time plus the tolerance year, constituting the time column range of the matrix. If a certain citation occurs in the second year after the literature is withdrawn, then this time is marked as "+2" offset year in the row vector. The constructed time matrix is a two-dimensional array structure, where each element represents the year difference or positional relationship between the "citation behavior time" and the "citation reference range". For example, if the citation behavior occurs in 2019 and the withdrawal time of this literature is 2021, then its relative position is 2 years before withdrawal, corresponding to column 2021 and row 2019 in the matrix, and the filled value is "-2". After all positions in the time matrix are filled, the time matrix structure is constructed, and time region screening and visual judgment can be performed on it. Finally, the citation behavior time matrix is obtained.
[0064] The mis-citation path identification sub-module is used to compare the citation time and the detection duration range according to the citation behavior time matrix, identify the mis-citation path, extract the corresponding node numbers and path numbers, and obtain the mis-citation section identification value;
[0065] For each row value in the time matrix, read its citation time and match and judge it with the vector of the valid citation time range recorded in the matrix column vector. When the citation time is not included in the valid citation interval vector, that is, the column index does not appear for the corresponding year of this row, then record the path number to which the current row belongs and the citation node number, and identify it as a mis-citation record. In actual implementation, the citation path number is defined by the traversal order of the path structure, and the node number is determined by the sequence number of the node in the path structure. For example, the citation path number is P05, the citation node number is N03, the citation time is 2023, the withdrawal time of the literature is 2021, the tolerance is 1 year, and the valid citation interval is from 2017 to 2022. Then this citation behavior is in the invalid section. Therefore, P05 and N03 are marked into the mis-citation path set. In the actual structure record, the mis-citation records form a binary index set according to the path number and the node number, which is convenient for subsequent path structure elimination operations. Statistically summarize all records that meet the mis-citation criteria, output the total number of mis-citation records, the number of paths, the number of nodes, and generate the mis-citation section identification value.
[0066] Please refer to Figure 3 and the path truncation module includes:
[0067] The path elimination sub-module is used to eliminate the mis-citation path according to the mis-citation section identification value, extract the number of nodes and the number of citation edges at each level in the remaining paths, and obtain the structure data set after path elimination;
[0068] First, based on the path numbers and node numbers recorded in the mis-citation section identification values generated by the previous module, perform an identification and filtering operation on the path set. When performing path filtering, logically screen the path data set according to the numbers, and reconstruct the number mapping table of the remaining paths to ensure the integrity of the subsequent index structure. Then, traverse the structure of the remaining path set, layer by layer read the number of nodes and the corresponding number of reference edges in each propagation layer. In path number P01, layer L1 contains 10 nodes and 12 reference edges. In P02, layer L2 contains 15 nodes and 22 reference edges, and so on. The structure parameters of all paths will be aggregated into a unified path layer statistical structure for subsequent extraction of density values and span values. As shown in the data listed in Table 2, after clearly identifying the layer number, the number of nodes, and the number of reference edges, output the structure set data and obtain the structure data set after path elimination.
[0069] The structure parameter extraction sub-module is used to calculate the node density value in each layer based on the structure data set after path elimination, according to the number of nodes and the number of reference edges in each layer of the remaining paths, extract the layer span value of each layer, and obtain the path layer density and span values.
[0070] In actual operation, the node density value is defined as the ratio of the number of reference edges in this layer to the number of nodes, indicating the degree of information interaction in this layer. In layer L1, there are 12 reference edges and 10 nodes, so the density is 1.2. In L2, there are 22 reference edges and 15 nodes, and the density is 1.467. In L3, there are 10 reference edges and 8 nodes, and the density is 1.25. For the calculation of the layer span value, it refers to the number difference between the starting node and the ending node in this layer plus 1. In Table 2, the span of L1 is 3, L2 is 5, and L3 is 2. This parameter reflects the layer extension range of the path in the graph spectrum. Aggregate the node density and span values into a structure vector to provide support parameters for the subsequent structure constraint degree, and finally obtain the path layer density and span values.
[0071] Table 2 Path Structure Statistical Table
[0072] ;
[0073] Table 2 lists the structure parameter data of typical paths in each layer.
[0074] The structure constraint construction sub-module is used to calculate the propagation structure constraint degree value according to the path layer density and span values, according to the node density value and layer span value of each layer, using the formula:
[0075] ;
[0076] Calculate the propagation structure constraint degree value;
[0077] Among them, represents the propagation structure constraint degree value, represents the number of reference edges of the th layer, represents the number of nodes of the th layer, represents the average number of nodes of all valid layers, and are the maximum level number and the minimum level number in the path respectively, represents the total number of valid levels in the path, is the index variable of the path level;
[0078] According to the path level density and span value, according to the node density value and level span value of each level, use the formula:
[0079] ;
[0080] Calculate the propagation structure constraint degree value. In the implementation of the calculation, substitute the parameters in Table 2 into the formula for numerical calculation. The number of reference edges of each level is {12, 22, 10}, and the number of nodes is {10, 15, 8}. First, calculate the average number of nodes , the number of layers , then:
[0081] ;
[0082] ;
[0083] ;
[0084] ;
[0085] Among them, the propagation structure constraint degree value is a quantitative index used to measure whether the propagation structure in a single reference path has stability, convergence and distribution rationality. The larger the parameter value, the more concentrated the node connections in the path, the more balanced the hierarchical structure, the stronger the propagation density, and the more controllable the propagation characteristics. The smaller the value, the larger the propagation level span, the looser the structure or the skewed distribution of nodes between layers in the path, which is likely to form a hidden danger of misinformation diffusion. The result shows that when the current path structure has a relatively compact level span and there are certain fluctuations in the node density, its propagation structure constraint degree value is 0.854. This value is used as a measurement factor for structural integrity in the analysis of the stability of the propagation path structure, and finally the propagation structure constraint degree value is obtained.
[0086] Please refer to Figure 4 , the node extraction module includes:
[0087] The frequency change extraction sub-module is used to extract the citation frequency of each node in the path within consecutive time periods according to the propagation structure constraint value, construct the citation frequency time series corresponding to each node, calculate the citation frequency increment between adjacent time periods, and obtain the node citation frequency change series;
[0088] According to the propagation structure constraint value, extract the citation frequency of each node in the path within consecutive time periods, construct the citation frequency time series corresponding to each node. First, obtain the citation data set of each path. At each node, record its citation frequency in each time period. Then, according to the citation records of each node, generate the time series of each node. This series shows the change in the citation frequency of the node in each time period. For example, the citation frequency of node A is 5 times from time period T1 to T2, and 8 times from T3 to T4, while that of node B is 2 times from T1 to T2 and 6 times from T3 to T4. By constructing the time series of the node citation frequency, the system will perform frequency increment analysis on each node, calculate the citation frequency increment between adjacent time periods of each node, and then obtain the frequency change series of each node, as shown in Table 3. The table records the citation frequencies of each node in different time periods and indicates the citation frequency increment of each node. Finally, obtain the node citation frequency change series of each node.
[0089] Table 3 Node Citation Frequency Time Series Table
[0090] ;
[0091] Table 3 lists the citation frequencies and frequency increments of nodes A, B, and C in different time periods.
[0092] The transition node identification sub-module is used to analyze according to the node citation frequency change series and the frequency increment of each node within consecutive time periods to identify the transition nodes, and obtain the hierarchical transition node proportion value according to the number of transition nodes and the number of nodes in each path level;
[0093] According to the sequence of node citation frequency changes, analyze based on the frequency increment of each node within consecutive time periods to identify transition nodes. First, the system compares the frequency increments within adjacent time periods according to the frequency change sequence of each node. If the increment value of a certain node between time periods exceeds the set threshold, it is considered that the node has undergone a transition. Transition nodes usually show a sharp increase in the number of citations. The frequency increment of node A is 3, the increment of node B is 4, and the increment of node C is 4. Assuming the set threshold is 3, then nodes A, B, and C are all transition nodes. Next, the system will further calculate the proportion of transition nodes in each layer based on the number of transition nodes and the total number of nodes in each path layer. This proportion reflects the distribution of transition nodes within that layer. Assuming there are 10 nodes in layer L1, and 2 of them are transition nodes, then the proportion of transition nodes in this layer is 0.2. While in layer L2, there are 12 nodes, and 3 of them are transition nodes, then the proportion of transition nodes is 0.25. In this way, the system can obtain the proportion of transition nodes in each layer and finally obtain the value of the layer transition node proportion.
[0094] The node extraction sub-module is used to calculate the diffusion degree and obtain the paper influence level according to the layer transition node proportion value, based on the number of transition nodes and the propagation layer in each path of the retracted paper, using the formula:
[0095] ;
[0096] Calculate the diffusion degree and obtain the paper influence level;
[0097] Among them, represents the diffusion degree of the paper path, represents the number of transition nodes, represents the total number of propagation layers, represents the total number of nodes in the path, represents the maximum layer span of the transition nodes, is the index of the paper number;
[0098] According to the layer transition node proportion value, calculate the diffusion degree and obtain the paper influence level according to the number of transition nodes and the propagation layer in each path of the retracted paper, using the formula:
[0099] ;
[0100] Calculate the diffusion degree and obtain the paper influence level. First, the system calculates based on the number of transition nodes, the propagation layer, the total number of nodes in the path, and the maximum layer span of the transition nodes in each path. Assume that the number of transition nodes in the path is 5, the number of propagation layers is 4, the total number of nodes is 20, and the maximum layer span of the transition nodes is 2, then:
[0101] ;
[0102] Among them, the diffusion degree of a paper reflects the influence of the transition nodes in the paper's dissemination path. Transition nodes refer to the nodes where the citation frequency changes drastically, and these nodes are usually the "key" nodes with greater influence in the dissemination path. The number of levels and the maximum level span are factors that measure the influence diffusion in the dissemination path. The more levels there are, the wider the spread of influence; and the greater the level span of the transition nodes, the deeper the spread of influence. By combining these factors, the ultimate goal of the formula is to evaluate the diffusion degree of the paper in the path, reflecting the breadth, depth, and influence of its dissemination. The calculation result shows that the obtained diffusion degree is 1.5, and this value is used to measure the influence degree of the paper in the dissemination path. A higher diffusion degree value indicates that the paper has a higher influence in the network, and finally, the paper's influence level is obtained.
[0103] Please refer to Figure 5 , the density sorting module includes:
[0104] A path set extraction sub-module, which is used to extract the dissemination path set corresponding to each paper retraction entity according to the paper influence level, identify the node numbers in each path, and obtain path node data;
[0105] According to the paper influence level, extract the dissemination path set corresponding to each paper retraction entity. The system first determines the corresponding dissemination path according to the influence level of the paper retraction entity, and then assigns a path set to each retraction entity. Each path set consists of multiple paths, and each path contains multiple nodes. The system further extracts the node numbers in each path to ensure the uniqueness and correctness of the nodes. The identification of the node numbers is paired with the level data of each path. The node numbers in each path are sorted according to their positions in the path to ensure that the order of the nodes in the path is consistent with the dissemination path. For example, assume that the dissemination paths involved in the retracted paper A include paths P1, P2, and P3, where path P1 contains nodes N1, N2, N3, path P2 contains nodes N2, N4, N5, and path P3 contains nodes N3, N5, N6. The node numbers extracted by the system are N1, N2, N3, N4, N5, N6, and a path node data set is created for these nodes. During this process, the system can accurately extract all the paths and their node data corresponding to each retracted paper and obtain path node data.
[0106] A node frequency statistics sub-module, which is used to calculate the occurrence frequency of the same node in each path according to the path node data, obtain the ratio between the number of duplicate nodes and the number of path nodes, and generate node frequency statistics data;
[0107] Based on the path node data, calculate the occurrence frequency of the same node in each path. The system analyzes each path, counts the repetition frequency of each node in the path, and focuses on calculating the frequency of each node in multiple paths. Through this frequency statistics, the system can understand the importance and influence of each node in the overall propagation path. If node N2 appears once in path P1 and once in path P2, then the frequency of node N2 is 2. If path P1 contains 3 nodes and P2 contains 3 nodes, the system calculates the occurrence ratio of node N2 in these two paths as 2 / 6 = 0.33. Similarly, for node N3, it appears once in path P1 and once in path P3, with a frequency of 2, and the total number of nodes in path P1 and P3 is also 3, and the frequency ratio is also calculated as 2 / 6 = 0.33. Through these steps, the system will generate a frequency statistics dataset for each node, and the system further calculates the proportion of each node in the path set to obtain the node frequency statistics data.
[0108] The priority sorting sub-module is used to calculate the priority score of the retraction entity according to the node frequency statistics data, combined with the diffusion degree index, using the formula:
[0109] ;
[0110] Calculate the priority score of the retraction entity, sort and obtain the list of retraction entities;
[0111] Among them, is the priority score of the retraction entity, is the node number, is the path number, is the total number of nodes, is the total number of paths, is the node 's occurrence frequency, is the path 's sum of occurrence frequencies of all nodes, is the node 's weight, is the node 's propagation level number, is the node 's diffusion degree index, is the maximum value among all node diffusion degrees;
[0112] According to the node frequency statistics data, combined with the diffusion degree index, using the formula:
[0113] ;
[0114] Calculate the priority score of the retraction entity. The system first calculates the frequency value 、weight value 、propagation level number and the degree of diffusion . Through these values, the system can evaluate the importance and priority of each node during the propagation process. For example, assume that the frequency of node N2 is 4, the weight is 1.5, the number of propagation levels is 3, and the degree of diffusion is 0.4. The sum of the frequencies of all nodes is 6. Then the priority score of node N2 is calculated as follows:
[0115] ;
[0116] Among them, the priority score represents the relative importance or priority of each retracted paper entity in its set of propagation paths. This score is calculated by comprehensively considering the frequency, weight, propagation level, and degree of diffusion of the node, and finally reflects the influence of the retracted entity in the academic communication graph. Its main purpose is to prioritize the retracted papers among multiple propagation paths to ensure that during the push of retraction statements, those paper entities that have a greater impact on academic communication and a higher spread are processed first. The higher the score of a retracted entity, the more high-frequency, widely spread, and high-level nodes it contains in its propagation path, indicating that the retracted paper has a higher propagation influence and requires an earlier and prioritized release of the retraction statement. The system will calculate the priority scores of all retracted entities and sort them according to the scores to generate a list of retracted entities, which will provide a basis for the push priority of retraction statements.
[0117] Table 4 Data table for calculating node priority scores
[0118] ;
[0119] As shown in Table 4, the table shows parameters such as different node numbers, their frequencies, weights, the number of propagation levels, and the degree of diffusion, as well as the priority scores of each node calculated by the formula.
[0120] Please refer to Figure 6 . The statement push module includes:
[0121] A coverage quantity extraction sub-module, which is used to extract the set of propagation paths corresponding to each retracted entity according to the list of retracted entities, identify the path coverage quantity of each path, and generate path coverage data;
[0122] First, it is necessary to obtain the set of dissemination paths for each retraction entity according to the retraction entity list. For each retraction entity, the system extracts all the associated paths of the entity, and then identifies the specific nodes and structures of each path. The key to this process is the statistics of the path coverage quantity, which is obtained by comparing the relevance of all nodes in each path in the retraction entity graph. For example, assume that a retraction entity involves dissemination paths containing 5 different paths, and the system will identify the coverage quantity of each path respectively. The path coverage quantity is the intersection of all nodes in these paths, and the interconnection situation and frequency between nodes are considered during the calculation. The basic data used in this process includes information such as the path node data, citation data of each retraction entity, and the occurrence frequency of each node within the path. Combining these data, path coverage data is finally obtained, providing basic data support for subsequent operations. Through the execution of this sub-module, the system generates a path coverage data set, which provides the core basis for subsequent node attribution calculation and push batch division.
[0123] The affiliated institution number calculation sub-module is used to obtain the number of affiliated institutions of each retraction entity in the dissemination graph by extracting the number of affiliated institutions of the nodes associated with each path according to the path coverage data, and generate the data of the number of affiliated institutions of the nodes.
[0124] First, based on the path coverage data, the node information associated with each path is extracted, and the institution to which each node belongs is identified from it. The number of affiliated institutions of the nodes is obtained by calculating the number of institutions to which these nodes belong. For example, if path A includes nodes N1, N2, and N3, and N1 and N2 belong to institutions I1 and I2 respectively, and N3 belongs to I3, then the number of affiliated institutions of this path is 3. This step needs to ensure the integrity of the path coverage data and accurately extract the affiliated institutions of each node by comparing the relevance of each node. During the calculation process, the system verifies the attribution of each node and conducts statistics through database queries and set intersections. Finally, a data set containing the number of affiliated institutions of the nodes corresponding to each retraction entity is generated, providing the required institution quantity data for the division of the push batch. In practical applications, ensuring the accuracy of the calculation of the number of affiliated institutions of the nodes is crucial for the success of subsequent push tasks.
[0125] The push batch division sub-module is used to calculate the coverage range of each retraction entity in the dissemination graph according to the data of the number of affiliated institutions of the nodes, combined with the preset path coverage level threshold, divide the retraction entities into multiple push batches, and monitor the acceptance and processing status of each entity's statement, generating the statement push grouping result.
[0126] First, the system calculates the coverage range of each retraction entity in the dissemination graph based on the aforementioned data of the number of affiliated institutions of nodes and in combination with a preset path coverage level threshold. The coverage range is determined by comparing the ratio of the number of affiliated institutions of nodes to the number of path coverages. If this ratio is greater than a certain set threshold, it is considered that the retraction entity has a high coverage range. The system will automatically divide the retraction entities into multiple push batches according to this determination. For example, if the coverage range value of a certain retraction entity is greater than 0.8 and it belongs to a high-priority path, this entity will be included in the first push list. On the other hand, if the coverage range value is low, it will be divided into secondary push batches. By sorting the push priorities of each entity, the finally generated push grouping data helps relevant personnel to accurately process retraction statements, ensuring that each retraction entity can be processed in a timely manner. The core of push batch division is path coverage data and the number of affiliated institutions of nodes. The combination of these two can accurately reflect the actual influence range of retraction entities, thus optimizing the efficiency and effect of statement pushing.
[0127] Table 5 Data Table for Dividing Push Batches of Retraction Entities
[0128] ;
[0129] As shown in Table 5, the table shows the number of path coverages, the number of affiliated institutions of nodes, the coverage range value, and the push batches of retraction entities.
[0130] The above is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A literature misquotation identification and dissemination tracking system based on retracted papers, characterized in that The system comprises: The miscitation identification module is used to obtain the publication and withdrawal time of the document, set the effective citation detection time interval, extract the citation time of each citation behavior in the citation chain, compare the citation time with the effective citation detection time interval, identify the miscitation path, and obtain the miscitation segment identification value; A path truncation module is used to remove the misleading path according to the misleading segment identification value, extract the number of nodes and the number of reference edges in each layer of the remaining path, calculate the node density value and the layer span value of each layer, and obtain the propagation structure constraint value; A node extraction module is used to extract the citation frequency change value of each node in the path in a continuous time period according to the propagation structure constraint value, identify the transition nodes and analyze the transition node ratio of each layer, calculate the diffusion degree according to the number of transition nodes and propagation level of the retracted paper in each path, and obtain the paper impact level; The density sorting module is used to extract the propagation path set corresponding to each paper retraction entity according to the paper impact level, calculate the frequency and proportion of each node in the path, sort each paper retraction entity path set based on the diffusion degree, and obtain the retraction entity list.
2. The literature misquotation identification and dissemination tracking system based on the retracted papers according to claim 1, wherein The mis-citation segment identification value includes the number of citations that exceed the time limit, the total number of legal citations, and the mis-citation path marking matrix. The propagation structure constraint value includes the citation sparsity coefficient, the structure distribution span, and the node level composition ratio. The paper impact level is specifically the diffusion index, the path transition frequency, and the cross-layer node activity rate. The retracted entity list includes the entity number, ranking, and path overlap weight.
3. The literature misquotation identification and dissemination tracking system based on the retracted papers according to claim 1, wherein The misleading identification module comprises: The time interval setting submodule is used to obtain the publication and withdrawal time of the document, obtain the effective time of the citation, and set the effective citation detection time interval; A time matrix construction submodule is used to extract the citation time data of each citation behavior in the citation chain according to the effective citation detection time interval, take the time range of the cited document as the column vector, take the citation behavior time as the row vector, construct a time matrix, and obtain the citation behavior time matrix; The misreference path identification submodule is used to compare the reference time and the detection duration interval according to the reference behavior time matrix, identify the misreference path, extract the corresponding node number and path number, and obtain the misreference section identification value.
4. The literature misquotation identification and dissemination tracking system based on retracted papers according to claim 3, wherein The path truncation module comprises: A path elimination submodule is used to eliminate misleading paths according to the misleading segment identification value, extract the number of nodes and the number of reference edges at each level in the remaining paths, and obtain a structure data set after path elimination; A structural parameter extraction submodule is used to calculate the node density value in each level based on the structure data set after path elimination and the number of nodes and reference edges in each level in the remaining path, extract the level span value of each layer, and obtain the path level density and span value; The structural constraint construction submodule is used to calculate the propagation structure constraint value according to the path level density and span value, and according to the node density value and level span value of each level.
5. The literature misquotation identification and dissemination tracking system based on retracted papers according to claim 4, wherein The node extraction module comprises: The frequency change extraction sub-module is used to extract the citation frequency of each node in the path within consecutive time periods according to the propagation structure constraint value, construct the citation frequency time series corresponding to each node, calculate the citation frequency increment between adjacent time periods, and obtain the node citation frequency change sequence; The transition node identification sub-module is used to analyze according to the node citation frequency change sequence and the frequency increment of each node within consecutive time periods, identify the transition nodes, and obtain the hierarchical transition node proportion value according to the number of transition nodes and the number of nodes in each path level; The node extraction sub-module is used to calculate the diffusion degree according to the hierarchical transition node proportion value, the number of transition nodes and the propagation level of the retracted paper in each path, and obtain the paper influence level.
6. The literature misquotation identification and dissemination tracking system based on retracted papers according to claim 5, wherein The density sorting module includes: The path set extraction sub-module is used to extract the propagation path set corresponding to each paper retraction entity according to the paper influence level, identify the node numbers in each path, and obtain the path node data; The node frequency statistics sub-module is used to calculate the occurrence frequency of the same node in each path according to the path node data, obtain the proportion between the number of repeated nodes and the number of path nodes, and generate the node frequency statistics data; The priority sorting sub-module is used to calculate the priority score of the retraction entity according to the node frequency statistics data and in combination with the diffusion degree index, sort and obtain the retraction entity list.
7. The literature misquotation identification and dissemination tracking system based on retracted papers according to claim 1, wherein The system further includes: The statement push module is used to extract the path coverage number and the number of node belonging institutions corresponding to each retraction entity according to the retraction entity list, calculate the coverage range of each retraction entity in the propagation map, compare with the preset path coverage level threshold, divide the retraction entities into multiple push batches, and monitor the acceptance and processing status of each entity for the statement, and obtain the statement push grouping result; The statement push grouping result specifically refers to the push batch number, the institutional response status, and the statement receiving node proportion.
8. The misquotation identification and dissemination tracking system for papers after retraction according to claim 7, wherein The statement push module includes: The coverage number extraction sub-module is used to extract the propagation path set corresponding to each retraction entity according to the retraction entity list, identify the path coverage number of each path, and generate the path coverage data; The belonging institution number calculation sub-module is used to obtain the number of node belonging institutions of each retraction entity in the propagation map by extracting the number of node belonging institutions associated with each path according to the path coverage data, and generate the node belonging institution number data; The push batch division sub-module is used to calculate the coverage range of each retraction entity in the propagation map according to the node belonging institution number data and in combination with the preset path coverage level threshold, divide the retraction entities into multiple push batches, and monitor the acceptance and processing status of each entity for the statement, and generate the statement push grouping result.