Adaptive hyperbolic exponential-based scientific paper influence evaluation method and system
By using the adaptive transcendence index algorithm and fitting the probability density function with a large language model and maximum likelihood estimation, the problem of unfairness and complexity in paper evaluation in existing technologies is solved, and the fair comparison and real-time impact assessment of scientific papers across different disciplines is realized.
Patent Information
- Application Number
- CN202311600014.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2043-11-28
AI Technical Summary
Existing technologies are insufficient to accurately reflect the contributions of a single scientific paper across different disciplines in real time, and traditional evaluation indicators suffer from unfair comparisons and computational complexity.
The adaptive transcendence index algorithm is adopted to determine the paper topic through a large language model, retrieve papers in the same field in real time, calculate the adaptive transcendence index, and use the maximum likelihood estimation method to fit the probability density function to generate a continuous exponential distribution, thereby achieving adaptiveness and interpretability evaluation.
It enables fair comparisons across different disciplines. The adaptive transcendence index is adaptive and interpretable, and can calculate the impact of papers in real time, avoiding evaluation bias caused by different citation counts.
Smart Images

Figure CN117573807B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of bibliometrics, and particularly relates to a quality evaluation algorithm and system for scientific papers. BACKGROUND
[0002] Bibliometrics is a discipline that uses mathematical and statistical methods to quantitatively analyze literature, especially scientific papers. In the 1960s, with the proposal of the Science Citation Index (SCI) by Eugene Garfield, the citation network analysis by Derek John de Solla Price, and the definition of bibliometrics by Alan Pritchard, bibliometrics was born. Since its inception, the theories and methods of bibliometrics have been constantly updated with the development of technology and social needs. Among them, the number of citations has been widely cited as a classic quantitative evaluation indicator. It is the number of times a paper or academic achievement is cited by other papers and academic achievements. A higher number of citations usually has a higher influence. However, since the citation number evaluation method based on citation analysis method sets some assumptions and premises at the beginning, this method is also subject to certain limitations in actual use. For example, the citation number evaluation method considers all citations as the same, which ignores the differences brought by different citations. Obviously, being cited by a high-quality literature and being cited by a relatively low-quality literature should be different in value, and the citation number indicator treats them equally, only measuring the influence by the number of citations.
[0003] Based on the above shortcomings, different scholars at home and abroad have tried to improve the citation number indicator. The most mainstream improvement idea is to use the PageRank sorting algorithm to score citations in order to try to eliminate or reduce the bias brought by different citations to some extent. PageRank is an algorithm developed by Larry Page and Sergey Brin of Google Company, which is used to evaluate the importance and ranking of web pages. PageRank algorithm considers that the importance of a web page is determined by the importance of other web pages. Papers can be compared to web pages, literature citations can be compared to web page links, and literature citations can be compared to web page links. Then the mathematical expression is as shown in formula 1:
[0004]
[0005] Wherein, PR(A) is the PageRank value of the paper, is the cited literature of the paper is the number of literature citations of the paper is the total number of papers This is the damping factor, which is usually set to 0.85. For citing papers in the paper Thesis The rank value contributed to paper A is also known as the mini-PageRank. Although PageRank-based citation analysis methods are designed to consider equal importance, the algorithm requires datasets with more than two levels of citation relationships to achieve stable performance. In practical applications, it is often difficult to obtain data on multi-level citations, which can negatively impact the algorithm's performance. Furthermore, while PageRank converges to a stable non-normalized value after multiple iterations, the mathematical nature of non-normalization can lead to significant differences in PageRank values for papers making similar contributions across different fields.
[0006] Some scholars also use the impact factor of the conference proceedings / journal in which the paper was published (see Equation 2) as an indicator for evaluating the paper.
[0007]
[0008] in, Total citation frequency This represents the total number of publications. However, in reality, papers published in a conference proceedings / journal often encompass multiple research areas, and some multidisciplinary conference proceedings / journals even span multiple disciplines, making this value unsuitable for direct interdisciplinary comparisons. To overcome this issue, the Chinese Academy of Sciences adopts the Journal Transcendence Index as an evaluation index for current journals, with the mathematical expression:
[0009]
[0010] in, For journals On the topic The surpassing index, Document type Indicates journal The theme is , type The index represents a set of papers. The formula implies the probability that a randomly selected paper from a journal has a higher citation count than a paper from another journal with the same topic and document type. This calculation method avoids the inconsistency between the numerator and denominator and effectively addresses the skewness problem. However, the calculation process is discretized, and due to significant differences in citation counts, cases may arise where papers with different citation counts have the same transcendence index. Furthermore, it should be noted that this index is primarily designed for journals and is not suitable for quantitatively evaluating individual papers.
[0011] The existing evaluation scheme relies on simple citation times, or is based on PageRank algorithm, or directly uses conference set / journal impact factor for evaluation, which is difficult to accurately reflect the contribution of a single paper to the field in real time. Within the existing cognitive range, there is no real-time updateable adaptive transcendence index evaluation algorithm and system for a single paper. SUMMARY
[0012] The purpose of the present application is to overcome the deficiencies in the prior art, fill the gaps in related technology, and provide a scientific paper influence real-time evaluation algorithm and system based on adaptive transcendence index, which sets up four modules of paper theme determination module, real-time retrieval module, adaptive transcendence index calculation module and visualization output module. The paper theme determination module analyzes the given paper title and abstract based on a large language model to obtain one or more paper themes. The real-time retrieval module obtains the relevant data of the paper from the publicly available database according to the obtained paper theme. The adaptive transcendence index calculation module calculates the adaptive transcendence index of the specified paper according to the returned relevant data. The visualization output module uses Python language to render a PDF file with pictures and text and a clean interface based on the previously obtained and calculated paper related data.
[0013] The purpose of the present application is realized by the following technical solutions:
[0014] A scientific paper influence evaluation method based on adaptive transcendence index, comprising:
[0015] Determine the paper theme; obtain the theme of the paper using a large language model according to the title, abstract and preset prompt of the specified paper;
[0016] Obtain real-time retrieval results according to the paper theme; retrieve the paper set of the same theme as the given paper in the publicly available paper database through the paper theme, obtain the most relevant paper and the corresponding citation times of each paper in the paper set; ;
[0017] Calculate the adaptive transcendence index according to the retrieval results; calculate the discrete citation frequency distribution of k papers and regard it as a probability mass function, i.e. , according to the selected paper and the corresponding citation times of each paper ; , where represents the number of papers with citation times; use maximum likelihood estimation method to fit , to obtain the probability density function (PDF) of continuous exponential distribution ; wherein f(x) represents the probability density at value x, are the results of maximum likelihood estimation, representing the scale parameter that controls the scaling and the shift parameter that controls the displacement position, respectively; calculate The definite integral on the interval The adaptive transcendental index The adaptive transcendental index is an embodiment of the influence of the paper, and its value range is 0 to 1, 0 represents that the paper has no influence in the field, and 1 represents that the paper is very recognized in the field.
[0018] The application also provides a scientific paper influence evaluation system based on an adaptive transcendental index, comprising:
[0019] A paper theme determination module is configured to obtain the theme of the paper by using a large language model according to the title, abstract and preset prompt of the specified paper;
[0020] A real-time retrieval module is configured to retrieve a set of papers belonging to the same theme as the given paper in a publicly accessible paper database in real time according to the paper theme, and obtain the most relevant k papers in the set and the corresponding number of citations of each paper. An adaptive transcendental index calculation module is configured to calculate the discrete citation frequency distribution of the k papers according to the selected number of citations of each paper , and regard it as a probability mass function, that is,
[0021] , wherein represents the number of papers with n citations; fit using maximum likelihood estimation to obtain the probability density function of the continuous exponential distribution ; wherein, f(x) represents the probability density at value x, are the results of maximum likelihood estimation, representing the scale parameter that controls the scaling and the shift parameter that controls the displacement position, respectively; calculate The definite integral on the interval The adaptive transcendental index ; A visual output module is configured to render a PDF file in real time by using Python language according to the number of citations of the paper obtained and calculated previously and the adaptive transcendental index.
[0022]
[0023] The application further provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the scientific paper influence evaluation method based on adaptive transcendence index when executing the program.
[0024] The application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is executable on a processor to implement the steps of the scientific paper influence evaluation method based on adaptive transcendence index.
[0025] Compared with the prior art, the technical scheme of the application has the beneficial effects that:
[0026] 1. The adaptive transcendence index has adaptability, and can generate the transcendence index of a specified article in the same field without manually pre-specifying keywords.
[0027] 2. The adaptive transcendence index has good mathematical properties. The application converts the probability mass function into the probability density function, and then utilizes the continuity of the probability density function to avoid the case that different citation times correspond to the same transcendence index.
[0028] 3. The adaptive transcendence index has interpretability, and the adaptive transcendence index has high interpretability, i.e., the probability that the citation times of the article are greater than the citation times of any article in the same field in the same year.
[0029] 4. The application overcomes the unfair comparison phenomenon of traditional evaluation indexes (such as citation times and journal impact factors) in different fields to a certain extent. The adaptive transcendence index can adaptively fit the probability density function for different disciplines, so that the adaptive transcendence indexes of papers in different disciplines can be compared more fairly.
[0030] 5. The adaptive transcendence index is calculated based on the citation times, has the characteristics of small calculation cost, and can obtain the adaptive transcendence index of the paper in real time. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 FIG. 1 is a flowchart of the scientific paper influence evaluation algorithm based on adaptive transcendence index provided in Embodiment 1;
[0032] Figure 2 FIG. 2 is a working schematic diagram of the scientific paper influence evaluation system based on adaptive transcendence index provided in Embodiment 2;
[0033] Figure 3 FIG. 3 is a PDF page screenshot output by the scientific paper influence evaluation system. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0035] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0036] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and it should also be understood that when the terms "comprise" and / or "include" are used in the specification, there is a presence of a feature, step, operation, device, component, and / or combinations thereof.
[0037] Term explanation:
[0038] Prompt: A type of input used to instruct an AI model on what action to take or output to generate when performing a specific task.
[0039] Large language model: A large language model refers to a deep learning model trained using a large amount of text data.
[0040] Database: A repository that organizes, stores, and manages data according to a data structure.
[0041] Probability mass function: A function in probability theory used to describe the probability distribution of a discrete random variable. For a discrete random variable, the probability mass function gives the probability of each possible value.
[0042] Probability density function: A function in probability theory used to describe the probability distribution of a continuous random variable. For a continuous random variable, the probability density function represents the probability density of the random variable within a certain range of values.
[0043] Maximum likelihood estimation: A commonly used method of parameter estimation, used to estimate the parameter value that is most likely to produce the observed data from the observed data.
[0044] Definite integral: Can be seen as the "cumulative summation" of a function over a given interval.
[0045] Example 1
[0046] like Figure 1 As shown in this embodiment, a scientific paper impact evaluation algorithm based on an adaptive transcendence index includes:
[0047] S1. Determine the paper's topic; the process of determining the paper's topic involves using a large language model to obtain the topic based on the given paper's title, abstract, and preset prompts. The table below shows the prompts used to determine the paper's topic. These prompts, along with the paper's title and abstract, are fed into the large language model; the titles and abstracts are enclosed in curly braces. This embodiment uses either gpt-3.5-turbo or gpt-4 as the large language model.
[0048] S2. Obtain real-time search results based on the paper topic; the specific process is as follows: based on the paper topic, search in real-time in an open-source paper database for a set of papers belonging to the same topic as the given paper, and obtain the most relevant papers within that set. The paper and each paper Corresponding citation count .
[0049] S3. Calculate the adaptive transcendence index based on the search results; the specific process is as follows: based on the selected... The paper and each paper Corresponding citation count The discrete citation frequency distribution of k papers is calculated. This is treated as a probability mass function, i.e. ,in Indicates having The number of cited papers. Taking into account... The number of cited papers generally exhibits an exponential decay relationship with the number of citations, which can be fitted using the maximum likelihood estimation method. This yields the probability density function (PDF) of the continuous exponential distribution. Where f(x) represents the probability density function at the value x. The results of the maximum likelihood estimation method fitting are given, where represent the scale parameter controlling scaling and the position parameter controlling displacement, respectively. Finally, In the interval definite integrals on This is the adaptive transcendence index we are looking for.
[0050] The advantages of the above technical solution are that the adaptive transcendence index proposed in this invention is adaptive, capable of adaptively fitting probability density functions for different disciplines, allowing for fair comparison of adaptive transcendence indices across different disciplines. This invention cleverly transforms the probability mass function into a probability density function, and then utilizes the continuity of the probability density function to avoid the situation where different citation counts correspond to the same transcendence index, as is present in existing transcendence indices. The adaptive transcendence index proposed in this invention has high interpretability, meaning it represents the probability that a paper has a higher citation count than any other paper in the same year and field.
[0051] Example 2
[0052] See Figure 2 This embodiment provides a scientific paper impact evaluation system based on an adaptive transcendence index, including:
[0053] The paper topic determination module is configured to: use a large language model to analyze and obtain the topic of one or more papers based on the title, abstract and preset prompts of the specified papers;
[0054] The real-time retrieval module is configured to: based on the paper's topic, retrieve a set of papers belonging to the same topic as the given paper in an open-source paper database in real time, and obtain the most relevant papers within that set. The paper and each paper Corresponding citation count ;
[0055] The adaptive transcendental exponent calculation module is configured to: calculate based on the selected... The paper and each paper Corresponding citation count The discrete citation frequency distribution of k papers is calculated. This is treated as a probability mass function, i.e. ,in Indicates having The number of cited papers. Fitting the maximum likelihood estimation method. This yields the probability density function of the continuous exponential distribution. .in, This represents the probability density function at the value x. The results of the maximum likelihood estimation method fitting are given, where represent the scale parameter controlling scaling and the position parameter controlling displacement, respectively. Calculation In the interval The definite integral over the given surface yields the adaptive transcendence exponent. ;
[0056] The visualization output module is configured to render a PDF file with illustrations and a neat interface in real time according to the previously obtained and calculated paper-related data by using Python language. Figure 3 As shown in FIG. 8, it is a PDF page screenshot rendered by the scientific paper influence real-time evaluation system, and the PDF contains the title, author, publication date, aFNCSI index and visualization chart, etc.
[0057] The paper theme determination module and the real-time retrieval module designed in the embodiment can be used to match the existing papers in any database, the retrieval algorithm has low overhead and small server occupancy pressure, and the latest retrieval data can be returned in real time, thereby effectively ensuring the timeliness and accuracy of the original data. The adaptive transcendence index calculation module has low algorithm complexity, can quickly fit the frequency distribution curve of the number of citations of papers in different fields, and obtain the adaptive transcendence index. The visualization output module can visually display the abstract statistical data and numerical results in the PDF file, thereby effectively enhancing the actual experience of users.
[0058] Embodiment 3
[0059] Preferably, the embodiment also provides a specific implementation of an electronic device capable of implementing all steps of the scientific paper influence evaluation method based on the adaptive transcendence index in the above embodiments, and the electronic device specifically includes the following contents:
[0060] a processor, a memory, a communications interface and a bus;
[0061] The processor, the memory and the communications interface complete the communication among each other through the bus; the communications interface is used to realize the information transmission between the server-side device, the metering device and the user-side device and other related devices.
[0062] The processor is used to call the computer program in the memory, and the processor executes the computer program to realize all steps of the scientific paper influence evaluation method based on the adaptive transcendence index in the above embodiments.
[0063] The embodiment of the present application also provides a computer readable storage medium capable of implementing all steps of the scientific paper influence evaluation method based on the adaptive transcendence index in the above embodiments, and the computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize all steps of the scientific paper influence evaluation method based on the adaptive transcendence index in the above embodiments.
[0064] Those skilled in the art will appreciate that embodiments of the application can be readily used as software, hardware, or a combination of software and hardware. In a software embodiment, the methods can be tangibly embodied in a computer-readable storage medium having stored
[0065] The present application is described in reference to the flowchart and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0066] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0068] Those skilled in the art will appreciate that implementing all or part of the methods described above in the embodiments can be accomplished by way of computer program instructions arranged to process the relevant hardware, the program being stored on a computer-readable storage medium. The storage medium can be a magnetic disk, optical disk, Read-Only Memory (ROM) or Random Access Memory (RAM), etc.
[0069] The present application is not limited to the embodiments described above. The above description of the specific embodiments is intended to describe and illustrate the technical solutions of the present application, and the specific embodiments described above are merely illustrative and not restrictive. Without departing from the purpose of the present application and the scope protected by the claims, those of ordinary skill in the art can make many forms of specific changes under the inspiration of the present application, and these all belong to the protection scope of the present application.
Claims
1. A scientific paper impact evaluation method based on adaptive beyond exponential, characterized in that, The method comprises the following steps: determining a paper theme; obtaining the theme of the paper by using a large language model according to a title, an abstract and a preset prompt of the paper; obtaining real-time search results according to the paper theme; obtaining a set of papers of the same theme as the given paper by searching in real time in a publicly accessible paper database through the paper theme, and obtaining the most relevant papers in the set of papers papers and each paper corresponding number of citations ; According to the search results, an adaptive over-index is calculated; according to the selected papers and each paper corresponding number of citations , the discrete citation frequency distribution of k papers is calculated and regarded as a probability mass function, i.e. , wherein represents the number of papers with citations using maximum likelihood estimation method , to obtain a probability density function (PDF) of continuous exponential distribution ; wherein f(x) represents the probability density at value x, are the results of maximum likelihood estimation method, respectively representing the scale parameter of control scaling and the parameter of control displacement position; Computing the Definite Integral on an Interval Adaptive Exponential ; Represents the number of citations for a given paper.
2. A scientific paper impact evaluation system based on adaptive beyond index, characterized in that, The method comprises the following steps: A paper theme determination module is configured to obtain the theme of the paper by using a large language model according to a title, an abstract and a preset prompt of the paper. a real-time retrieval module for retrieving, in real time, a set of papers belonging to the same topic as the given paper from a publicly accessible database of papers according to the topic of the paper, and obtaining the most relevant papers in the set ; an adaptive over-indexing module for calculating, for each of the k papers the number of papers with the number of papers with the number of papers with where the number of papers with the number of papers with Fitting using maximum likelihood estimation , resulting in a probability density function of the continuous exponential distribution ; where, denotes the probability density at value x, are the results of the maximum likelihood estimation fitting, respectively the scale parameter controlling the scaling and the shift position parameter; calculate the definite integral over the interval , resulting in the adaptive hyper-exponential , denotes the number of citations of the specified paper; A visual output module is configured to render a PDF file in real time by using Python according to the previously obtained and calculated number of citations of the paper and the adaptive surpassing index.
3. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to realize the steps in the scientific paper influence evaluation method based on the adaptive surpassing index according to claim 1.
4. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps in the scientific paper influence evaluation method based on the adaptive surpassing index according to claim 1.
Citation Information
Patent Citations
Journal influence evaluation method based on academic big data
CN106484839A
Determination method for paper influence
CN116595177A