A paper screening method, system and related device

By evaluating the weight and combination of the target paper in the cited paper, the problem of low accuracy in the existing technology of screening representative papers by relying on the citation frequency is solved, and more accurate paper screening is achieved.

CN120030154BActive Publication Date: 2025-07-22INST OF MEDICAL INFORMATION CHINESE ACAD OF MEDICAL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510520049.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the prior art, when screening representative papers based on the frequency of citations, it is difficult to fully reflect the degree of innovation of the paper content, resulting in a low screening accuracy.

Method used

By calculating the weight and combination of the knowledge units in the target paper in the cited paper, the degree of popularity and novelty of the paper is comprehensively evaluated and representative papers are selected.

Benefits of technology

It improves the screening accuracy of representative papers and can more comprehensively reflect the degree and importance of the papers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030154B_ABST
    Figure CN120030154B_ABST
Patent Text Reader

Abstract

The present application discloses a method, a system and related devices for screening papers, relating to the field of data processing, including: obtaining knowledge units in a target paper, where the knowledge units are key contents in the target paper; calculating the weight of a knowledge unit for a citing paper based on the position of the knowledge unit in the citing paper, calculating the total weight of the knowledge units in the target paper and taking it as the weight of the target paper for the citing paper; obtaining combinations of knowledge units, calculating the prevalence of each combination in a paper collection and calculating the novelty of each combination; calculating the total novelty of the combinations in the target paper and taking it as the novelty of the target paper; taking the product of the weight of the target paper and the novelty of the target paper as the originality of the target paper, and selecting representative papers according to the originality. The present application calculates the originality based on the weight and novelty of the papers and screens the papers, which can improve the screening accuracy of representative papers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and particularly to a method and system for screening papers and related devices. Background Art

[0002] Due to some current requirements, scholars are often required to provide a certain number range of representative academic papers.

[0003] Currently, the main methods for screening papers are to directly or indirectly use the citation frequency of papers for screening. For example, calculating the H-index of scholars or comparing the actual citation frequency of papers with the citation frequency baseline and other indirect methods. Since the citation frequency of papers cannot fully reflect the innovation degree of papers, relying solely on a single indicator of the citation frequency of papers for screening has the problem of low accuracy in screening representative papers. Summary of the Invention

[0004] In view of the above problems, this application provides a method and system for screening papers and related devices to achieve the purpose of improving the accuracy of screening representative papers. The specific solutions are as follows:

[0005] In the first aspect of this application, a method for screening papers is provided. The method for screening papers includes:

[0006] Obtain at least two knowledge units in the target paper, where the target paper is a published paper of the user in the paper collection, the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge unit is the key content in the target paper;

[0007] For each knowledge unit in the target paper, calculate the weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper, where the citing paper is a paper that cites the target paper;

[0008] Calculate the total weight of the knowledge units in the target paper, and use the total weight as the weight of the target paper for the citing paper;

[0009] Obtain at least one combination pair of the knowledge units in the target paper, calculate the prevalence of each combination pair in the total combination pairs of the paper collection, and calculate the novelty of each combination pair based on the prevalence, where the total combination pairs are: in the paper collection, the sum of the non-repeated combination pairs of the knowledge units in each paper;

[0010] Calculate the total novelty of the combination pairs in the target paper, and use the total novelty as the novelty of the target paper;

[0011] Take the product of the weight of the target paper and the novelty level of the target paper as the originality level of the target paper;

[0012] After obtaining the originality levels of the published papers in the paper collection, select some of the published papers from the published papers as the representative papers of the user according to the originality levels.

[0013] In a possible implementation, calculating the weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper includes:

[0014] Calculate the position weights of the knowledge unit in each of the citing papers respectively, and standardize the respective position weights.

[0015] Take the sum of the standardized position weights of the knowledge unit as the weight of the knowledge unit for the citing paper.

[0016] In a possible implementation, the calculating the position weights of the knowledge unit in each of the citing papers respectively includes:

[0017] For each knowledge unit in each of the citing papers:

[0018] Count the first frequency of the knowledge unit appearing in the title of the citing paper and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper;

[0019] Obtain the first position weight of the title and the second position weight of the abstract paragraph;

[0020] Calculate the first product of the first frequency and the first position weight, and calculate the second product of the second frequency and the second position weight;

[0021] Take the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.

[0022] In a possible implementation, at least the first knowledge unit and the second knowledge unit are included in the combination pair, and calculating the prevalence of each combination pair in the total combination pairs of the paper collection includes:

[0023] Count the target combination pair total, where the target combination pair total is: the sum of the numbers of the combination pairs of the knowledge units in each paper in the paper collection;

[0024] Obtain the co-occurrence number of the combination pair in the target combination pair total;

[0025] Count the number of first combination pairs that only contain the first knowledge unit in the total of the combination pairs, and the number of second combination pairs that only contain the second knowledge unit in the total of the combination pairs;

[0026] Calculate the third product of the co-occurrence quantity and the total of the combination pairs, and the fourth product of the number of the first combination pairs and the number of the second combination pairs;

[0027] Take the ratio of the third product to the fourth product as the prevalence of the combination pair in the total of the combination pairs in the paper collection, where the third product is the numerator and the fourth product is the denominator.

[0028] In a possible implementation, calculating the novelty degree of each combination pair based on the prevalence includes:

[0029] Normalize the data of the prevalence of each combination pair;

[0030] Calculate the novelty degree of each combination pair based on the prevalence after data normalization.

[0031] The second aspect of this application provides a paper screening system, and the paper screening system includes:

[0032] A knowledge extraction unit, configured to obtain at least two knowledge units in a target paper, where the target paper is a published paper of a user in a paper collection, the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge unit is the key content in the target paper;

[0033] A first calculation unit, configured to, for each knowledge unit in the target paper, calculate the weight of the knowledge unit for a citing paper based on the position of the knowledge unit in the citing paper, where the citing paper is a paper that cites the target paper;

[0034] A weight calculation unit, configured to calculate the total weight of the knowledge units in the target paper, and take the total weight as the weight of the target paper for the citing paper;

[0035] A second calculation unit, configured to obtain at least one combination pair of the knowledge units in the target paper, calculate the prevalence of each combination pair in the total of the combination pairs in the paper collection, and calculate the novelty degree of each combination pair based on the prevalence, where the total of the combination pairs is: the sum of non-repeating combination pairs of the knowledge units in each paper in the paper collection;

[0036] A combination calculation unit, configured to calculate the total novelty degree of the combination pairs in the target paper, and take the total novelty degree as the novelty degree of the target paper;

[0037] A paper evaluation unit, configured to use the product of the weight of the target paper and the novelty degree of the target paper as the originality degree of the target paper;

[0038] A paper screening unit, configured to, after obtaining the originality degrees of the published papers in the paper collection, select some of the published papers as the representative papers of the user according to the originality degrees.

[0039] In a possible implementation, the first calculation unit may include:

[0040] A position weight calculation sub-unit, configured to calculate the position weights of the knowledge unit in each of the citing papers respectively, and perform data standardization on each of the position weights;

[0041] A weight sum calculation sub-unit, configured to use the sum of the standardized position weights of the knowledge unit as the weight of the knowledge unit for the citing paper.

[0042] In a possible implementation, the position weight calculation sub-unit is specifically configured as:

[0043] For each knowledge unit in each of the citing papers: count the first frequency of the knowledge unit appearing in the title of the citing paper and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper; obtain the first position weight of the title and the second position weight of the abstract paragraph; calculate the first product of the first frequency and the first position weight, calculate the second product of the second frequency and the second position weight; use the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.

[0044] A third aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:

[0045] The memory is used to store a computer program;

[0046] The processor is used to execute the computer program so that the electronic device can implement the paper screening method according to the first aspect or any implementation manner of the first aspect.

[0047] A fourth aspect of the present application provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the paper screening method according to the first aspect or any implementation manner of the first aspect.

[0048] With the above technical solutions, the present application provides a method, a system and related devices for screening papers. This method evaluates the creativity of papers from two perspectives, namely the weight of the target paper for the citing papers and the novelty of the target paper among the papers published by the user and all the reference papers. First, knowledge units that are regarded as key content are selected from the target paper. Regarding the weight of the target paper for the citing papers, based on the positions of the knowledge units in the citing papers, this method calculates the weight of each knowledge unit for the citing papers, and takes the sum of the weights of the knowledge units in the target paper as the weight of the target paper for the citing papers. Regarding the novelty of the target paper among the papers published by the user and the reference papers, this method obtains at least one combination pair of knowledge units in the target paper, calculates the prevalence of each combination pair in the total combination pairs of the paper collection respectively, then calculates the novelty of each combination pair in the total combination pairs of the paper collection, and takes the sum of the novelties of each combination pair in the target paper as the novelty of the target paper among the papers published by the user and the reference papers. Finally, the product of the weight of the target paper and the novelty of the target paper is taken as the originality of the target paper, and representative papers of the user are selected from the published papers in the paper collection according to the originality. This method does not screen papers based on the citation frequency of the papers, but comprehensively screens papers from multiple aspects such as the weight of the paper for the citing papers and the novelty of the paper among the papers published by the user and all the reference papers. The weight of the paper for the citing papers represents the importance of the paper for the papers published after it, while the novelty of the paper represents the novelty of the paper for the papers published by the user and the reference papers published before it. Therefore, the papers selected by this method based on the originality calculated from the weight and novelty of the papers can effectively improve the screening accuracy of the representative papers. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0050] Figure 1 It is a schematic flowchart of a method for screening papers provided by an embodiment of the present application;

[0051] Figure 2 It is a schematic structural diagram of a system for screening papers provided by an embodiment of the present application;

[0052] Figure 3 It is a hardware structure block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application.

[0054] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will understand that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.

[0055] The terms "first", "second", etc. in the specification of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0056] In many scenarios, how scholars select representative papers from all the papers they have published is a problem that many scholars need to face. Representative papers can reflect the academic achievements of scholars in a certain field.

[0057] Currently, the main methods for screening papers are to directly or indirectly screen papers using the citation frequency of the papers. However, the citation frequency of papers is used to screen papers from the dimension of the influence of the papers, and it is difficult to comprehensively reflect the originality or innovation degree of the content of the papers. For example, a paper with very novel content may not attract attention due to publication channels or language limitations, and thus has a low citation frequency. Another example is some review papers, whose content is generally a summary and sorting of the existing achievements of other scholars and does not produce novel content, but usually can obtain a high citation frequency.

[0058] Therefore, as the product of research activities, academic papers are the carriers of knowledge content. Screening representative papers only relying on the citation frequency of papers only considers the citation situation of papers and does not consider the innovation degree of the content of the papers themselves, and the screening accuracy is not high.

[0059] To solve the above problems, an embodiment of the present application provides a method for screening papers. This method turns to evaluate and screen papers from the content level of the papers, and comprehensively evaluates the originality of the paper content by the content novelty of the paper in the papers published by the user and the papers published before it, and the importance of the paper in the papers published after it, so as to screen out representative papers. The method for screening papers in the embodiment of the present application will be introduced in detail below with reference to the accompanying drawings.

[0060] Referring to Figure 1 , Figure 1 FIG. is a schematic flowchart of a method for screening papers provided by an embodiment of the present application. As Figure 1 shown, a data processing method provided by an embodiment of the present application may include steps S10 to S16, and these steps will be described in detail below.

[0061] S10. Obtain at least two knowledge units in the target paper. The target paper is a published paper of the user in the paper collection. The paper collection includes at least one published paper of the user and at least one reference paper of the published paper. The knowledge unit is the key content in the target paper.

[0062] Among them, the knowledge unit may be a concept, statement, word or phrase, term, law, etc. used in this solution to represent specific content. The key content in the target paper may refer to the content that has a high correlation with the research content or research direction of the target paper. Therefore, the knowledge unit may be a keyword, subject word or phrase in the target paper. In another alternative embodiment, the knowledge unit may also be a key sentence, paragraph, etc. in the target paper.

[0063] The knowledge unit can be directly selected fixedly. For example, directly select the keywords located after the abstract and before the text in the target paper as the knowledge unit in this embodiment, or extract the keywords, subject words or phrases in the target paper through a third-party extraction tool and use them as the knowledge unit. When the knowledge unit is a phrase, there may be a semantic relationship between the multiple words that make up the phrase, such as the subject-object relationship.

[0064] When extracting knowledge units, the third-party extraction tool can not only directly extract single words as knowledge units, but also extract a complete semantic phrase and select some words from the semantic phrase to form a phrase as a knowledge unit to more accurately reflect the knowledge content in the paper. For example, if the extraction tool can separately extract two words, A and B, from the target paper, then A can be used as a knowledge unit and B can be used as a knowledge unit. The extraction tool can also extract the phrase "A treats B" with semantic relationships from the target paper, where A is the subject, "treats" is the predicate, and B is the object. Then, A and B can be selected as words, and the phrase formed by A - B can be used as a knowledge unit. At this time, there is a subject-object relationship between A and B.

[0065] The paper collection can be a collection that includes multiple published papers of the user and multiple reference papers of the published papers. Among them, some reference papers of the published papers can be selected, or all reference papers of the published papers can be selected. In this paper collection, multiple published papers of the user can belong to the same field or research direction, and can also belong to the same publication time period. For example, all papers published by the user within a certain five-year period or all papers published within one year.

[0066] S11. For each knowledge unit in the target paper, calculate the weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper, where the citing paper is the paper that cites the target paper;

[0067] S12. Calculate the total weight of the knowledge units in the target paper, and use the total weight as the weight of the target paper for the citing paper.

[0068] Among them, the weight of the knowledge unit for the citing paper can represent: the usefulness or importance of the knowledge unit of the target paper in all citing papers. And the weight of the target paper for the citing paper can represent: the influence of the target paper on the citing paper or the usefulness or importance of the content of the target paper for the citing paper, and can also represent the usefulness or importance of the target paper in a field.

[0069] In this embodiment, the citing papers of the target paper can be counted first, and then the weight of each knowledge unit in the target paper for the citing paper can be calculated. Finally, the weights of each knowledge unit in the target paper for the citing paper are summed to obtain the weight of the target paper for the citing paper.

[0070] Specifically, the specific calculation process of the weight of the knowledge unit for the citing paper can be as shown in Step 1 and Step 2:

[0071] Step 1: Calculate the position weights of the knowledge unit in each citing paper respectively, and standardize the data of each position weight;

[0072] Step 2: Use the sum of the standardized position weights of each knowledge unit as the weight of the knowledge unit for the citing paper.

[0073] Among them, the position weight of a knowledge unit in a citing paper can be: the degree of usefulness of the knowledge unit for the citing paper calculated based on the appearance position of the knowledge unit in a citing paper of the target paper. For example, if a knowledge unit appears in the title of a citing paper, it indicates that the knowledge unit is highly relevant to the content of the citing paper, and the importance of the knowledge unit for the citing paper is relatively high. In this embodiment, for each knowledge unit in the target paper, this embodiment can first calculate the position weight of the knowledge unit for each citing paper according to the position of the knowledge unit in the target paper in each citing paper, and finally use the sum of all the position weights of the knowledge unit as the weight of the knowledge unit for the citing paper.

[0074] Specifically, the specific calculation process of the position weight of a knowledge unit in a citing paper can be as shown in Steps 3 to 6:

[0075] For each knowledge unit in each citing paper:

[0076] Step 3: Count the first frequency of the knowledge unit appearing in the title of the citing paper and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper;

[0077] Step 4: Obtain the first position weight of the title and the second position weight of the abstract paragraph;

[0078] Step 5: Calculate the first product of the first frequency and the first position weight, and calculate the second product of the second frequency and the second position weight;

[0079] Step 6: Use the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.

[0080] Among them, the formula form of the above content can be expressed as:

[0081]

[0082] Among them, can represent the knowledge unit; can represent the knowledge unit For the position weight of the citing paper, can represent the position weight. In this embodiment can include the first position weight and the second position weight , can represent the total number of position weight types, can represent a summation variable. In this embodiment, can take a value of 2; can represent a knowledge unit and the frequency of occurrence in the citing paper. In this embodiment can include a first frequency and a second frequency .

[0083] The first frequency can be the number of times the knowledge unit appears in the title of the citing paper, and the second frequency can be the number of times the knowledge unit appears in the abstract paragraph of the citing paper. The first position weight is the preset weight of the title, and the second position weight is the preset weight of the abstract paragraph. In this embodiment, the value of the first position weight can be 2, and the value of the second position weight can be 1.

[0084] Of course, in another alternative embodiment, when calculating the position weight of a knowledge unit in the citing paper, the statistical positions of the knowledge unit in the citing paper can include not only the title and the abstract, but also other positions in the citing paper, such as the summary paragraphs of each chapter in the citing paper.

[0085] Since for a knowledge unit, its importance can vary depending on its appearance position in the citing paper, and its importance can also vary depending on the number of times it appears in different positions of the citing paper. Therefore, this embodiment not only considers the appearance position of the knowledge unit in the citing paper, but also considers the number of times the knowledge unit appears in different positions of the citing paper, and performs a weighted sum calculation on the number of times and the preset weights corresponding to different appearance positions to obtain the position weight of the knowledge unit for the citing paper. For example, for knowledge unit A, it appears 1 time in the title of a citing paper and the preset weight corresponding to the title is 2, and it appears 3 times in the abstract paragraph of the citing paper and the preset weight corresponding to the abstract paragraph is 1, then the position weight of knowledge unit A for this citing paper is 5 (5 = 1×2 + 3×1).

[0086] In another alternative embodiment, the total weights of the citing paper in different positions can be directly set. When a knowledge unit appears in different positions of the citing paper, the total weights corresponding to each position can be summed up, and then the average value is calculated and used as the position weight of the knowledge unit in the citing paper.

[0087] Furthermore, when there is only one citing paper for the target paper, the position weight of the knowledge unit in the citing paper can be directly used as the weight of the knowledge unit for the citing paper. When there are multiple citing papers for the target paper, the sum of the position weights of the knowledge unit in each citing paper can be used as the weight of the knowledge unit for the citing paper, and its formula can be as follows:

[0088]

[0089] Among them, can represent the knowledge unit for the weight of the citing paper; can represent the total number of citing papers; can represent the knowledge unit the positional weight in the citing paper after data standardization.

[0090] The formula for data standardization can be shown as follows:

[0091]

[0092] Among them, can represent the knowledge unit the positional weight in the citing paper after data standardization; can represent the knowledge unit the positional weight in the citing paper; can represent the knowledge unit the maximum value among the positional weights of each position.

[0093] Since the lengths of citing papers are different, the lengths of the bibliographic entries (title + abstract) of citing papers can correspondingly vary. In a citing paper with a longer abstract, the knowledge unit can appear repeatedly in the abstract, resulting in a larger positional weight of the knowledge unit in the citing paper. While in a citing paper with a shorter abstract, the knowledge unit can appear only once in the abstract, resulting in a smaller positional weight of the knowledge unit in the citing paper. Therefore, the length of the citing paper can affect the weight of the knowledge unit for the citing paper. Thus, in this embodiment, to reduce the influence brought by some outliers, data standardization processing needs to be performed before calculating the weight of the knowledge unit for the citing paper.

[0094] After this embodiment obtains the weight of each knowledge unit in the target paper for the citing paper, the weight of the target paper for the citing paper can be obtained by calculating the sum of the weights of all knowledge units in the target paper. Its calculation formula can be shown as follows:

[0095]

[0096] Among them, can represent the target paper for the weight of the citing paper; can represent the target paper the total number of knowledge units; can represent the knowledge unit for the weight of the citing paper.

[0097] In this embodiment, by calculating the weight of the target paper with respect to the citing papers, since the citing papers are the papers that cite the target paper, the weight of the target paper with respect to the citing papers can reflect the importance of the content in the target paper for the citing papers published after the target paper. It is precisely because the content in the target paper has reference significance or novel content that other papers do not have that the citing papers choose to refer to the target paper. Therefore, the weight of the target paper calculated in this embodiment can well reflect the innovation of the content of the target paper.

[0098] S13. Obtain at least one combination pair of knowledge units in the target paper, calculate the prevalence of each combination pair in the total combination pairs of the paper collection, and calculate the novelty of each combination pair based on the prevalence. The total combination pairs are: the sum of the non-repeated combination pairs of the knowledge units in each paper in the paper collection;

[0099] S14. Calculate the total novelty of the combination pairs in the target paper, and use the total novelty as the novelty of the target paper.

[0100] Among them, there may not be a strict order of execution between step S11 - step S12 and step S13 - step S14. In an alternative embodiment, step S11 - step S12 can be executed first to calculate the weight of the target paper, and then step S13 - step S14 can be executed to calculate the novelty of the target paper. In another alternative embodiment, step S13 - step S14 can be executed first to calculate the novelty of the target paper, and then step S11 - step S12 can be executed to calculate the weight of the target paper.

[0101] The combination pair can be: a phrase containing at least two knowledge units. In the target paper, if the knowledge units are only keywords or subject words, multiple combination pairs of the knowledge units in the target paper can be obtained by pairwise combination; if the knowledge units are only phrases, each phrase can be directly used as a combination pair, thereby obtaining multiple combination pairs of the knowledge units in the target paper; if the knowledge units include both keywords or subject words and phrases, some combination pairs can be obtained by pairwise combination of the keywords or subject words first, and then each phrase can be used as a combination pair to obtain another part of the combination pairs. Finally, the part of the combination pairs and the other part of the combination pairs are used together as multiple combination pairs of the knowledge units in the target paper. Of course, in this embodiment, the words in the phrase can also be extracted first, and then pairwise combined with the keywords or subject words to finally obtain multiple combination pairs of the knowledge units in the target paper.

[0102] Since there are multiple papers in a paper collection, and each paper may have multiple pairs of knowledge units, and there may be duplicate pairs of combinations. In this embodiment, the total number of pairs of combinations can be: the sum of non-duplicate pairs of combinations in each paper of the paper collection. For example, the paper collection includes Paper A and Paper B. The pairs of combinations of knowledge units in Paper A include: a-b, a-c, and b-c. The pairs of combinations of knowledge units in Paper B include: a-b, a-d, and b-d. Since the pair of combinations a-b appears repeatedly, the total number of pairs of combinations in the paper collection is 5, not 6.

[0103] The prevalence of a pair of combinations can be: the extent to which the pair of combinations is widespread in the paper collection or the extent to which the pair of combinations is adopted or used in the paper collection. Then the novelty of a pair of combinations can be: the uniqueness of the pair of combinations in the paper collection or the extent to which the pair of combinations does not appear or is not used in the paper collection.

[0104] Since directly counting the number of occurrences of the knowledge units of the target paper in the paper collection is too straightforward and it is difficult to well measure the novelty of the target paper. Therefore, in this embodiment, the knowledge units of the target paper are transformed into the form of pairs of combinations. First, calculate the prevalence of the pairs of combinations in the paper collection, then calculate the novelty of the pairs of combinations, and finally use the sum of the novelties of each pair of combinations in the target paper as the novelty of the target paper.

[0105] Specifically, each pair of combinations includes at least a first knowledge unit and a second knowledge unit. The specific calculation process of the prevalence of the pairs of combinations in the target paper can be as shown in Steps Seven to Eleven:

[0106] Step Seven: Statistically calculate the total number of target pairs of combinations. The total number of target pairs of combinations is: the sum of the numbers of pairs of combinations of knowledge units in each paper in the paper collection;

[0107] Step Eight: Obtain the co-occurrence number of the pair of combinations in the total number of target pairs of combinations;

[0108] Step Nine: Statistically calculate the number of the first pairs of combinations that only include the first knowledge unit in the total number of pairs of combinations, and the number of the second pairs of combinations that only include the second knowledge unit in the total number of pairs of combinations;

[0109] Step Ten: Calculate the third product of the co-occurrence number and the total number of pairs of combinations, and the fourth product of the number of the first pairs of combinations and the number of the second pairs of combinations;

[0110] Step Eleven: Use the ratio of the third product to the fourth product as the prevalence of the pair of combinations in the total number of pairs of combinations in the paper collection, where the third product is the numerator and the fourth product is the denominator.

[0111] Among them, the total number of target combination pairs can be: the sum of the numbers of each combination pair of knowledge units in each paper in the paper collection. Different from the total number of combination pairs, the total number of target combination pairs includes duplicate combination pairs. For example, the paper collection includes Paper A and Paper B. The combination pairs of knowledge units in Paper A include: a-b, a-c, and b-c. The combination pairs of knowledge units in Paper B include: a-b, a-d, and b-d. Even if the combination pair a-b occurs repeatedly, it is still counted in the quantity. Therefore, the total number of target combination pairs in the paper collection is 6. If it is the total number of combination pairs, the duplicate combination pairs are not counted, and the total number of combination pairs is 5.

[0112] The co-occurrence quantity of a combination pair in the total number of target combination pairs is: the quantity of the combination pair in the total number of target combination pairs. Since the co-occurrence quantity of a combination pair represents the number of times the two knowledge units that make up the combination pair appear in a paper at the same time, therefore, the duplicate combination pairs need to be counted in the total number of target combination pairs in this embodiment, otherwise it may affect the calculation of the prevalence of the combination pair. For example, the paper collection includes Paper A and Paper B. The combination pairs of knowledge units in Paper A include: a-b, a-c, and b-c. The combination pairs of knowledge units in Paper B include: a-b, a-d, and b-d. If the duplicate combination pairs are not counted, the co-occurrence times of the combination pair a-b are only 1, but the combination pair a-b appears in both Paper A and Paper B, and the actual co-occurrence times of the combination pair a-b should be 2. If calculated with the co-occurrence times of 1, the prevalence of the combination pair a-b is relatively low, but the actual prevalence of the combination pair a-b should be relatively high.

[0113] The first combination pair can be: a combination pair in which one knowledge unit in the combination pair is the first knowledge unit and does not include the second knowledge unit. Then the quantity of the first combination pair can be the quantity of the first combination pair in the total number of combination pairs. The second combination pair can be: a combination pair in which one knowledge unit in the combination pair is the second knowledge unit and does not include the first knowledge unit. Then the quantity of the second combination pair can be the quantity of the second combination pair in the total number of combination pairs.

[0114] Furthermore, the formula form of the above content can be as follows:

[0115]

[0116] Among them, can represent the prevalence of the combination pair composed of knowledge unit and knowledge unit . The greater the prevalence, the more prevalent the combination pair composed of knowledge unit and knowledge unit in the paper collection, and the less novel it is; can represent the prevalence of the combination pair composed of knowledge unit and knowledge unit The co-occurrence quantity of the formed combination pair in the total of target combination pairs; Can represent the number of the first combination pairs that only contain knowledge units in the total of combination pairs ; Can represent the number of the second combination pairs that only contain knowledge units in the total of combination pairs ; Can represent the total of combination pairs of the thesis collection.

[0117] After calculating the prevalence of all combination pairs in the target thesis in this embodiment, this embodiment needs to first normalize the prevalence of each combination pair in the target thesis (linear function normalization), and then calculate the novelty degree of each combination pair in the target thesis based on the prevalence after data normalization. Specifically, the formula for data normalization can be shown as follows:

[0118]

[0119] Wherein, Can represent the prevalence of the combination pair formed by knowledge unit and knowledge unit after data normalization; Can be the minimum value of the prevalence among all combination pairs in the target thesis; Can be the maximum value of the prevalence among all combination pairs in the target thesis.

[0120] After this embodiment adjusts the prevalence of all combination pairs in the target thesis to the range of [0, 1] through data normalization, in order to calculate the novelty degree of the combination pair based on the prevalence after data normalization of the combination pair, the calculation formula can be shown as follows:

[0121]

[0122] Wherein, Can represent the novelty degree of the combination pair formed by knowledge unit and knowledge unit ; Can represent the prevalence of the combination pair formed by knowledge unit and knowledge unit after data normalization.

[0123] Thus, this embodiment can obtain the novelty degree of all combination pairs in the target thesis, then calculate the sum of the novelty degrees of all combination pairs in the target thesis, and take the sum of the novelty degrees as the novelty degree of the target thesis. Its formula form can be shown as follows:

[0124]

[0125] Among them, can represent the novelty level of the target paper ; can represent the total number of knowledge units of the target paper . If the knowledge units are keywords or subject terms, there are combinations correspondingly. In another alternative embodiment, if the knowledge units are phrases, there are combinations correspondingly; can represent the novelty level of the combination pair composed of knowledge unit and knowledge unit .

[0126] Since in this embodiment, the paper collection is the collection of the user's published papers and all reference papers of the published papers, and the target paper is one of the published papers, the reference papers are the papers published before the target paper. Therefore, in this embodiment, by calculating the novelty level of the target paper in the paper collection, it is indicated that the target paper has novelty compared with the previously published papers, so as to reflect the innovation of the content of the target paper.

[0127] S15. Take the product of the weight of the target paper and the novelty level of the target paper as the originality level of the target paper;

[0128] S16. After obtaining the originality levels of the published papers in the paper collection, select some published papers from the published papers as the representative papers of the user according to the originality levels.

[0129] Among them, the originality level of the target paper can represent: the uniqueness of the content of the target paper, the degree of difference from the existing achievements. In this embodiment, a creativity measurement model of the target paper can be designed, and the creativity measurement model is used to calculate the originality level of the target paper. Specifically, in this embodiment, the paper collection and the citing papers can be directly input into the creativity measurement model, and the creativity measurement model calculates the weight of the target paper for the citing papers, the novelty level of the target paper, and the originality level of the target paper. Finally, the originality level of the target paper output by the creativity measurement model is obtained. In this embodiment, the originality level of the target paper can also be calculated by inputting the calculated weight of the target paper for the citing papers and the novelty level of the target paper into the creativity measurement model. Among them, the formula representation form of the creativity measurement model can be as follows:

[0130]

[0131] Among them, can represent the originality level of the target paper ; can represent the novelty level of the target paper ; Can represent the target paper Regarding the weight of the citing papers.

[0132] After obtaining the originality of all the published papers in the paper collection in this embodiment, all the published papers can be sorted according to the originality, and some of the published papers with high originality can be selected as the representative papers of the user.

[0133] A paper screening method provided by the present application evaluates the creativity of papers from two perspectives, namely the weight of the target paper for the citing papers and the novelty of the target paper among the user's published papers and all reference papers. First, select the knowledge units that are the key content from the target paper. Regarding the weight of the target paper for the citing papers, this method calculates the weight of each knowledge unit for the citing papers based on the position of the knowledge unit in the citing papers, and takes the sum of the weights of the knowledge units in the target paper as the weight of the target paper for the citing papers. Regarding the novelty of the target paper among the user's published papers and reference papers, this method obtains at least one combination pair of knowledge units in the target paper, calculates the prevalence of each combination pair in the total combination pairs of the paper collection respectively, then calculates the novelty of each combination pair in the total combination pairs of the paper collection, and takes the sum of the novelties of each combination pair in the target paper as the novelty of the target paper among the user's published papers and reference papers. Finally, take the product of the weight of the target paper and the novelty of the target paper as the originality of the target paper, and screen out the representative papers of the user from the published papers in the paper collection according to the originality. This method does not screen papers based on the citation frequency of the papers, but comprehensively screens papers from multiple aspects such as the weight of the paper for the citing papers and the novelty of the paper among the user's published papers and all reference papers. The weight of the paper for the citing papers represents the importance of the paper for the papers published after it, while the novelty of the paper represents the novelty of the paper for the user's published papers and the reference papers published before it. Therefore, this method uses the papers screened based on the originality calculated from the weight and novelty of the papers as the representative papers, which can effectively improve the screening accuracy of the representative papers.

[0134] The above introduced a paper screening method provided by the embodiments of the present application. Next, a system applying the above paper screening method will be introduced.

[0135] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a paper screening system provided by the embodiments of the present application. As Figure 2 shown, the paper screening system includes:

[0136] The knowledge extraction unit 100 is used to obtain at least two knowledge units from the target paper, where the target paper is a published paper of the user in the paper collection. The paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge unit is the key content in the target paper;

[0137] The first calculation unit 110 is used to calculate the weight of each knowledge unit in the target paper for the citing paper based on the position of the knowledge unit in the citing paper, where the citing paper is a paper that cites the target paper;

[0138] The weight calculation unit 120 is used to calculate the total weight of the knowledge units in the target paper and use the total weight as the weight of the target paper for the citing paper;

[0139] The second calculation unit 130 is used to obtain at least one combination pair of the knowledge units in the target paper, calculate the prevalence of each combination pair in the total combination pairs of the paper collection, and calculate the novelty of each combination pair based on the prevalence. The total combination pairs are: the sum of the non-repeated combination pairs of the knowledge units in each paper in the paper collection;

[0140] The combination calculation unit 140 is used to calculate the total novelty of the combination pairs in the target paper and use the total novelty as the novelty of the target paper;

[0141] The paper evaluation unit 150 is used to use the product of the weight of the target paper and the novelty of the target paper as the originality of the target paper;

[0142] The paper screening unit 160 is used to select some of the published papers as the representative papers of the user from the published papers according to the originality after obtaining the originality of the published papers in the paper collection.

[0143] In a possible implementation, the first calculation unit 110 may include:

[0144] The position weight calculation sub-unit is used to calculate the position weights of the knowledge units in each citing paper respectively and standardize the data of each position weight;

[0145] The total weight calculation sub-unit is used to use the sum of the standardized position weights of the knowledge units as the weight of the knowledge unit for the citing paper.

[0146] In a possible implementation, the position weight calculation sub-unit may be specifically configured as:

[0147] For each knowledge unit in each citing paper: count the first frequency of the knowledge unit appearing in the title of the citing paper, and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper; obtain the first position weight of the title and the second position weight of the abstract paragraph; calculate the first product of the first frequency and the first position weight, and calculate the second product of the second frequency and the second position weight; take the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.

[0148] In a possible implementation, the combination pair at least includes a first knowledge unit and a second knowledge unit. The second calculation unit 130 calculates the prevalence of each combination pair in the total combination pairs of the paper collection, which can be specifically configured as:

[0149] Count the total target combination pairs, where the total target combination pairs are: in the paper collection, the sum of the number of combination pairs of knowledge units in each paper; obtain the co-occurrence number of the combination pair in the total target combination pairs; count the number of first combination pairs that only contain the first knowledge unit in the total combination pairs, and the number of second combination pairs that only contain the second knowledge unit in the total combination pairs; calculate the third product of the co-occurrence number and the total combination pairs, and the fourth product of the number of first combination pairs and the number of second combination pairs; take the ratio of the third product and the fourth product as the prevalence of the combination pair in the total combination pairs of the paper collection, where the third product is the numerator and the fourth product is the denominator.

[0150] In a possible implementation, the second calculation unit 130 calculates the novelty of each combination pair based on the prevalence, which can be specifically configured as:

[0151] Normalize the data of the prevalence of each combination pair; calculate the novelty of each combination pair based on the prevalence after data normalization.

[0152] This application embodiment also provides an electronic device. Refer to Figure 3 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in this application embodiment. The electronic device in this application embodiment may include but is not limited to fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 3 The shown electronic device is only an example and should not bring any limitations to the functions and usage scopes of this application embodiment.

[0153] As Figure 3As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage device 308 into the random access memory (RAM) 303. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.

[0154] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a memory card, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0155] In an embodiment of the present application, there is also provided a computer program product including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any one of the paper screening methods provided in the embodiments of the present application.

[0156] In an embodiment of the present application, there is also provided a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device is enabled to implement any one of the paper screening methods provided in the embodiments of the present application.

[0157] In addition, it should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the system embodiment drawings provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which may be specifically implemented as one or more communication buses or signal lines.

[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can easily be implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, software program implementation is a better embodiment in more cases. Based on such understanding, the technical solution of this application, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0159] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0160] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)), etc.

[0161] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments.

[0162] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the users and the authorization of the users should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0163] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for screening papers, characterized in that, The described paper screening method includes: Obtaining at least two knowledge units in a target paper, where the target paper is a published paper of a user in a paper collection, the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge unit is the key content in the target paper; For each knowledge unit in the target paper, calculate the weight of the knowledge unit for a citing paper based on the position of the knowledge unit in the citing paper, where the citing paper is a paper that cites the target paper; Calculate the total weight of the knowledge units in the target paper, and use the total weight as the weight of the target paper for the citing paper; Obtain at least one combination pair of knowledge units in the target paper, calculate the prevalence of each combination pair in the total combination pairs of the paper collection, and calculate the novelty of each combination pair based on the prevalence, where the total combination pairs is: the sum of the non-repeated combination pairs of knowledge units in each paper in the paper collection; Calculate the total novelty of the combination pairs in the target paper, and use the total novelty as the novelty of the target paper; Take the product of the weight of the target paper and the novelty of the target paper as the originality of the target paper; After obtaining the originality of the published papers in the paper collection, select some published papers from the published papers as the representative papers of the user according to the originality; Wherein, at least a first knowledge unit and a second knowledge unit are included in the combination pair, and calculating the prevalence of each combination pair in the total combination pairs of the paper collection and calculating the novelty of each combination pair based on the prevalence includes: Statistical target combination pair total, where the target combination pair total is: the sum of the number of combination pairs of knowledge units in each paper in the paper collection; Obtain the co-occurrence number of the combination pair in the target combination pair total; Statistical the number of first combination pairs that only contain the first knowledge unit in the total combination pairs, and the number of second combination pairs that only contain the second knowledge unit in the total combination pairs; Calculate the third product of the co-occurrence number and the total combination pairs, and the fourth product of the number of the first combination pairs and the number of the second combination pairs; Take the ratio of the third product and the fourth product as the prevalence of the combination pair in the total combination pairs of the paper collection, where the third product is the numerator and the fourth product is the denominator; Normalize the prevalence of each combination pair, and convert the prevalence of each combination pair into a first value, where the first value falls within the target value range; Calculate a second value for each combination pair based on the first value of each combination pair, and use the second value of each combination pair as the novelty of each combination pair, where the sum of the values of the first value and the second value is the length of the target value range.

2. The paper screening method according to claim 1, characterized in that Calculating the weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper includes: Calculating the position weights of the knowledge unit in each of the citing papers respectively, and normalizing the respective position weights; Taking the sum of the normalized position weights of the knowledge unit as the weight of the knowledge unit for the citing paper.

3. The paper screening method according to claim 2, characterized in that, The step of respectively calculating the position weights of the knowledge unit in each of the citing papers includes: For each knowledge unit in each of the citing papers: Counting the first frequency of the knowledge unit appearing in the title of the citing paper and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper; Obtaining the first position weight of the title and the second position weight of the abstract paragraph; Calculating the first product of the first frequency and the first position weight, and calculating the second product of the second frequency and the second position weight; Taking the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.

4. A paper screening system, characterized in that, The paper screening system includes: A knowledge extraction unit configured to obtain at least two knowledge units from a target paper, where the target paper is a published paper of a user in a paper collection, the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge unit is the key content in the target paper; A first calculation unit configured to calculate, for each knowledge unit in the target paper, the weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper, where the citing paper is a paper that cites the target paper; A weight calculation unit configured to calculate the total weight of the knowledge units in the target paper, and taking the total weight as the weight of the target paper for the citing paper; A second calculation unit configured to obtain at least one combination pair of the knowledge units in the target paper, calculate the prevalence of each combination pair in the total combination pairs of the paper collection, and calculate the novelty of each combination pair based on the prevalence, where the total combination pairs is: the sum of the non-repeated combination pairs of the knowledge units in each paper in the paper collection; A combination calculation unit configured to calculate the total novelty of the combination pairs in the target paper, and taking the total novelty as the novelty of the target paper; A paper evaluation unit configured to take the product of the weight of the target paper and the novelty of the target paper as the originality of the target paper; A paper screening unit configured to, after obtaining the originality of the published papers in the paper collection, select some published papers from the published papers as the representative papers of the user according to the originality; Wherein, the combination pair includes at least a first knowledge unit and a second knowledge unit, and the second calculation unit is specifically configured as: Statistically calculate the total number of target combination pairs. The total number of target combination pairs is: in the paper collection, the total number of combination pairs of knowledge units in each paper; obtain the co-occurrence number of the combination pairs in the total number of target combination pairs; statistically calculate the number of first combination pairs that only contain the first knowledge unit in the total number of combination pairs, and the number of second combination pairs that only contain the second knowledge unit in the total number of combination pairs; calculate the third product of the co-occurrence number and the total number of combination pairs, and the fourth product of the number of first combination pairs and the number of second combination pairs; use the ratio of the third product to the fourth product as the prevalence of the combination pairs in the total number of combination pairs in the paper collection, where the third product is the numerator and the fourth product is the denominator. Normalize the data of the prevalence of each of the combination pairs, and convert the prevalence of each of the combination pairs into a first value, where the first value falls within the target value range; calculate a second value for each of the combination pairs based on the first value of each of the combination pairs, and use the second value of each of the combination pairs as the novelty of each of the combination pairs, where the sum of the values of the first value and the second value is the length of the target value range.

5. The paper screening system according to claim 4, wherein The first calculation unit includes: A position weight calculation sub-unit, configured to calculate the position weights of the knowledge units in each of the citing papers respectively, and normalize the data of each of the position weights. A weight sum calculation sub-unit, configured to use the sum of the normalized position weights of the knowledge units as the weight of the knowledge units for the citing papers.

6. The paper screening system according to claim 5, wherein The position weight calculation sub-unit is specifically configured as: For each knowledge unit in each of the citing papers: statistically calculate the first frequency of the knowledge unit appearing in the title of the citing paper, and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper; obtain the first position weight of the title and the second position weight of the abstract paragraph. Calculate the first product of the first frequency and the first position weight, and calculate the second product of the second frequency and the second position weight. Use the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.

7. An electronic device, characterized in that, Comprising at least one processor and a memory connected to the processor, wherein: The memory is used to store a computer program. The processor is used to execute the computer program so that the electronic device can implement the paper screening method described in any one of claims 1 to 3.

8. A computer program product, characterized in that, Comprising computer-readable instructions, when the computer-readable instructions run on an electronic device, the electronic device is enabled to implement the paper screening method described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Academic paper-based technological frontier index calculation method and system

    CN108614867A

  • Paper recommendation method based on knowledge graph and graph neural network

    CN118364139A