Paper screening method, system and related device
By calculating the position weights and combination pairs of key knowledge units in the cited papers, evaluating the originality of the papers, solving the problem of low accuracy of citation frequency screening in the existing technology, and achieving more accurate representative paper screening.
Patent Information
- Application Number
- CN202510520049.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In the prior art, the screening rate of papers cannot be fully reflected in the degree of innovation of papers, resulting in low accuracy in screening of representative papers.
By obtaining the key knowledge units in the target paper, calculating their position weights in the cited paper, and combining the combination of knowledge units to the degree of universality in the paper collection, calculating the novelty of the paper, and finally evaluating the originality of the paper based on the product of the weight and novelty, thereby screening out representative papers.
It improves the screening accuracy of representative papers and can more comprehensively reflect the degree and importance of the papers.
Smart Images

Figure CN120030154A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a paper screening method, system and related devices. Background Art
[0002] Due to some of the current demands, scholars are often asked to provide a certain range of representative academic papers.
[0003] At present, the main method of paper screening is to directly use the citation frequency of the paper for screening or indirectly use the citation frequency of the paper for screening, such as calculating the H-index of scholars or comparing the actual citation frequency of the paper with the frequency baseline and other indirect methods. Since the citation frequency of the paper cannot fully reflect the degree of innovation of the paper, there is a problem of low accuracy in screening representative papers if only the citation frequency of the paper is used for screening. Summary of the invention
[0004] In view of the above problems, this application provides a paper screening method, system and related devices to achieve the purpose of improving the screening accuracy of representative papers. The specific scheme is as follows:
[0005] The first aspect of the present application provides a paper screening method, the paper screening method comprising:
[0006] Acquire at least two knowledge units in a target paper, wherein the target paper is a published paper of a user in a paper collection, wherein the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge units are key contents in the target paper;
[0007] For each knowledge unit in the target paper, based on the position of the knowledge unit in the citing paper, the weight of the knowledge unit for the citing paper is calculated, and the citing paper is a paper that cites the target paper;
[0008] Calculating the sum of weights of the knowledge units in the target paper, and using the sum of weights as the weight of the target paper for the citing paper;
[0009] Obtain at least one combination pair of knowledge units in the target paper, calculate the prevalence of each combination pair in the sum of the combination pairs in the paper collection, and calculate the novelty of each combination pair based on the prevalence, wherein the sum of the combination pairs is: the sum of non-repeating combination pairs of knowledge units in each paper in the paper collection;
[0010] Calculating the sum of novelty levels of the combination pairs in the target paper, and taking the sum of novelty levels as the novelty level of the target paper;
[0011] The product of the weight of the target paper and the novelty of the target paper is taken as the originality of the target paper;
[0012] After obtaining the originality levels of the published papers in the paper collection, some of the published papers are selected from the published papers as representative papers of the user according to the originality levels.
[0013] In a possible implementation, the calculating the weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper includes:
[0014] Calculating the position weights of the knowledge units in the citing papers respectively, and performing data standardization on the position weights;
[0015] The sum of the standardized position weights of the knowledge unit is taken as the weight of the knowledge unit for the citing paper.
[0016] In a possible implementation, respectively calculating the position weights of the knowledge units in the citing papers includes:
[0017] For each of the knowledge units in each of the citing papers:
[0018] Counting the first frequency of the knowledge unit appearing in the title of the citing paper and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper;
[0019] Obtaining a first position weight of the title and a second position weight of the abstract paragraph;
[0020] Calculating a first product of the first frequency and the first position weight, and calculating a second product of the second frequency and the second position weight;
[0021] The sum of the first product and the second product is used as the position weight of the knowledge unit for the citing paper.
[0022] In a possible implementation, the combination pair includes at least a first knowledge unit and a second knowledge unit, and the calculating the prevalence of each combination pair in the total combination pairs of the paper collection includes:
[0023] Counting the total number of target combination pairs, the total number of target combination pairs being: the total number of combination pairs of knowledge units in each paper in the paper collection;
[0024] Obtain the co-occurrence number of the combination pair in the total of the target combination pairs;
[0025] Counting the number of first combination pairs that only include the first knowledge unit in the sum of the combination pairs, and the number of second combination pairs that only include the second knowledge unit in the sum of the combination pairs;
[0026] Calculate a third product of the co-occurrence quantity and the sum of the combination pairs, and a fourth product of the first combination pair quantity and the second combination pair quantity;
[0027] The ratio of the third product to the fourth product is used as the prevalence of the combination pair in the total combination pairs of the paper collection, wherein the third product is the numerator and the fourth product is the denominator.
[0028] In a possible implementation, calculating the novelty of each of the combination pairs based on the prevalence includes:
[0029] Normalizing the prevalence of each of the combination pairs;
[0030] The novelty of each of the combination pairs is calculated based on the prevalence of the normalized data.
[0031] A second aspect of the present application provides a paper screening system, the paper screening system comprising:
[0032] A knowledge extraction unit, used for obtaining at least two knowledge units in a target paper, wherein the target paper is a published paper of a user in a paper collection, wherein the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge units are key contents in the target paper;
[0033] A first calculation unit is used to calculate, for each knowledge unit in the target paper, a weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper, wherein the citing paper is a paper that cites the target paper;
[0034] A weight calculation unit, used to calculate the sum of weights of the knowledge units in the target paper, and use the sum of weights as the weight of the target paper for the citing paper;
[0035] A second calculation unit is used to obtain at least one combination pair of knowledge units in the target paper, calculate the prevalence of each combination pair in the sum of the combination pairs in the paper collection, and calculate the novelty of each combination pair based on the prevalence, wherein the sum of the combination pairs is: the sum of non-repeating combination pairs of knowledge units in each paper in the paper collection;
[0036] A combination calculation unit, used to calculate the sum of novelty degrees of the combination pairs in the target paper, and use the sum of novelty degrees as the novelty degree of the target paper;
[0037] A paper evaluation unit, used for taking the product of the weight of the target paper and the novelty of the target paper as the originality of the target paper;
[0038] The paper screening unit is used to select some published papers from the published papers as representative papers of the user according to the originality after obtaining the originality of the published papers in the paper collection.
[0039] In a possible implementation, the first computing unit may include:
[0040] A position weight calculation subunit, used to respectively calculate the position weight of the knowledge unit in each of the citing papers, and perform data standardization on each of the position weights;
[0041] The weight sum calculation subunit is used to take the sum of the standardized position weights of the knowledge unit as the weight of the knowledge unit for the citing paper.
[0042] In a possible implementation, the position weight calculation subunit is specifically configured as follows:
[0043] For each of the knowledge units in each of the cited papers: count the first frequency of the knowledge unit appearing in the title of the citing paper, and the second frequency of the knowledge unit appearing in the abstract of the citing paper; obtain the first position weight of the title and the second position weight of the abstract; calculate the first product of the first frequency and the first position weight, and calculate the second product of the second frequency and the second position weight; take the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.
[0044] A third aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0045] The memory is used to store computer programs;
[0046] The processor is used to execute the computer program so that the electronic device can implement the paper screening method of the above-mentioned first aspect or any implementation method of the first aspect.
[0047] The fourth aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the paper screening method of the above-mentioned first aspect or any implementation of the first aspect.
[0048] By means of the above technical scheme, the present application provides a paper screening method, system and related devices. The method evaluates the degree of creativity of the paper from two perspectives, namely, the weight of the target paper for the citing paper and the novelty of the target paper in the user's published papers and all reference papers. First, the knowledge unit as the key content is selected from the target paper. For the weight of the target paper for the citing paper, this method calculates the weight of each knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper, and takes the sum of the weights of the knowledge units in the target paper as the weight of the target paper for the citing paper. For the novelty of the target paper in the user's published papers and reference papers, this method obtains at least one combination pair of knowledge units in the target paper, calculates the prevalence of each combination pair in the sum of the combination pairs in the paper collection, and then calculates the novelty of each combination pair in the sum of the combination pairs in the paper collection, and takes the sum of the novelty of each combination pair in the target paper as the novelty of the target paper for the user's published papers and reference papers. Finally, the product of the target paper's weight and the target paper's novelty is used as the target paper's originality, and the user's representative papers are screened from the published papers in the paper collection based on the originality. This method does not screen papers based on the frequency of citations, but rather comprehensively screens papers based on the weight of the paper for citing papers and the novelty of the paper in the user's published papers and all reference papers. The weight of the paper for citing papers indicates the importance of the paper to the papers published after it, while the novelty of the paper indicates the novelty of the paper to the user's published papers and the reference papers published before it. Therefore, this method uses the papers screened based on the originality calculated based on the weight of the paper and the novelty as representative papers, which can effectively improve the accuracy of screening representative papers. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0050] Figure 1 A schematic diagram of a process for screening papers provided in an embodiment of the present application;
[0051] Figure 2 A schematic diagram of the structure of a paper screening system provided in an embodiment of the present application;
[0052] Figure 3 A hardware structure block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0054] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0055] The terms "first", "second" etc. in the specification of the application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable in appropriate circumstances, and this is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0056] In many scenarios, how scholars select representative papers from all the papers they have published is a problem that many scholars need to face. Representative papers can reflect the academic achievements of scholars in a certain field.
[0057] At present, the main method of paper screening is to directly use the citation frequency of the paper for screening or indirectly use the citation frequency of the paper for screening. The citation frequency of the paper is to screen the paper from the dimension of the paper's influence, which is difficult to fully reflect the originality or innovation of the paper content. For example, the content of a paper is very novel, but due to the limitation of the publication channel or language, it has not attracted attention, so the citation frequency of the paper is low. For example, some review papers, whose content is generally a summary and sorting of the existing achievements of other scholars, do not produce novel content, but usually can obtain a higher citation frequency.
[0058] Therefore, academic papers, as the product of research activities, are the carriers of knowledge content. Screening representative papers based solely on the frequency of citations of the papers only considers the citations of the papers, but does not consider the degree of innovation of the content of the papers themselves. Therefore, the screening accuracy is not high.
[0059] To solve the above problems, an embodiment of the present application provides a method for screening papers. This method evaluates and screens papers from the content level of the papers, and comprehensively evaluates the originality of the paper content by the content novelty of the paper in the user's published papers and the papers published before it, and the importance of the paper in the papers published after it, so as to screen out representative papers. The method for screening papers in the embodiment of the present application will be introduced in detail below with reference to the accompanying drawings.
[0060] Referring to Figure 1 , Figure 1 FIG. is a schematic flowchart of a method for screening papers provided by an embodiment of the present application. As Figure 1 shown, a data processing method provided by an embodiment of the present application may include steps S10 to S16, and these steps will be described in detail below.
[0061] S10. Obtain at least two knowledge units in the target paper. The target paper is a published paper of the user in the paper collection. The paper collection includes at least one published paper of the user and at least one reference paper of the published paper. The knowledge unit is the key content in the target paper.
[0062] Among them, the knowledge unit may be a concept, statement, word or phrase, term, law, etc. used in this solution to represent specific content. The key content in the target paper may refer to: content that has a high correlation with the research content or research direction of the target paper. Therefore, the knowledge unit may be a keyword, subject word or phrase in the target paper. In another alternative embodiment, the knowledge unit may also be a key sentence, paragraph, etc. in the target paper.
[0063] The knowledge unit can be directly selected fixedly. For example, the keywords located after the abstract and before the text in the target paper are directly selected as the knowledge unit in this embodiment. It is also possible to extract keywords, subject words or phrases in the target paper through a third-party extraction tool and use them as the knowledge unit. When the knowledge unit is a phrase, there may be a semantic relationship between the multiple words that make up the phrase, such as the subject-object relationship.
[0064] When extracting knowledge units, third-party extraction tools can not only extract single words directly as knowledge units, but also extract a complete semantic phrase and select some words from the semantic phrase to form a phrase as a knowledge unit to more accurately reflect the knowledge content in the paper. For example, the extraction tool can extract two words A and B separately from the target paper, then A can be used as a knowledge unit, and B can be used as a knowledge unit. The extraction tool can also extract the phrase "A treats B" that contains a semantic relationship from the target paper. A is the subject, "treatment" is the predicate, and B is the object. Then A and B can be selected as words, and the phrase composed of AB can be used as a knowledge unit. At this time, there is a subject-object relationship between A and B.
[0065] A paper collection can be a collection of multiple published papers of a user and multiple reference papers of the published papers, where you can select some reference papers of the published papers or all reference papers of the published papers. In this paper collection, multiple published papers of a user can belong to the same field or research direction, and can also belong to the same publication time period, such as all papers published by a user within five years or all papers published within one year.
[0066] S11. For each knowledge unit in the target paper, the weight of the knowledge unit for the citing paper is calculated based on the position of the knowledge unit in the citing paper. The citing paper is the paper that cites the target paper.
[0067] S12. Calculate the sum of the weights of the knowledge units in the target paper, and use the sum of the weights as the weight of the target paper for the citing paper.
[0068] The weight of a knowledge unit for citing papers can be expressed as: the usefulness or importance of the knowledge unit of the target paper in all citing papers. The weight of a target paper for citing papers can be expressed as: the influence of the target paper on citing papers or the usefulness or importance of the content of the target paper on citing papers, or the usefulness or importance of the target paper in a field.
[0069] In this embodiment, the citing papers of the target paper can be counted first, and then the weight of each knowledge unit in the target paper for the citing papers can be calculated. Finally, the weight of each knowledge unit in the target paper for the citing papers can be summed to obtain the weight of the target paper for the citing papers.
[0070] Specifically, the specific calculation process of the weight of the knowledge unit for citing papers can be shown in steps 1 and 2:
[0071] Step 1: Calculate the position weights of knowledge units in each citing paper and standardize the data of each position weight;
[0072] Step 2: Use the sum of the standardized position weights of each knowledge unit as the weight of the knowledge unit for the citing paper.
[0073] Among them, the position weight of a knowledge unit in a citing paper can be: the degree of usefulness of the knowledge unit for the citing paper calculated based on the occurrence position of the knowledge unit in a citing paper of the target paper. For example, a knowledge unit that appears in the title of a citing paper indicates that the knowledge unit is highly relevant to the content of the citing paper, so the importance of the knowledge unit for the citing paper is relatively high. In this embodiment, for each knowledge unit in the target paper, this embodiment can first calculate the position weight of the knowledge unit for each citing paper according to the position of the knowledge unit in the target paper in each citing paper, and finally use the sum of all the position weights of the knowledge unit as the weight of the knowledge unit for the citing paper.
[0074] Specifically, the specific calculation process of the position weight of a knowledge unit in a citing paper can be as shown in Steps 3 to 6:
[0075] For each knowledge unit in each citing paper:
[0076] Step 3: Count the first frequency of the knowledge unit appearing in the title of the citing paper and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper;
[0077] Step 4: Obtain the first position weight of the title and the second position weight of the abstract paragraph;
[0078] Step 5: Calculate the first product of the first frequency and the first position weight, and calculate the second product of the second frequency and the second position weight;
[0079] Step 6: Use the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.
[0080] Among them, the formula form of the above content can be expressed as:
[0081]
[0082] Among them, can represent the knowledge unit; can represent the knowledge unit For the position weight of the citing paper, can represent the position weight. In this embodiment can include the first position weight and the second position weight , can represent the total number of position weight types, can represent the summed variable, in this embodiment, The value of can be 2; Can represent knowledge units The frequency of occurrence in citing papers, in this example Can include the first frequency and the second frequency .
[0083] The first frequency may be the number of times the knowledge unit appears in the title of the citing paper, and the second frequency may be the number of times the knowledge unit appears in the abstract of the citing paper. The first position weight is the preset weight of the title, and the second position weight is the preset weight of the abstract. In this embodiment, the value of the first position weight may be 2, and the value of the second position weight may be 1.
[0084] Of course, in another optional embodiment, when calculating the position weight of the knowledge unit in the citing paper, the statistical position of the knowledge unit in the citing paper may include not only the title and abstract, but also other positions in the citing paper, such as the summary paragraphs of each chapter in the citing paper.
[0085] Since for a knowledge unit, its importance may vary depending on the position of its appearance in the citing paper, and its importance may also vary depending on the number of times it appears in different positions in the citing paper. Therefore, this embodiment not only considers the position of the knowledge unit in the citing paper, but also considers the number of times the knowledge unit appears in different positions in the citing paper, and performs a weighted sum calculation on the number of appearances and the preset weights corresponding to the different positions of appearance to obtain the position weight of the knowledge unit for the citing paper. For example, knowledge unit A appears once in the title of a citing paper and the preset weight corresponding to the title is 2, and appears 3 times in the abstract paragraph of the citing paper and the preset weight corresponding to the abstract paragraph is 1, then the position weight of knowledge unit A for the citing paper is 5 (5=1×2+3×1).
[0086] In another optional embodiment, the total weight of citing papers at different positions can be directly set. When a knowledge unit appears in different positions of citing papers, the total weights corresponding to each position can be summed up, and then the average value can be calculated and used as the position weight of the knowledge unit in the citing papers.
[0087] Furthermore, when the target paper has only one citing paper, the position weight of the knowledge unit in the citing paper can be directly used as the weight of the knowledge unit for the citing paper. When the target paper has multiple citing papers, the sum of the position weights of the knowledge unit in each citing paper can be used as the weight of the knowledge unit for the citing paper, and the formula can be as follows:
[0088]
[0089] in, Can represent knowledge units The weight of citing papers; It can represent the total number of citing papers; Can represent knowledge units The position weight in citing papers after data normalization.
[0090] The formula for data normalization can be shown as follows:
[0091]
[0092] in, Can represent knowledge units The position weight in citing papers after data normalization; Can represent knowledge units Position weight in citing papers; Can represent knowledge units The maximum value among the weights of each position.
[0093] Since the length of citing papers varies, the length of the titles (title + abstract) of citing papers may vary accordingly. In a citing paper with a longer abstract, the knowledge unit may appear repeatedly in the abstract, resulting in a greater weight of the knowledge unit in the citing paper. In a citing paper with a shorter abstract, the knowledge unit may appear only once in the abstract, resulting in a smaller weight of the knowledge unit in the citing paper. Therefore, the length of citing papers can affect the weight of the knowledge unit for citing papers. In order to reduce the impact of some outliers, this embodiment requires data standardization before calculating the weight of the knowledge unit for citing papers.
[0094] After obtaining the weight of each knowledge unit in the target paper for the citing paper in this embodiment, the weight of the target paper for the citing paper can be obtained by calculating the sum of the weights of all knowledge units in the target paper. The calculation formula can be as follows:
[0095]
[0096] in, Can indicate target paper The weight of citing papers; Can indicate target paper The total number of knowledge units; Can represent knowledge units The weight of citing papers.
[0097] In this embodiment, by calculating the weight of the target paper with respect to the citing papers, since the citing papers are the papers that cite the target paper, the weight of the target paper with respect to the citing papers can reflect the importance of the content in the target paper for the citing papers published after the target paper. It is precisely because the content in the target paper has reference significance or novel content that other papers do not have that the citing papers choose to refer to the target paper. Therefore, the weight of the target paper calculated in this embodiment can well reflect the innovation of the content of the target paper.
[0098] S13. Obtain at least one combination pair of knowledge units in the target paper, calculate the prevalence of each combination pair in the total combination pairs of the paper collection, and calculate the novelty of each combination pair based on the prevalence. The total combination pairs are: the sum of the non-repeated combination pairs of knowledge units in each paper in the paper collection.
[0099] S14. Calculate the total novelty of the combination pairs in the target paper, and use the total novelty as the novelty of the target paper.
[0100] Among them, there may not be a strict order of execution between steps S11 - S12 and steps S13 - S14. In an alternative embodiment, steps S11 - S12 can be executed first to calculate the weight of the target paper, and then steps S13 - S14 can be executed to calculate the novelty of the target paper. In another alternative embodiment, steps S13 - S14 can be executed first to calculate the novelty of the target paper, and then steps S11 - S12 can be executed to calculate the weight of the target paper.
[0101] The combination pair can be: a phrase containing at least two knowledge units. In the target paper, if the knowledge units are only keywords or subject terms, multiple combination pairs of knowledge units in the target paper can be obtained through pairwise combination; if the knowledge units are only phrases, each phrase can be directly used as a combination pair, thereby obtaining multiple combination pairs of knowledge units in the target paper; if the knowledge units include both keywords or subject terms and phrases, some combination pairs can be obtained by pairwise combination of the keywords or subject terms first, and then each phrase can be used as a combination pair to obtain another part of the combination pairs. Finally, the part of the combination pairs and the other part of the combination pairs are used together as multiple combination pairs of knowledge units in the target paper. Of course, in this embodiment, the words in the phrase can also be extracted first, and then pairwise combined with the keywords or subject terms to finally obtain multiple combination pairs of knowledge units in the target paper.
[0102] Since the paper collection contains multiple papers, each paper may have multiple combinations of knowledge units, and there may be repeated combinations. In this embodiment, the total number of combinations may be: the sum of the non-repeating combinations in each paper in the paper collection. For example, the paper collection contains paper A and paper B. The combination pairs of knowledge units in paper A include: ab, ac, and bc, and the combination pairs of knowledge units in paper B include: ab, ad, and bd. Since the combination pair ab appears repeatedly, the total number of combinations in the paper collection is 5, not 6.
[0103] The prevalence of a combination pair can be: the extent to which the combination pair is widespread in the collection of papers or the extent to which the combination pair is adopted or used in the collection of papers. The novelty of a combination pair can be: the extent to which the combination pair is unique in the collection of papers or the extent to which the combination pair does not appear or is not used in the collection of papers.
[0104] Since directly counting the number of occurrences of the target paper's knowledge units in the paper collection is too direct and difficult to measure the target paper's novelty well, this embodiment chooses to convert the target paper's knowledge units into combination pairs, first by calculating the prevalence of the combination pairs in the paper collection, then calculating the novelty of the combination pairs, and finally taking the sum of the novelty of each combination pair in the target paper as the novelty of the target paper.
[0105] Specifically, each combination pair includes at least the first knowledge unit and the second knowledge unit. The specific calculation process of the prevalence of the combination pair in the target paper can be shown in steps 7 to 11:
[0106] Step 7: Count the total number of target combination pairs. The total number of target combination pairs is: the total number of combination pairs of knowledge units in each paper in the paper collection;
[0107] Step 8: Obtain the number of co-occurrences of the combination pair in the total of the target combination pair;
[0108] Step 9: Count the number of first-group pairs that only contain the first knowledge unit in the sum of the combination pairs, and the number of second-group pairs that only contain the second knowledge unit in the sum of the combination pairs;
[0109] Step 10: Calculate the third product of the number of co-occurrences and the sum of the combination pairs, and the fourth product of the number of pairs in the first group and the number of pairs in the second group;
[0110] Step 11: Take the ratio of the third product to the fourth product as the prevalence of the combination pair in the total number of combination pairs in the paper collection, where the third product is the numerator and the fourth product is the denominator.
[0111] The target combination pair sum can be: the sum of the number of each combination pair of knowledge units in each paper in the paper collection. The difference between the target combination pair sum and the combination pair sum is that the target combination pair sum includes repeated combination pairs. For example, the paper collection contains paper A and paper B. The combination pairs of knowledge units in paper A include: ab, ac and bc, and the combination pairs of knowledge units in paper B include: ab, ad and bd. Even if the combination pair ab is repeated, it is still counted in the number. Therefore, the target combination pair sum of the paper collection is 6. If it is the combination pair sum, the repeated combination pairs are not counted in the number, and the combination pair sum is 5.
[0112] The number of co-occurrences of a combination pair in the total of target combination pairs is: the number of combination pairs in the total of target combination pairs. Since the number of co-occurrences of a combination pair represents the number of times the two knowledge units constituting the combination pair appear in a paper at the same time, repeated combination pairs need to be counted in the total of target combination pairs in this embodiment, otherwise the calculation of the prevalence of the combination pairs may be affected. For example, the paper collection contains paper A and paper B. The combination pairs of knowledge units in paper A include: ab, ac and bc, and the combination pairs of knowledge units in paper B include: ab, ad and bd. If repeated combination pairs are not counted, the number of co-occurrences of combination pair ab is only 1, but combination pair ab appears in both paper A and paper B, and the actual number of co-occurrences of combination pair ab should be 2. If the co-occurrence number is calculated as 1, the prevalence of combination pair ab is low, but the actual prevalence of combination pair ab should be high.
[0113] The first combination pairs may be: a combination pair in which one knowledge unit is the first knowledge unit, and does not include a combination pair of the second knowledge unit, then the number of first combination pairs may be the number of first combination pairs in the total number of combination pairs, and the second combination pairs may be: a combination pair in which one knowledge unit is the second knowledge unit, and does not include a combination pair of the first knowledge unit, then the number of second combination pairs may be the number of second combination pairs in the total number of combination pairs.
[0114] Furthermore, the formula form of the above content can be as follows:
[0115]
[0116] in, Can be represented by knowledge units and knowledge units The prevalence of the formed combination pairs. The greater the prevalence, the more knowledge units and knowledge units The more common the constructed pair is in the collection of papers, the less novel it is; Can be represented by knowledge units and knowledge units The number of co-occurrences of the constructed combination pairs in the total number of target combination pairs; It can be said that the sum of the combination pairs only contains knowledge units The number of pairs in the first group; It can be said that the sum of the combination pairs only contains knowledge units The number of pairs in the second group; It can represent the sum of the combined pairs of a collection of papers.
[0117] After calculating the prevalence of all combination pairs in the target paper, this embodiment needs to first normalize the prevalence of each combination pair in the target paper (linear function normalization), and then calculate the novelty of each combination pair in the target paper based on the prevalence after data normalization. Specifically, the formula for data normalization can be as follows:
[0118]
[0119] in, It can represent the normalized data by knowledge unit and knowledge units the prevalence of the formed pairs; It can be the minimum value of prevalence among all the combination pairs in the target paper; It can be the maximum value of the prevalence of all combination pairs in the target paper.
[0120] In this embodiment, the prevalence of all combination pairs in the target paper can be adjusted to between [0, 1] by data normalization, so as to calculate the novelty of the combination pair by the prevalence of the normalized combination pair data. The calculation formula can be as follows:
[0121]
[0122] in, Can be represented by knowledge units and knowledge units The novelty of the formed pair; It can represent the normalized data by knowledge unit and knowledge units The prevalence of the formed pairs.
[0123] Thus, this embodiment can obtain the novelty of all combination pairs in the target paper, then calculate the sum of the novelty of all combination pairs in the target paper, and use the sum of the novelty as the novelty of the target paper. The formula form can be as follows:
[0124]
[0125] in, Can represent target paper the novelty of Can represent target paper The total number of knowledge units, if the knowledge unit is a keyword or subject word, the corresponding In another optional embodiment, if the knowledge unit is a phrase, the corresponding Combination pairs; Can be represented by knowledge units and knowledge units The novelty of the formed pair.
[0126] In this embodiment, the paper collection is a collection of the user's published papers and all reference papers of the published papers, the target paper is one of the published papers, and the reference paper is a paper published before the target paper. Therefore, this embodiment calculates the novelty of the target paper in the paper collection to indicate the novelty of the target paper compared with previously published papers, so as to reflect the innovation of the content of the target paper.
[0127] S15. The product of the weight of the target paper and the novelty of the target paper is taken as the originality of the target paper;
[0128] S16. After obtaining the originality of the published papers in the paper collection, select some of the published papers as representative papers of the user according to the originality.
[0129] Among them, the originality of the target paper can be expressed as: the uniqueness of the content of the target paper, and the degree to which it differs from existing achievements. This embodiment can design a creativity measurement model for the target paper, and use the creativity measurement model to calculate the originality of the target paper. Specifically, this embodiment can directly input the paper collection and citing papers into the creativity measurement model. The creativity measurement model calculates the weight of the target paper for the citing papers, the novelty of the target paper, and the originality of the target paper, and finally obtains the originality of the target paper output by the creativity measurement model. This embodiment can also calculate the originality of the target paper by inputting the calculated weight of the target paper for the citing papers and the novelty of the target paper into the creativity measurement model, wherein the formula representation of the creativity measurement model can be as follows:
[0130]
[0131] in, Can represent target paper The degree of originality; Can represent target paper the novelty of Can represent target paper The weight of citing papers.
[0132] After obtaining the originality of all published papers in the paper collection, this embodiment can sort all published papers according to the originality, and select some published papers with high originality as representative papers of the user.
[0133] The present application provides a paper screening method, which evaluates the degree of creativity of the paper from two perspectives, namely, the weight of the target paper for the citing paper and the novelty of the target paper in the user's published papers and all reference papers. First, the knowledge unit as the key content is selected from the target paper. For the weight of the target paper for the citing paper, this method calculates the weight of each knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper, and takes the sum of the weights of the knowledge units in the target paper as the weight of the target paper for the citing paper. For the novelty of the target paper in the user's published papers and reference papers, this method obtains at least one combination pair of knowledge units in the target paper, calculates the prevalence of each combination pair in the sum of the combination pairs of the paper collection, and then calculates the novelty of each combination pair in the sum of the combination pairs of the paper collection, and takes the sum of the novelty of each combination pair in the target paper as the novelty of the target paper for the user's published papers and reference papers. Finally, the product of the weight of the target paper and the novelty of the target paper is used as the originality of the target paper, and the representative papers of the user are screened out from the published papers in the paper collection according to the originality. This method does not screen papers based on the frequency of citations, but rather comprehensively screens papers based on the weight of the paper for citing papers and the novelty of the paper among the user's published papers and all reference papers. The weight of the paper for citing papers indicates the importance of the paper to the papers published after it, while the novelty of the paper indicates the novelty of the paper to the user's published papers and the reference papers published before it. Therefore, this method uses the papers selected based on the degree of originality calculated based on the weight of the paper and the novelty as representative papers, which can effectively improve the accuracy of screening representative papers.
[0134] The above introduces a paper screening method provided by an embodiment of the present application, and the following will introduce a system applying the above paper screening method.
[0135] See also Figure 2 , Figure 2 This is a schematic diagram of the structure of a paper screening system provided in an embodiment of the present application. Figure 2 As shown, the paper screening system includes:
[0136] The knowledge extraction unit 100 is used to obtain at least two knowledge units in a target paper, where the target paper is a published paper of a user in a paper collection, and the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge unit is the key content of the target paper;
[0137] A first calculation unit 110 is used to calculate, for each knowledge unit in the target paper, the weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper, where the citing paper is the paper that cites the target paper;
[0138] A weight calculation unit 120, used to calculate the sum of weights of the knowledge units in the target paper, and use the sum of weights as the weight of the target paper for the citing paper;
[0139] The second calculation unit 130 is used to obtain at least one combination pair of knowledge units in the target paper, calculate the prevalence of each combination pair in the total combination pairs of the paper collection, and calculate the novelty of each combination pair based on the prevalence, and the total combination pair is: the sum of non-repeated combination pairs of knowledge units in each paper in the paper collection;
[0140] A combination calculation unit 140 is used to calculate the sum of novelty degrees of the combination pairs in the target paper, and use the sum of novelty degrees as the novelty degree of the target paper;
[0141] The paper evaluation unit 150 is used to take the product of the weight of the target paper and the novelty of the target paper as the originality of the target paper;
[0142] The paper screening unit 160 is used to select some published papers as representative papers of the user from the published papers according to the originality after obtaining the originality of the published papers in the paper collection.
[0143] In a possible implementation, the first computing unit 110 may include:
[0144] The position weight calculation subunit is used to calculate the position weight of each knowledge unit in each citing paper and perform data standardization on each position weight;
[0145] The weight sum calculation subunit is used to take the sum of the standardized position weights of each knowledge unit as the weight of the knowledge unit for citing papers.
[0146] In a possible implementation, the position weight calculation subunit may be specifically configured as follows:
[0147] For each knowledge unit in each citing paper: count the first frequency of the knowledge unit appearing in the title of the citing paper, and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper; obtain the first position weight of the title and the second position weight of the abstract paragraph; calculate the first product of the first frequency and the first position weight, and calculate the second product of the second frequency and the second position weight; take the sum of the first product and the second product as the position weight of the knowledge unit for the citing paper.
[0148] In a possible implementation, the combination pair includes at least the first knowledge unit and the second knowledge unit. The second calculation unit 130 calculates the prevalence of each combination pair in the total number of combination pairs in the paper collection, which can be specifically configured as follows:
[0149] Count the total number of target combination pairs, which is: the sum of the number of combination pairs of knowledge units in each paper in the paper collection; obtain the co-occurrence number of combination pairs in the total number of target combination pairs; count the number of first combination pairs that only contain the first knowledge unit in the total number of combination pairs, and the number of second combination pairs that only contain the second knowledge unit in the total number of combination pairs; calculate the third product of the co-occurrence number and the total number of combination pairs, and the fourth product of the first number of combination pairs and the second number of combination pairs; take the ratio of the third product to the fourth product as the prevalence of the combination pair in the total number of combination pairs in the paper collection, where the third product is the numerator and the fourth product is the denominator.
[0150] In a possible implementation, the second calculation unit 130 calculates the novelty of each combination pair based on the prevalence, which can be specifically configured as follows:
[0151] The prevalence of each combination pair is normalized; and the novelty of each combination pair is calculated based on the prevalence after normalization.
[0152] The present application also provides an electronic device in an embodiment. Figure 3 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 3 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0153] like Figure 3As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 to a random access memory (RAM) 303. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0154] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a memory card, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0155] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the paper screening methods provided in the embodiments of the present application.
[0156] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any paper screening method provided in the embodiment of the present application.
[0157] It should also be noted that the system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the system embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0158] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0159] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0160] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
[0161] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0162] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0163] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A paper screening method, characterized in that: The paper screening method includes: Acquire at least two knowledge units in a target paper, wherein the target paper is a published paper of a user in a paper collection, wherein the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge units are key contents in the target paper; For each knowledge unit in the target paper, based on the position of the knowledge unit in the citing paper, the weight of the knowledge unit for the citing paper is calculated, and the citing paper is a paper that cites the target paper; Calculating the sum of weights of the knowledge units in the target paper, and using the sum of weights as the weight of the target paper for the citing paper; Obtain at least one combination pair of knowledge units in the target paper, calculate the prevalence of each combination pair in the sum of the combination pairs in the paper collection, and calculate the novelty of each combination pair based on the prevalence, wherein the sum of the combination pairs is: the sum of non-repeating combination pairs of knowledge units in each paper in the paper collection; Calculating the sum of novelty levels of the combination pairs in the target paper, and taking the sum of novelty levels as the novelty level of the target paper; The product of the weight of the target paper and the novelty of the target paper is taken as the originality of the target paper; After obtaining the originality levels of the published papers in the paper collection, some of the published papers are selected from the published papers as representative papers of the user according to the originality levels.
2. The paper screening method according to claim 1, characterized in that: The calculating the weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper includes: Calculating the position weights of the knowledge units in each of the citing papers respectively, and performing data standardization on each of the position weights; The sum of the standardized position weights of the knowledge unit is taken as the weight of the knowledge unit for the citing paper.
3. The paper screening method according to claim 2, characterized in that: The respectively calculating the position weights of the knowledge units in the citing papers comprises: For each of the knowledge units in each of the citing papers: Counting the first frequency of the knowledge unit appearing in the title of the citing paper and the second frequency of the knowledge unit appearing in the abstract paragraph of the citing paper; Obtaining a first position weight of the title and a second position weight of the abstract paragraph; Calculating a first product of the first frequency and the first position weight, and calculating a second product of the second frequency and the second position weight; The sum of the first product and the second product is used as the position weight of the knowledge unit for the citing paper.
4. The paper screening method according to claim 1, characterized in that: The combination pair includes at least a first knowledge unit and a second knowledge unit, and the calculating the prevalence of each combination pair in the total combination pairs of the paper collection includes: Counting the total number of target combination pairs, the total number of target combination pairs being: the total number of combination pairs of knowledge units in each paper in the paper collection; Obtain the co-occurrence number of the combination pair in the total of the target combination pairs; Counting the number of first combination pairs that only include the first knowledge unit in the sum of the combination pairs, and the number of second combination pairs that only include the second knowledge unit in the sum of the combination pairs; Calculate a third product of the co-occurrence quantity and the sum of the combination pairs, and a fourth product of the first combination pair quantity and the second combination pair quantity; The ratio of the third product to the fourth product is used as the prevalence of the combination pair in the total combination pairs of the paper collection, wherein the third product is the numerator and the fourth product is the denominator.
5. The paper screening method according to claim 1 or 4, characterized in that: The calculating the novelty of each of the combination pairs based on the prevalence comprises: Normalizing the prevalence of each of the combination pairs; The novelty of each of the combination pairs is calculated based on the prevalence of the normalized data.
6. A paper screening system, characterized in that: The paper screening system includes: A knowledge extraction unit, used for obtaining at least two knowledge units in a target paper, wherein the target paper is a published paper of a user in a paper collection, wherein the paper collection includes at least one published paper of the user and at least one reference paper of the published paper, and the knowledge units are key contents in the target paper; A first calculation unit is used to calculate, for each knowledge unit in the target paper, a weight of the knowledge unit for the citing paper based on the position of the knowledge unit in the citing paper, wherein the citing paper is a paper that cites the target paper; A weight calculation unit, used to calculate the sum of weights of the knowledge units in the target paper, and use the sum of weights as the weight of the target paper for the citing paper; A second calculation unit is used to obtain at least one combination pair of knowledge units in the target paper, calculate the prevalence of each combination pair in the sum of the combination pairs in the paper collection, and calculate the novelty of each combination pair based on the prevalence, wherein the sum of the combination pairs is: the sum of non-repeating combination pairs of knowledge units in each paper in the paper collection; A combination calculation unit, used to calculate the sum of novelty degrees of the combination pairs in the target paper, and use the sum of novelty degrees as the novelty degree of the target paper; A paper evaluation unit, used for taking the product of the weight of the target paper and the novelty of the target paper as the originality of the target paper; The paper screening unit is used to select some published papers from the published papers as representative papers of the user according to the originality after obtaining the originality of the published papers in the paper collection.
7. The paper screening system according to claim 6, characterized in that: The first computing unit includes: A position weight calculation subunit, used to respectively calculate the position weight of the knowledge unit in each of the citing papers, and perform data standardization on each of the position weights; The weight sum calculation subunit is used to take the sum of the standardized position weights of the knowledge unit as the weight of the knowledge unit for the citing paper.
8. The paper screening system according to claim 7, characterized in that: The position weight calculation subunit is specifically configured as follows: For each of the knowledge units in each of the cited papers: counting the first frequency of the knowledge unit appearing in the title of the cited paper and the second frequency of the knowledge unit appearing in the abstract of the cited paper; obtaining the first position weight of the title and the second position weight of the abstract; Calculating a first product of the first frequency and the first position weight, and calculating a second product of the second frequency and the second position weight; The sum of the first product and the second product is used as the position weight of the knowledge unit for the citing paper.
9. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the paper screening method as described in any one of claims 1 to 5.
10. A computer program product, characterized in that It includes computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the paper screening method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Academic paper-based technological frontier index calculation method and system
CN108614867A
Method for calculating scholarship transboundary flow
CN111105327A
New thesis recommendation method and system based on distinguishable themes
CN116628350A
Urban scientific and technological innovation power assessment method based on mapping knowledge domain
CN116739085A
Paper recommendation method based on knowledge graph and graph neural network
CN118364139A