Document information pushing method, device and system, and storage medium
By comprehensively considering download count, citation frequency, and H5 index, the overall influence of the literature is evaluated, which solves the problem of the single source of literature in the literature search system and realizes the diversified push and comprehensive recommendation of high-quality academic literature.
Patent Information
- Application Number
- CN202311865782.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2043-12-29
AI Technical Summary
In existing literature search systems, when filtering high-quality literature solely based on H5 indexes, it is easy to miss high-quality literature from less well-known journals, resulting in one-sided and incomplete literature information being pushed, leading to a poor user experience.
By comprehensively considering three indicators—download count, citation frequency, and H5 index—the comprehensive impact index parameters for each article are determined, and the articles with the greatest comprehensive impact are recommended to improve the diversity and comprehensiveness of the literature sources.
It has diversified the sources of literature, and the literature pushed has good peer reviews, high professionalism and high scholar recognition, which has improved the user push experience and the comprehensiveness and objectivity of literature recommendations.
Smart Images

Figure CN117880353B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to methods, apparatus, systems and storage media for pushing document information. Background Technology
[0002] With the rapid increase in the number of published documents, literature from different journals is stored in a literature search system to enable users to comprehensively search for literature. Given the large number of documents stored in such systems, to facilitate users' quick access to high-quality literature with high interest, a filter is typically applied to high-quality documents from various journals that meet the H5 index criteria, and these H5 index-based high-quality documents are automatically recommended to users. However, the high-quality documents selected using the single-dimensional static indicator of the H5 index usually come from well-known journals, resulting in a relatively singular source of literature. This can easily lead to the omission of high-quality documents from less well-known journals, resulting in a one-sided and incomplete recommendation of literature information and a poor user experience. Summary of the Invention
[0003] This invention provides a method, apparatus, system, and storage medium for pushing literature information, aiming to at least solve the technical problem that current literature push methods often rely on a single source of literature, easily overlooking high-quality literature from lesser-known journals, thus resulting in a one-sided and incomplete push of literature. The technical solution of this invention is as follows:
[0004] According to a first aspect of the present invention, a method for pushing literature information is provided. The method includes: determining the target number of downloads for each literature, determining the target citation frequency for each literature, and determining the target H5 index of the journal to which each literature belongs; determining the comprehensive impact index parameter of each literature based on the target number of downloads, target citation frequency, and target H5 index; the comprehensive impact index parameter of each literature being positively correlated with the target number of downloads, target citation frequency, and target H5 index of each literature; identifying the literature with the largest comprehensive impact index parameter among a preset number of literatures as target literatures; and pushing the literature information of the target literatures.
[0005] In one possible implementation, determining the target download count for each document includes: obtaining the actual download count of each document within a first preset time period; determining the maximum and minimum actual download counts among the various actual download counts for each document; and normalizing the various actual download counts for each document based on the maximum and minimum actual download counts to obtain the target download counts for each document.
[0006] In another possible implementation, the target citation frequency of each document is determined, including: obtaining the actual number of citations of each document within a first preset time period; determining the maximum and minimum actual citation counts among the actual citation counts of each document; and normalizing the actual citation counts of each document based on the maximum and minimum actual citation counts to obtain the target citation counts for each document.
[0007] In another possible implementation, the target H5 index of the journal to which each article belongs is determined, including: obtaining the actual H5 index of the journal to which each article belongs; normalizing the absolute median difference of each actual H5 index to obtain the target weights corresponding to each actual H5 index; weighting the differences between each actual H5 index and each weight to determine the weighted indexes of each actual H5 index; for any target weight, any weight difference represents the difference between 1 and any target weight; and normalizing each weighted index based on the frequency of the maximum and minimum weighted indices to obtain the target H5 indexes corresponding to the journal to which each article belongs.
[0008] In another possible implementation, the normalized absolute median difference of each actual H5 index is processed to obtain the target weights corresponding to each actual H5 index. This includes: arranging the actual H5 indices of each journal to which each article belongs in descending order; determining the absolute median difference of each actual H5 index based on the actual H5 index in the middle of the order; and normalizing each absolute median difference based on the largest and smallest absolute median differences to obtain the target weights.
[0009] In another possible implementation, based on the actual H5 index at the middle position, the absolute median difference corresponding to each actual H5 index is determined, including: when the number of journals N to which each article belongs is odd, that is, when the number of actual H5 indices N is odd, the middle actual H5 index is the actual H5 index at the (1+N) / 2 position; when the number of journals N to which each article belongs is even, that is, when the number of actual H5 indices N is even, the middle actual H5 index is the average of the actual H5 indices at the (1+N / 2) and (N / 2) positions; the absolute value of the difference between each actual H5 index and the middle actual H5 index is determined as the absolute median difference corresponding to each actual H5 index.
[0010] In another possible implementation, each document refers to documents published within a second preset time interval from the current time in the journal. The comprehensive impact index parameters for each document are determined based on its target download count, target citation frequency, and target H5 index. This includes: for any document among all documents, the target download count, target citation frequency, and target H5 index are weighted and summed with a first weight, a second weight, and a third weight, respectively, to obtain the comprehensive impact index parameters for that document; both the first and third weights are greater than the second weight; and the sum of the first, second, and third weights is 1.
[0011] According to a second aspect of the present invention, a document information push device is provided, comprising: a determining unit configured to determine the target number of downloads for each document, the target citation frequency for each document, and the target H5 index of the journal to which each document belongs; an evaluation unit configured to determine the comprehensive impact index parameters of each document based on the target number of downloads, the target citation frequency, and the target H5 index; wherein the comprehensive impact index parameters of each document are positively correlated with the target number of downloads, the target citation frequency, and the target H5 index; and a push unit configured to determine the document with the largest comprehensive impact index parameter among a preset number of documents as target documents, and to push the document information of the target documents.
[0012] In one possible implementation, the determining unit is specifically configured to: obtain the actual number of downloads of each document within a first preset time period; determine the maximum and minimum actual download counts among the actual download counts of each document; and, based on the maximum and minimum actual download counts, normalize the actual download counts of each document to obtain the target download counts corresponding to each document.
[0013] In another possible implementation, the determining unit is specifically configured to: obtain the actual number of citations of each document within a first preset time period; determine the maximum and minimum actual citation counts among the actual citation counts of each document; and normalize the actual citation counts of each document based on the maximum and minimum actual citation counts to obtain the target citation counts corresponding to each document.
[0014] In another possible implementation, the determining unit is specifically configured as follows: obtaining the actual H5 index of the journal to which each article belongs; performing normalized absolute median difference processing on each actual H5 index to obtain each target weight corresponding to each actual H5 index; weighting the difference between each actual H5 index and each weight to determine each weighted index of each actual H5 index; for any target weight, any weight difference represents the difference between 1 and any target weight; and based on the frequency of the maximum and minimum weighted indices among the weighted indices, performing normalization processing on each weighted index to obtain each target H5 index corresponding to the journal to which each article belongs.
[0015] In another possible implementation, the determining unit is further configured as follows: arranging the actual H5 indices of the journals to which each article belongs in descending order; determining the absolute median differences corresponding to each actual H5 index based on the actual H5 index located in the middle of the order; and normalizing each absolute median difference based on the largest and smallest absolute median differences to obtain each target weight.
[0016] In another possible implementation, the determining unit is further configured as follows: when the number N of journals to which each article belongs is odd, that is, when the number N of each actual H5 index is odd, the middle actual H5 index is the actual H5 index at position (1+N) / 2; when the number N of journals to which each article belongs is even, that is, when the number of each actual H5 index is even, the middle actual H5 index is the average of the actual H5 indices at positions (1+N / 2) and (N / 2); the absolute value of the difference between each actual H5 index and the middle actual H5 index is determined as the absolute median difference corresponding to each actual H5 index.
[0017] In another possible implementation, each document refers to the documents published in journals within a second preset time interval from the current time. The evaluation unit is specifically configured as follows: for any document in each document, the target download count, target citation frequency, and target H5 index of any document are weighted and summed with the first weight, the second weight, and the third weight, respectively, to obtain the comprehensive impact index parameter of any document; the first weight and the third weight are both greater than the second weight; the sum of the first weight, the second weight, and the third weight is 1.
[0018] According to a third aspect of the present invention, a document retrieval system is provided, including a processor and a memory for storing processor-executable instructions; wherein the processor is configured to execute the document information push method described in the first aspect and any possible implementation thereof.
[0019] According to a fourth aspect of the present invention, an electronic device is provided, comprising: a processor and a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to implement a document information push method as described in the first aspect and any possible implementation thereof.
[0020] According to a fifth aspect of the present invention, a computer-readable storage medium is provided, on which instructions are stored, such that when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the document information push method of the first aspect.
[0021] According to a sixth aspect of the present disclosure, a computer program product is provided, the computer program product including computer instructions, which, when executed on an electronic device, cause the electronic device to perform the document information push method of the first aspect and any possible implementation thereof.
[0022] The technical solution provided by the embodiments of the present invention brings at least the following beneficial effects: Based on three different dimensions of indicators—download count, citation frequency, and H5 index—the comprehensive impact index of each document is evaluated. This ensures that the comprehensive impact index parameter of each document not only represents the influence of the static bibliometric indicator, the H5 index of the journal to which the document belongs, but also the influence of dynamic bibliometric indicators such as download count and citation count. Furthermore, a predetermined number of documents with the highest comprehensive impact index are selected as target documents and pushed to users. This ensures that the target documents pushed to users are from diverse sources, have good peer reviews, are highly professional, and have high scholarly recognition. This avoids omitting high-quality documents from less well-known journals during the push process, ensuring the comprehensiveness and objectivity of the pushed documents, thereby improving the user experience.
[0023] The above-mentioned literature information push method integrates various bibliometric indicators of the literature content itself, and adjusts the comprehensive impact index parameters of each literature according to the above-mentioned different dimensions of indicators, so as to improve the coverage of literature sources and thus automatically recommend high-quality academic literature in various disciplines to users.
[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0026] Figure 1This is a schematic diagram illustrating a document information push system according to an exemplary embodiment;
[0027] Figure 2 This is a flowchart illustrating a document information push method according to an exemplary embodiment;
[0028] Figure 3 This is a schematic diagram illustrating a document information push process according to an exemplary embodiment;
[0029] Figure 4 This is a block diagram illustrating a document information push device according to an exemplary embodiment;
[0030] Figure 5 This is a schematic diagram of an electronic device according to an exemplary embodiment. Detailed Implementation
[0031] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0032] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0033] Before providing a detailed introduction to the document information push method provided in the embodiments of this application, let's first briefly introduce the application scenarios and implementation architecture involved in the embodiments of this application.
[0034] First, a brief introduction to the application scenarios involved in this application will be given.
[0035] With the rapid increase in the number of published documents, literature from different journals is stored in a literature search system to enable users to comprehensively search for literature. Given the large number of documents stored in such systems, to facilitate users' quick access to high-quality literature with high interest, a filter is typically applied to high-quality documents from various journals that meet the H5 index criteria, and these H5 index-based high-quality documents are automatically recommended to users. However, the high-quality documents selected using the single-dimensional static indicator of the H5 index usually come from well-known journals, resulting in a relatively singular source of literature. This can easily lead to the omission of high-quality documents from less well-known journals, resulting in a one-sided and incomplete recommendation of literature information and a poor user experience.
[0036] Currently, there are two main problems: how to rank academic literature for multiple purposes and how to improve the coverage of literature sources.
[0037] Specifically, in the field of academic literature, researchers typically need to consider multiple bibliometric metrics, such as download volume, the journal's H5 index, and citation frequency, to assess the quality and impact of a document. However, current ranking based on a single bibliometric metric may not provide a comprehensive evaluation, and differences in data dimensions between metrics can lead to inaccurate rankings. Therefore, the challenge of multi-objective ranking of academic literature is how to integrate multiple bibliometric metrics to obtain more comprehensive and accurate evaluation results, thereby improving literature recommendation methods.
[0038] When extracting the H5 index of journals containing literature within a subject area, improper handling of the weighting parameters of the journal's H5 index may result in some journals having excessively high or low H5 indices, leading to excessively high recall rates for single journal sources and reducing the diversity of literature sources. The key to this problem lies in how to adjust the weighting parameters of the journal's H5 index to improve the comprehensiveness and objectivity of literature recommendations while ensuring coverage of a wider range of literature sources.
[0039] To address the aforementioned issues, this application provides a method for pushing literature information. Based on three different dimensions—download count, citation frequency, and H5 index—it evaluates the comprehensive impact index of each article. This ensures that the comprehensive impact index parameter represents not only the influence of the static bibliometric indicator (H5 index) of the journal to which the article belongs, but also the influence of dynamic bibliometric indicators such as download count and citation count. Furthermore, a predetermined number of articles with the highest comprehensive impact index are selected as target articles and pushed to users. This ensures that the target articles pushed to users are from diverse sources, have good peer reviews, are highly professional, and have high scholarly recognition. This avoids omitting high-quality articles from less well-known journals during the push process, guaranteeing the comprehensiveness and objectivity of the pushed articles, thereby improving the user experience.
[0040] Based on this, the above-mentioned literature information push method integrates various bibliometric indicators of the literature content itself, and adjusts the comprehensive impact index parameters of each literature according to the above-mentioned different dimensions of indicators, so as to improve the coverage of literature sources and thus automatically recommend high-quality academic literature of various disciplines to users.
[0041] Secondly, the implementation architecture involved in this application will be briefly introduced below.
[0042] Figure 1 This is a schematic diagram of a document information push system 10 provided in this disclosure. For example... Figure 1 As shown, the document information push system includes a server 101 and a user terminal 102, and the server 101 and the user terminal 102 can establish a connection through a wired network or a wireless network.
[0043] In some embodiments, a user requests access to the document information push system on the user terminal 102. The server 101 accepts the document access request and simultaneously obtains the target download count, target citation frequency, and target H5 index of each document. Based on the target download count, target citation frequency, and target H5 index of each document, the server determines the comprehensive impact index parameters of each document. The server then identifies the document with the highest comprehensive impact index parameters among a preset number of documents as the target document and pushes the document information of the target document to the user terminal 102, so that the user terminal 102 can display the relevant document information of the target document to the user.
[0044] In other embodiments, a document push system is installed on the user terminal 102, corresponding to a document retrieval application on the terminal. Users can then use the document retrieval application on the corresponding terminal to perform the document information push process described in the above example.
[0045] In some embodiments, server 101 includes or is connected to a database, where various literature information can be stored. User terminals can access the literature information in the database through server 101.
[0046] In other embodiments, server 101 may be a single server, or it may be a server cluster consisting of multiple servers. In some embodiments, the server cluster may also be a distributed cluster. This application does not limit the specific implementation of server 101.
[0047] The aforementioned user terminal 102 can all be understood as a terminal device. A terminal device can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phones, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, etc., that can install and use content community applications (such as Kuaishou). This disclosure does not impose any special restrictions on the specific form of the terminal device. It can interact with the user through one or more methods such as a keyboard, touchpad, touchscreen, remote control, voice interaction, or handwriting devices.
[0048] Optionally, the above Figure 1 In the document information push system shown, server 101 can be connected to at least one user terminal. This application does not limit the number or type of terminal devices.
[0049] The document information push method provided in this application embodiment can be applied to the aforementioned... Figure 1 The document information push system shown in the implementation architecture can also be applied to a user terminal or electronic device. For ease of understanding, the document information push method provided in this application will be described in detail below with reference to the accompanying drawings.
[0050] Figure 2 This is a flowchart illustrating a document information push method according to an exemplary embodiment, such as... Figure 2 As shown, this document information push method includes the following steps.
[0051] S21, determine the target number of downloads for each document, the target citation frequency for each document, and the target H5 index for the journal to which each document belongs.
[0052] In one implementation, to ensure the timeliness of the documents, each document is a document included in journals whose publication time is within a second preset time from the current time.
[0053] In some implementations, the target download count for each document is determined as follows: First, the actual download count of each document within a first preset time period is obtained, i.e., the actual download volume. From the actual download counts corresponding to all documents comprised of these documents, the maximum and minimum actual download counts are selected. Furthermore, for any document within each set of documents, the actual download count is normalized based on the maximum and minimum actual download counts to obtain the target download count for that document.
[0054] In some embodiments, the number of target downloads is determined based on a first model.
[0055] For example, taking the academic data of the past 30 days retrieved from the database of the literature information push system as an example, the process of normalizing the actual download volume of each article in a filtered subject-specific journal is explained as follows.
[0056] Based on the first model characterizing the variable relationship using the following formula (1), the target download count for each document is determined. This can be understood as the target download count for each document being the normalized value of the actual download volume of each document within the subject's document volume. For any given document, the obtained normalized download count is stored as a parameter of that document, namely the target download count, and used as the output of the first model.
[0057]
[0058] Among them, D i P represents the actual number of downloads of the i-th document within the first preset time period; 0 < i ≤ P N , where P N The number of documents in this discipline over the past 30 days; min(D) is the minimum actual download count of a single document among all documents within the first preset time period; max(D) is the maximum actual download count of a single document among all documents within the first preset time period; sigmoid(D) i ) represents the target number of downloads for the i-th document.
[0059] The aforementioned first preset time period can also be referred to as the preset period.
[0060] In other implementations, the target citation frequency for each document is determined as follows: The actual number of citations for each document within a first preset time period is obtained, and the maximum and minimum actual citation counts for each document are determined. For any given document, the actual citation count is normalized based on the maximum and minimum actual citation counts to obtain the target citation count for that document.
[0061] In some embodiments, the number of references to the target is determined based on a second model.
[0062] For example, taking the academic data of the past 30 days retrieved from the database of the literature information push system as an example, the process of normalizing the actual citation counts of each article in a filtered subject-specific journal is explained as follows.
[0063] Based on the second model that characterizes the relationship between variables using the following formula (2), the target citation count for each document is determined. This can be understood as the target download count for each document being the normalized value of the actual citation count of each document within the discipline's document volume. For any given document, the obtained normalized citation count result is stored as a parameter of that document, namely the target citation count, and used as the output of the second model.
[0064]
[0065] Where Ci is the actual citation frequency of the i-th document within the first preset time period, 0 < i ≤ P N , where P N The number of documents in this discipline over the past 30 days is given by sigmoid(C). min(C) represents the minimum actual citation count of a single document within the first preset time period, and max(C) represents the maximum actual citation count of a single document within the first preset time period. i ) represents the target citation frequency of the i-th document.
[0066] In other implementations, the target H5 index of the journal is determined as follows: The actual H5 index of the journal to which each article belongs is obtained, and the normalized absolute median difference is applied to each actual H5 index to obtain the target weights corresponding to each actual H5 index. For any actual H5 index corresponding to the journal to which an article belongs, the difference between 1 and the target weight of that actual H5 index is taken, and this difference is used as the weight difference of that actual H5 index. Then, for any actual H5 index corresponding to the journal to which an article belongs, the weighted sum of the actual H5 index and the corresponding weight difference is determined as the weighted index of each actual H5 index. Based on the frequencies of the largest and smallest weighted indices, each weighted index is normalized to obtain the target H5 index corresponding to the journal to which each article belongs.
[0067] In the above implementation, since the H5 values of different journals within the same discipline may vary significantly, the recommended results may be derived from a single source journal. Therefore, by adopting the absolute median difference algorithm, the interference of excessively high or low H5 values of the journals to which the literature belongs within the discipline is reduced, thus mitigating the diversity of data sources. At the same time, the influence of the H5 parameter indicators of the quality of journals within the discipline is preserved.
[0068] Because the actual H5 index of some well-known top journals is relatively static compared to download volume and citation frequency, and the normalization of H5 indexes for top journals significantly affects recommendation scores, this risk is mitigated by using a weighted approach of "1 - normalized absolute median difference" to reduce such randomness and increase the diversity of sources of excellent journals.
[0069] Specifically, the process of normalizing the absolute median difference of each actual H5 index to determine the target weight is as follows: First, arrange the actual H5 indices of each journal to which each article belongs in descending order. Then, determine the absolute median difference of each actual H5 index based on the actual H5 index in the middle of the order. Then, normalize each absolute median difference based on the largest and smallest absolute median differences to obtain the target weight.
[0070] Furthermore, based on the parity of the total number of journals to which each article belongs, the absolute median differences corresponding to each actual H5 index are determined. Specifically, when the number of journals to which each article belongs is odd, i.e., when the number N of each actual H5 index is odd, the middle actual H5 index is the actual H5 index at position (1+N) / 2; when the number N of journals to which each article belongs is even, i.e., when the number of each actual H5 index is even, the middle actual H5 index is the average of the actual H5 indices at positions (1+N / 2) and (N / 2); the absolute value of the difference between each actual H5 index and the middle actual H5 index is determined as the absolute median difference corresponding to each actual H5 index.
[0071] The aforementioned articles can be selected from professional journals within the discipline. The actual H5 index of the journal to which the article belongs can be retrieved from the H5 database (journal-h5).
[0072] As a specific implementation method, based on the third model, the target H5 index corresponding to each article is determined. The different actual H5 indices of the journals to which each article belongs within the second preset time period are arranged in descending order to obtain the H5 index set.
[0073] For example, S = {H51, H52, ..., H5} N}, S is the set of H5 exponents, H5 i Let N be the H5 index of the journal to which the i-th article belongs, and N be the number of journals to which each article belongs.
[0074] When the number of journals N is odd, the middle actual H5 index is determined based on the following formula (3).
[0075] M 0.5 =H5 (1+N) / 2 Formula (3).
[0076] Among them, M 0.5 This is the median of the H5 index set, i.e., the middle actual H5 index.
[0077] When the number of journals N is even, the middle actual H5 index is determined based on the following formula (4).
[0078]
[0079] Furthermore, based on formula (5), the absolute median difference corresponding to each actual H5 index is determined.
[0080] MAD i =|H5 i -M 0.5 | Formula (5).
[0081] Among them, MAD i This represents the absolute median difference in the actual H5 index of the journal to which the i-th article belongs within the second preset time period. The first preset time and the second preset time can be the same preset time or different preset times; this application does not specifically limit this.
[0082] Furthermore, the target weights corresponding to each actual H5 index are determined based on formula (6).
[0083]
[0084] Where, min(MAD) is the minimum absolute median difference among the journals to which each article belongs within the second preset time period, and max(MAD) is the maximum absolute median difference among the journals to which each article belongs within the second preset time period; Sigmoid(MAD) i ) is the normalized absolute median difference of the i-th document, which is the target weight corresponding to the actual H5 index of the i-th document.
[0085] δ(H5 i )={1-sigmoid(MAD i )}×H5 i Formula (7).
[0086]
[0087] Wherein, min[δ(H5)] is the minimum H5 value of each article within the second preset time period after weighting by journal, that is, the minimum weighted index among the weighted indices of each actual H5 index; max[δ(H5)] is the maximum H5 value of each article within the second preset time period after weighting by journal, that is, the maximum weighted index among the weighted indices of each actual H5 index; sigmoid[δ(H5)] is the maximum H5 value of each article within the second preset time period after weighting by journal, that is, the maximum weighted index among the weighted indices of each actual H5 index; sigmoid[δ(H5)] is the minimum ... i )] represents the target H5 index corresponding to the journal to which the i-th article belongs.
[0088] Based on the third model that characterizes the variable relationships using formulas (3) to (8) above, the target H5 index corresponding to the journal to which each article belongs is determined. This can be understood as the target H5 index corresponding to the journal to which each article belongs being the normalized value of the number of articles in the discipline. Furthermore, for any given article, the 1-normalized result of the actual H5 index is stored as a parameter of that article, namely the target H5 index, and used as the output of the third model.
[0089] S22. Based on the target download count, target citation frequency, and target H5 index of each document, determine the comprehensive impact index parameters of each document.
[0090] Among them, the comprehensive impact index parameters of each article are positively correlated with the target number of downloads, target citation frequency and target H5 index of each article.
[0091] As one implementation method, for any document in each literature, the target download count, target citation frequency, and target H5 index of that document are weighted and summed with the first weight, the second weight, and the third weight, respectively, to obtain the comprehensive influence index parameter of that document. The first weight and the third weight are both greater than the second weight; the sum of the first weight, the second weight, and the third weight is 1.
[0092] In one embodiment, considering the impact on the literature recommendation results and the increase in citation frequency within a first preset time period, the first weight, the second weight, and the third weight are adjusted based on the increase in citation frequency within the literature recommendation results and the first preset time period.
[0093] Optionally, the first and third weights are set to 45%, and the second weight is set to 10%.
[0094] For example, the process of determining the comprehensive impact index parameters of each article using formula (9) is explained below.
[0095] Score = sigmoid[δ(H5)] i )]×α+sigmoid(D i )×β+sigmoid(C i )×γ formula (9).
[0096] Where α, β and γ correspond to the third weight, the first weight and the second weight respectively; Score is the comprehensive influence index parameter.
[0097] S23 identifies the document with the highest preset comprehensive influence index parameter among all documents as the target document, and pushes the document information of the target document.
[0098] The above-mentioned literature information may include one or more of the following structured information: paper title, keywords, subject code, subject name, publication date, source, citation frequency, download frequency, author, chapter title, chapter content, references and cited content, etc.
[0099] As one implementation method, the target download count, target citation frequency, and target H5 index of the journals to which each of the above-mentioned documents belong can all be determined based on document data of different dimensions.
[0100] like Figure 3 As shown, to ensure the accuracy of the data in each document, the data for different dimensions of indicators were first cleaned to determine the actual indicator parameters for each dimension based on the subsequent document data. The actual indicator parameters include actual download count, actual citation frequency, and actual H5 index.
[0101] Specifically, academic journals and various dimensional indicators of publications from the past five years are extracted from academic literature databases. Simultaneously, based on pre-defined themes, data from journals in various disciplines are retrieved, and academic journal articles are categorized by subject, serving as the data source for each article. Furthermore, data cleaning is performed on each extracted article to ensure accuracy and streamline the dataset.
[0102] Furthermore, such as Figure 3 As shown, the actual download count, actual citation frequency, and actual H5 index for each document are obtained from the cleaned literature data. Typically, the actual H5 index of each journal to which each of these documents belongs is stored as an academic bibliometric indicator in this academic literature database.
[0103] Furthermore, such as Figure 3 As shown, the actual number of downloads for each document is input into the first model to obtain the target number of downloads for each document; the actual citation frequency for each document is input into the second model to obtain the target citation frequency for each document; and the actual H5 index for each document is input into the third model to obtain the target H5 index for each document.
[0104] Furthermore, such as Figure 3 As shown, the target download count, target citation frequency, and target H5 index of each document output by the three models are integrated after weighting and parameter tuning to obtain the comprehensive impact index parameters for the comprehensive evaluation of each document.
[0105] For example, the process of determining the comprehensive impact index parameters of each article is explained using the above formula (9).
[0106] Furthermore, such as Figure 3 As shown, the comprehensive impact index parameters of each article are arranged in descending order to obtain the comprehensive sequence.
[0107] Finally, the literature corresponding to the first preset number of comprehensive influence index parameters in the comprehensive sequence is determined as the target literature to be pushed to the user.
[0108] It is understandable that the target literature includes a predetermined number of documents.
[0109] Through the above implementation method, the comprehensive impact index of each article is evaluated based on three different dimensions: download count, citation frequency, and H5 index. This ensures that the comprehensive impact index parameter of each article not only represents the influence of the static bibliometric indicator (H5 index) of the journal to which the article belongs, but also the influence of dynamic bibliometric indicators such as download count and citation count. Furthermore, a predetermined number of articles with the highest comprehensive impact index are selected as target articles and pushed to users. This ensures that the target articles pushed to users are from diverse sources, have good peer reviews, are highly professional, and have high scholarly recognition. This avoids omitting high-quality articles from less well-known journals during the push process, guaranteeing the comprehensiveness and objectivity of the pushed articles, thereby improving the user experience.
[0110] Based on this, the above-mentioned literature information push method integrates various bibliometric indicators of the literature content itself, and adjusts the comprehensive impact index parameters of each literature according to the above-mentioned different dimensions of indicators, so as to improve the coverage of literature sources and thus automatically recommend high-quality academic literature of various disciplines to users.
[0111] In addition, the above-mentioned literature information push method can analyze every single document in the entire collection of documents to be analyzed, ensuring the accuracy of obtaining effective documents. This avoids the problem of low accuracy of obtaining effective documents due to users missing documents in the collection of documents to be analyzed because of objective reasons.
[0112] To achieve the above functions, the document information push device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the algorithmic steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0113] This disclosure also provides an embodiment such as Figure 4 The document information push device 40 shown includes: a determination unit 401, an evaluation unit 402, and a push unit 403.
[0114] The determination unit 401 is configured to determine the target download count, the target citation frequency, and the target H5 index of the journal to which each document belongs for each document; the evaluation unit 402 is configured to determine the comprehensive impact index parameters of each document based on the target download count, target citation frequency, and target H5 index; the comprehensive impact index parameters of each document are positively correlated with the target download count, target citation frequency, and target H5 index of each document; the push unit 403 is configured to determine the document with the largest number of comprehensive impact index parameters among the documents as the target document, and push the document information of the target document.
[0115] In one possible implementation, the determining unit 401 is specifically configured to: obtain the actual number of downloads of each document within a first preset time period; determine the maximum and minimum actual download counts among the actual download counts of each document; and normalize the actual download counts of each document based on the maximum and minimum actual download counts to obtain the target download counts corresponding to each document.
[0116] In another possible implementation, the determining unit 401 is specifically configured to: obtain the actual number of times each document is cited within a first preset time period; determine the maximum and minimum actual number of citations among the actual citations of each document; and normalize the actual citations of each document based on the maximum and minimum actual citations to obtain the target citation counts corresponding to each document.
[0117] In another possible implementation, the determining unit 401 is specifically configured to: obtain the actual H5 index of the journal to which each article belongs; perform normalized absolute median difference processing on each actual H5 index to obtain each target weight corresponding to each actual H5 index; weight the difference between each actual H5 index and each weight to determine each weighted index of each actual H5 index; for any target weight, any weight difference represents the difference between 1 and any target weight; and perform normalization processing on each weighted index based on the frequency of the maximum weighted index and the minimum weighted index among each weighted index to obtain each target H5 index corresponding to the journal to which each article belongs.
[0118] In another possible implementation, the determining unit 401 is further specifically configured to: arrange the actual H5 indices of the journals to which each article belongs in descending order; determine the absolute median difference corresponding to each actual H5 index based on the actual H5 index located in the middle of the order; and normalize each absolute median difference based on the largest and smallest absolute median differences to obtain each target weight.
[0119] In another possible implementation, the determining unit 401 is further specifically configured as follows: when the number of journals to which each article belongs is odd, that is, when the number N of each actual H5 index is odd, the middle actual H5 index is the actual H5 index at position (1+N) / 2; when the number N of journals to which each article belongs is even, that is, when the number of each actual H5 index is even, the middle actual H5 index is the average of the actual H5 indices at positions (1+N / 2) and (N / 2); and the absolute value of the difference between each actual H5 index and the middle actual H5 index is determined as the absolute median difference corresponding to each actual H5 index.
[0120] In another possible implementation, each document refers to documents published within a second preset time from the current time in journals; the evaluation unit 402 is specifically configured to: for any document in each document, the target download count, target citation frequency, and target H5 index of any document are weighted and summed with the first weight, the second weight, and the third weight, respectively, to obtain the comprehensive impact index parameter of any document; the first weight and the third weight are both greater than the second weight; the sum of the first weight, the second weight, and the third weight is 1.
[0121] Regarding the apparatus in the above embodiments, the specific manner in which each unit module performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0122] Figure 5 This is a schematic diagram of an online consultation device provided in this application. Figure 5 The online consultation device 50 may include at least one processor 501 and a memory 503 for storing processor-executable instructions. The processor 501 is configured to execute the instructions in the memory 503 to implement the literature information push method in the following embodiments.
[0123] In addition, the online consultation device 50 may also include a communication bus 502, at least one communication interface 504, an input device 506, and an output device 505.
[0124] Processor 501 may be a processor (central processing unit, CPU), microprocessor unit, ASIC, or one or more integrated circuits for controlling the execution of programs according to the present application.
[0125] The communication bus 502 may include a path for transmitting information between the aforementioned components.
[0126] Communication interface 504 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0127] Input device 506 is used to receive input signals and output device 505 is used to output signals.
[0128] Memory 503 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory may exist independently and be connected to the processing unit via a bus. Memory may also be integrated with the processing unit.
[0129] The memory 503 stores instructions for executing the scheme of this application, and the processor 501 controls the execution. The processor 501 executes the instructions stored in the memory 503 to implement the functions of the method of this application.
[0130] In a specific implementation, as one example, the processor 501 may include one or more CPUs, for example... Figure 5 CPU0 and CPU1 in the CPU.
[0131] In a specific implementation, as one example, the online consultation device 50 may include multiple processors, such as... Figure 5 Processors 501 and 507 are shown in the diagram. Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0132] This online consultation device, such as Figure 5 The diagram shows a processor 501 and a memory 503 for storing executable instructions of the processor 501. The processor 501 is configured to execute the executable instructions to implement the document information push method as described in any of the possible embodiments above. Since the same technical effects can be achieved, further details are omitted here to avoid repetition.
[0133] This application also provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of a document information push device or electronic device, the document information push device or electronic device is able to perform the document information push method as described in any of the possible embodiments above. And it can achieve the same technical effect; to avoid repetition, it will not be described again here.
[0134] This application also provides a computer program product, including a computer program or instructions, which are executed by a processor as described in any of the possible implementations of the document information push method above. This achieves the same technical effect, and to avoid repetition, it will not be described again here.
[0135] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0136] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for pushing literature information, characterized in that, The method includes: Determine the target number of downloads for each article, the target citation frequency for each article, and the target H5 index for the journal to which each article belongs; Based on the target download count, target citation frequency, and target H5 index of each document, the comprehensive impact index parameters of each document are determined; the comprehensive impact index parameters of each document are positively correlated with the target download count, target citation frequency, and target H5 index of each document. The document with the largest comprehensive impact index parameter among the preset number of documents is identified as the target document, and the document information of the target document is pushed. The determination of the target H5 index of the journal to which each article belongs includes: Obtain the actual H5 index of the journals to which each of the aforementioned articles belongs; The normalized absolute median difference of each actual H5 index is processed to obtain the target weights corresponding to each actual H5 index. The weighted sum of each actual H5 index and each weight difference is determined as each weighted index of the actual H5 index; for any target weight, any weight difference represents the difference between 1 and any target weight. Based on the frequency of the largest and smallest weighted indices among the weighted indices, the weighted indices are normalized to obtain the target H5 indices corresponding to the journals to which the articles belong. The step of normalizing the absolute median difference of each actual H5 index to obtain the target weights corresponding to each actual H5 index includes: arranging the actual H5 indices of the journals to which each article belongs in descending order; determining the absolute median difference corresponding to each actual H5 index based on the actual H5 index located in the middle of the order; and normalizing the absolute median difference based on the largest and smallest absolute median differences to obtain the target weights.
2. The method according to claim 1, characterized in that, Determining the target number of downloads for each document includes: Obtain the actual number of downloads for each document within a first preset time period; Determine the maximum and minimum actual download counts for each document. Based on the maximum and minimum actual download counts, the actual download counts of each document are normalized to obtain the target download counts for each document.
3. The method according to claim 1, characterized in that, Determining the target citation frequency for each document includes: Obtain the actual number of citations for each of the aforementioned documents within a first preset time period; Determine the maximum and minimum actual citation counts for each of the aforementioned documents; Based on the maximum and minimum actual citation counts, the actual citation counts of each document are normalized to obtain the target citation frequencies corresponding to each document.
4. The method according to claim 1, characterized in that, The step of determining the absolute median difference corresponding to each actual H5 index based on the actual H5 index located in the middle of the position includes: When the number of journals N to which each of the articles belongs is odd, the middle actual H5 index is the actual H5 index of the (1+N) / 2nd position. When the number of journals N to which each article belongs is even, the middle actual H5 index is the average of the actual H5 indices of the (1+N / 2)th and (N / 2)th positions. The absolute value of the difference between each actual H5 index and the middle actual H5 index is determined as the absolute median difference corresponding to each actual H5 index.
5. The method according to any one of claims 1 to 4, characterized in that, The aforementioned documents are those published within a second preset time interval from the current time; the determination of the comprehensive impact index parameters for each document based on its target download count, target citation frequency, and target H5 index includes: For any one of the aforementioned documents, the target download count, target citation frequency, and target H5 index of that document are weighted and summed with the first weight, the second weight, and the third weight, respectively, to obtain the comprehensive influence index parameter of that document; the first weight and the third weight are both greater than the second weight; the sum of the first weight, the second weight, and the third weight is 1.
6. A document information push device, characterized in that, The device includes: The unit is configured to determine the target number of downloads for each document, the target citation frequency for each document, and the target H5 index for the journal to which each document belongs. The evaluation unit is configured to determine the comprehensive impact index parameters of each document based on the target download count, target citation frequency, and target H5 index of each document; the comprehensive impact index parameters of each document are positively correlated with the target download count, target citation frequency, and target H5 index of each document; The push unit is configured to identify the document with the largest comprehensive impact index parameter among the comprehensive impact index parameters of the documents as the target document, and to push the document information of the target document. The determining unit is specifically configured to: obtain the actual H5 index of the journal to which each of the articles belongs; The normalized absolute median difference of each actual H5 index is processed to obtain the target weights corresponding to each actual H5 index. The weighted sum of each actual H5 index and each weight difference is determined as each weighted index of the actual H5 index; for any target weight, any weight difference represents the difference between 1 and any target weight. Based on the frequency of the largest and smallest weighted indices among the weighted indices, the weighted indices are normalized to obtain the target H5 indices corresponding to the journals to which the articles belong. Specifically, the determining unit is configured to: arrange the actual H5 indices of the journals to which each article belongs in descending order; determine the absolute median differences corresponding to each actual H5 index based on the actual H5 index located in the middle of the order; and normalize each absolute median difference based on the largest and smallest absolute median differences to obtain the target weights.
7. A document retrieval system, characterized in that, include: A processor and a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to implement the document information push method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the document information push method as described in any one of claims 1-5.
Citation Information
Patent Citations
Literature evaluation method for retrieval sorting, storage medium and terminal
CN115686432A