A prediction-based method for recommending academic book publishing topics

By using a prediction method based on user behavior data, and employing a self-attention layer and a Transformer model to predict changes in the popularity of academic book topics, this approach solves the problems of topic lag and data sparsity in existing technologies, and achieves more accurate topic recommendations.

CN116431896BActive Publication Date: 2025-12-02BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310145873.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2025-12-02
Estimated Expiration
2043-02-22

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict trending topics for academic books, and existing recommendation systems are insufficient in responding to reader interests, with sparse and lagging data leading to poor reliability in topic selection.

Method used

We employ a prediction method based on user behavior data, enhance the vector representation of topic selection data through a self-attention layer, and use the Transformer model for time series prediction to generate future changes in topic popularity.

Benefits of technology

It improves the timeliness and accuracy of academic book selection and recommendations, enhances the ability to respond to readers' interests, and improves the reliability of selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116431896B_ABST
    Figure CN116431896B_ABST
Patent Text Reader

Abstract

This invention discloses a prediction-based method for recommending academic book publishing topics. The method involves preprocessing user behavior records to extract usable fields and remove symbols and null values. A certain amount of topic data is predefined, consisting of keywords from frequently cited papers on a website. The user behavior data and predefined topic data, after semantic enhancement through a self-attention layer using a text vectorization model, are then compared with all user behavior data in terms of text similarity. The generated data is input into a layer-norm layer to calculate the mean and the corresponding value for each topic data point. Finally, all means are sorted to obtain recommended topic data. The time series corresponding to the obtained topic data is then predicted to determine the future popularity changes of the topic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of book publishing topic selection technology. In particular, it relates to a predictive method for recommending academic book topics, primarily aimed at publishing houses to assist staff in making academic book publishing decisions. Background Technology

[0002] With the increasing abundance of information available to people today, this vast and disorganized information presents significant challenges for publishers in selecting topics. Publishers must accurately select popular topics that resonate with readers' interests and proactively consider the future prospects of each topic.

[0003] Traditional methods of selecting book topics involve several approaches: 1. Market research; 2. Analyzing social and political developments; 3. Discovering topics from unsolicited submissions; and 4. Researching bibliographical analysis reports. However, all these methods rely on editors' accurate understanding of social developments and market changes. Furthermore, academic books also face challenges such as a smaller audience, difficulty in conducting market research, and rapid changes in subject matter.

[0004] The existing topic recommendation systems mostly rely on online public opinion data and the publishers' own book sales data. The data content is too complex and the academic information is too sparse, resulting in poor reliability of the academic book publishing topics obtained.

[0005] The invention patent application, with publication number CN114647675A and titled "A Book Publishing Topic Selection System Based on Academic Big Data," discloses the following modules: a user setting module for configuring the user's publishing institution, usage rights, and usage time limits; a subject setting module for users to personalize their subject settings; a topic recommendation module for recommending hot topics based on the user's subject settings; a topic setting module where users set topic selections, and after submission, the system generates a topic evaluation report and author recommendations; a topic evaluation module for evaluating user-selected or set topics and providing a topic evaluation report; an author recommendation module for extracting data from a core author database to recommend core academic authors for user-selected or set topics; a hot topic mining module for using data mining-based academic hot topic analysis methods to mine hot topics from massive amounts of academic literature; and an author mining module for using data statistical techniques to select core authors for the chosen topic from massive amounts of academic literature.

[0006] The above technology has the following problems:

[0007] First, while data related to academic literature is used, this data comes from academic literature databases. Although it can reflect current academic research trends to some extent, it is less effective in reflecting readers' interests. Furthermore, the publication cycle for papers is long, ranging from six months to a year from writing to publication. Therefore, published articles are lagging behind current research. Consequently, recommendation methods based on this data are less effective at capturing current hot topics.

[0008] Second: Existing technologies utilize bibliographic information statistical analysis tools to statistically analyze academic literature data to identify hot topics. While bibliographic information itself can summarize the main methods and content of a document to the greatest extent, it lacks a certain degree of overview of the overall field to which the article belongs. Therefore, methods that only perform statistical analysis on bibliographic information have certain informational gaps. Summary of the Invention

[0009] To address the aforementioned technical problems, the purpose of this invention is to provide a prediction-based method for recommending academic book publishing topics, which can solve the problem of selecting academic book publishing topics.

[0010] One technical solution adopted in this invention is:

[0011] A prediction-based method for recommending academic book publishing topics includes: user behavior data sourced from an academic paper system; predefined topic selection data conforming to certain logic; using a self-attention layer to enhance the semantic representation of the user behavior data and topic selection data vectors; calculating the average of text similarity data generated from the user behavior data and predefined topic selection data using a layer-norm layer, and then sorting the results to obtain recommended topic selection data; and using a transformer model to predict the popularity changes of the topic selection data.

[0012] (1) Data preprocessing

[0013] Preprocess existing user behavior data to extract content information from users' browsing and download records, including key information such as title, keywords, and access time.

[0014] (2) Hotspot discovery

[0015] The data is input into the hotspot discovery model as document = {t, k, d}, where k = (k1, k2, ...). Word vectors are used to represent the content title t and keywords k, and k and t are fused into a single vector using a self-attention mechanism. A predefined hotspot category is defined, and each category corresponds to a predefined topic selection data table e = {h, ​​r}, where h is the topic and r = (r1, r2, ...) are the secondary words contained in the topic. Vectors are used to represent h and r, and the resulting vectors are fused into e using a self-attention mechanism. Subsequently, the vector distance between the topic and e is calculated, and these distances are summed to form the hotspot value of the topic data.

[0016] (3) Hotspot prediction

[0017] Hotspot prediction is performed using a hotspot discovery algorithm, which generates sequential data of topic selection data, i.e., daily hotspot data (hot, index, time), where index is the popularity value, time is the time period, and h is the hotspot. By capturing the temporal (within a hotspot) and spatial (between hotspots) relationships, the future trend of hotspots can be predicted.

[0018] The specific steps are as follows:

[0019] (1) Data preprocessing

[0020] The input data is user behavior data;

[0021] The input features of the model are generally divided into three main categories: basic user information, user behavior sequences, and candidate items. User paper browsing history includes the user's unique identifier, user type, document ID, document title, browsing time, province where the user operated, document keywords, document subject number, document author, author's affiliation, and behavior type.

[0022] The user behavior data is structured as follows: USER_ID is the unique identifier for the user; USER_TYPE is the user type, categorized as individual or institutional group account; ARTICLE_ID is the document ID; DATATIME is the user's browsing time; ARTICLE_TITLE is the document title; KEYWORDS are the document keywords, consisting of multiple keywords separated by semicolons; AUTHOR is the document author, with multiple authors separated by semicolons; UNIT is the document author's affiliation, with multiple affiliations separated by semicolons; TYPE is the behavior type, with 2 indicating browsing; CLASSCODE is the document's subject classification number (Chinese Library Classification number); and PROVINCE is the province where the user's action occurred. The user behavior data is processed to remove special characters and spaces.

[0023] Extract the two fields, article_title and keywords, from the overall data to form a dictionary. The key in the dictionary is the content of article_title, and the data in the key is a list of the corresponding keywords.

[0024] (2) Hotspot discovery

[0025] Given a dataset D, each data point p i ={t, k, d} indicates that there are users who have commented on paper p. i A click operation was performed, where t is the title of the paper, and k = (k1, k2, ..., kk). n ) represents the keywords of the paper, and n represents the number of keywords. d represents the date of the user's action. Given dataset E, each data point e = {h, ​​r} represents a topic selection direction, where h is the topic, and r = (r1, r2, ..., r...). n The keywords are those included in the selected topics. The topic directions in dataset E are predefined. The source of the topic directions and the keywords included in the topics are the keywords from highly cited literature, and they are categorized into a certain topic direction. The same word can belong to multiple topic directions.

[0026] Use a pre-trained BERT model to obtain vector representations of the title, keywords, and words in the topic selection, t' = BERT(t)(k1, k2, ..., kk). n ) = BERT(ki), where t is the vector representation of the title, and ki is the vector representation of the i-th word in the keyword field. h' = BERT(h)(r1, r2, ..., r n ) = BERT(ri) where h is the vector representation of the topic and ri is the vector representation of the i-th word of the second-level vocabulary in the topic.

[0027] First, calculate the distance between the title vector and each keyword vector, and then normalize the distance to obtain the importance weight αi for each keyword:

[0028] si = cosin(t′, ki)

[0029] αi=softmax(si, s1, s2...sn)

[0030] The weighted vector representation of the title is obtained as follows:

[0031] t′=∑aiki

[0032] Finally, the popularity value is calculated by dividing the data into time granularities. Each paper in each time granularity is represented by a vector, and each topic direction is represented by a vector. The two constitute two two-dimensional tensors. The cosine similarity of the one-dimensional tensors in these two two-dimensional tensors is calculated pairwise to generate multiple values. The popularity value of each topic direction can be represented by these multiple values.

[0033] The norm of a vector is a measure of its size. The absolute value of a real number, the modulus of a complex number, and the length of a vector in three-dimensional space are all prototypes of the abstract concept of norm. Several commonly used vector norms are as follows:

[0034] (a) Zero norm, which is the number of non-zero elements in vector a;

[0035] (b)1 norm, which is the sum of the absolute values ​​of all elements in vector a;

[0036] (c)∞ norm, which is the maximum absolute value of the elements of vector a;

[0037] (d)-∞ norm, which is the minimum absolute value of the elements of vector a;

[0038] (e)p norm;

[0039] The popularity value of each topic direction is obtained by taking the L2 norm of multiple values ​​corresponding to each topic direction.

[0040] (3) Hotspot prediction

[0041] The generated hotspot data is constructed into a sequence of triples, where each triple represents the topic selection direction, the hotspot value, and the time. A Transformer model is then used for time series forecasting to predict the popularity changes of each topic selection direction over a future period. The forecast granularity can be daily, weekly, or monthly.

[0042] The technical effects of this invention are as follows: using user behavior data as input data enhances the real-time nature of generating topic selection data and performs time-series prediction to generate topics for a future period of time, making them more valuable for recommendations; by fusing document captions and keyword features through self-attention, it can better express document features and improve the accuracy of recommendations. Attached Figure Description

[0043] Figure 1 This is the overall framework of the method of the present invention;

[0044] Figure 2a These are the data field names in the example;

[0045] Figure 2b This is a data sample instance of the embodiment;

[0046] Figure 3 This is a schematic diagram of the self-attention mechanism of the present invention. Detailed Implementation

[0047] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.

[0048] Please see Figure 1 Specific embodiments of the present invention include:

[0049] (1) Data preprocessing

[0050] The input data is user behavior data.

[0051] The input features of the model are generally divided into three main categories: basic user information, user behavior sequences, and candidate items. User paper browsing history includes the user's unique identifier, user type, document ID, document title, browsing time, province where the user operated, document keywords, document subject number, document author, author's affiliation, and behavior type.

[0052] The user behavior data is processed as follows: USER_ID is the unique identifier for the user; USER_TYPE is the user type, categorized as individual or institutional group account; ARTICLE_ID is the document ID; DATATIME is the user's browsing time; ARTICLE_TITLE is the document title; KEYWORDS are the document keywords, consisting of multiple keywords separated by semicolons; AUTHOR is the document author, with multiple authors separated by semicolons; UNIT is the document author's affiliation, with multiple affiliations separated by semicolons; TYPE is the behavior type, with 2 indicating browsing; CLASSCODE is the document subject number, which is the Chinese Library Classification number; and PROVINCE is the province where the user's action occurred. Special characters and spaces are removed from the user behavior data. The `article_title` and `keywords` fields are extracted from the overall data to form a dictionary. The dictionary's key is the content of `article_title`, and the data within each key is a list of corresponding keywords. For example... Figure 2a and Figure 2b As shown.

[0053] (2) Hotspot discovery

[0054] Given a dataset D, each data point p i t, k, d indicate that there are users who have commented on paper p. i A click operation was performed, where t is the title of the paper, and k = (k1, k2, ..., kk). n) represents the keywords of the paper, and n represents the number of keywords. D represents the date of the user's behavior. Given a dataset E, each data point h, r represents a topic selection direction, h is the topic selection, and r = (r1, r2, ..., r2) n The keywords listed are secondary vocabulary included in the selected topics. A portion of the keywords are manually annotated in advance. These keywords are sourced from highly cited literature and categorized into a specific topic area. The same keyword may belong to multiple topic areas.

[0055] Use a pre-trained BERT model to obtain vector representations of the title, keywords, and words in the topic selection, t' = BERT(t)(k1, k2, ..., kk). n ) = BERT(ki), where t is the vector representation of the title, and ki is the vector representation of the i-th word in the keyword field. h' = BERT(h)(r1, r2, ..., r n ) = BERT(ri) where h is the vector representation of the topic and ri is the vector representation of the i-th word of the second-level vocabulary in the topic.

[0056] like Figure 3 As shown, the self-attention mechanism assigns greater weight to core words within the keyword list, enabling the text vector to contain more semantic information. First, the distance between the title vector and each keyword vector is calculated, and then the distance is normalized to obtain the importance weight αi for each keyword:

[0057] si = cosin(t′, ki)

[0058] αi=softmax(si, s1, s2...sn)

[0059] The weighted vector representation of the title is obtained as follows:

[0060] t′=∑αiki

[0061] Finally, the popularity value is calculated. Each paper is represented by a vector, and each topic direction is represented by a vector. The two constitute two two-dimensional tensors. The cosine similarity of the one-dimensional tensors in these two two-dimensional tensors is calculated pairwise, and finally multiple values ​​are generated. The popularity value of each topic direction can be represented by these multiple values.

[0062] The norm of a vector is a measure of its size. The absolute value of a real number, the modulus of a complex number, and the length of a vector in three-dimensional space are all prototypes of the abstract concept of norm. Several commonly used vector norms are as follows:

[0063] (a) Zero norm, which is the number of non-zero elements in vector a;

[0064] (b)1 norm, which is the sum of the absolute values ​​of all elements in vector a;

[0065] (c)∞ norm, which is the maximum absolute value of the elements of vector a;

[0066] (d)-∞ norm, which is the minimum absolute value of the elements of vector a;

[0067] (e)p norm.

[0068] The popularity value of each topic direction is obtained by taking the L2 norm of multiple values ​​corresponding to each topic direction.

[0069] (3) Hotspot prediction

[0070] The generated hotspot data is constructed into a sequence of triples, where each triple represents the topic selection direction, the hotspot value, and the time. A Transformer model is then used for time series forecasting to predict the popularity changes of each topic selection direction over a future period. The forecast granularity can be daily, weekly, or monthly.

[0071] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A prediction-based method for recommending academic book publishing topics, characterized in that, include: The data source is user behavior data from the academic paper system; Predefined topic selection data conforming to certain logic is used. A self-attention layer is used to enhance the semantic representation of user behavior data and topic selection data vectors. The text similarity data generated from user behavior data and predefined topic selection data is averaged through a layernorm layer and then sorted to obtain recommended topic selection data. The generated time series is predicted using a transformer model to obtain the popularity changes of topic selection data. Specifically, the following steps are included: (1) Data preprocessing Preprocess existing user behavior data to extract content information from users' browsing and download behavior records, including title, keywords, and access time information; (2) Hotspot discovery The data input hotspot discovery model is input as document = {t, k, d}, where k = (k1, k2, ..., kk). n For each content title t and keyword k, word vector representations are performed. k and t are then fused into a single vector using a self-attention mechanism. Predefined hot topic categories are established, with each category corresponding to a predefined topic selection data table e = {h, ​​r}, where h represents the topics and r = (r1, r2, ..., r...). n The second-level vocabulary included in the topic is represented by vectors h and r. The represented vectors are then fused to represent e through a self-attention mechanism. Subsequently, the vector distance between the topic and e is calculated and the distances are superimposed to form the popularity value of the topic data. In step (2), assuming a given dataset D, each data p i ={t, k, d} indicates that there are users who have commented on paper p. i A click operation was performed, where t is the title of the paper, and k = (k1, k2, ..., kk). n ) represents the keywords of the paper, n represents the number of keywords; d represents the date of user behavior; given dataset E, each data point e = {h, ​​r} represents a topic selection direction, h is the topic selection, r = (r1, r2, ..., r... n ) represents the keywords contained in the topic selection; the topic selection directions in dataset E are predefined. The source of the topic selection directions and the keywords contained in the topics are the keywords of literature with high citation counts, and they are classified into a certain topic selection direction. The same word can belong to multiple topic selection directions. A pre-trained BERT model is used to obtain vector representations of the title, keywords, and words in the topic selection, t' = BERT(t)(k1,k2,......k n ) = BERT(ki) where t is the vector representation of the title, and ki is the vector representation of the i-th word in the keyword field; h' = BERT(h)(r1, r2, ..., r n ) = BERT(ri) where h is the vector representation of the topic selection, and ri is the vector representation of the i-th word in the second-level vocabulary of the topic selection; First, calculate the distance between the title vector and each keyword vector, and then normalize the distance to obtain the importance weight αi for each keyword: si = cosin(t′, ki) αi=softmax(si, s1, s2...sn) The weighted vector representation of the title is obtained as follows: t′=∑αiki Finally, the popularity value is calculated. Each paper is represented by a vector, and each topic direction is represented by a vector. The two constitute two two-dimensional tensors. The cosine similarity of the one-dimensional tensors in these two two-dimensional tensors is calculated pairwise to generate multiple values. The popularity value of each topic direction can be represented by these multiple values. For each topic selection direction, take the L2 norm of multiple values ​​to obtain the popularity value of each topic selection direction; (3) Hotspot prediction Hotspot prediction is performed by generating sequence data of topic selection data based on the hotspot discovery algorithm, namely daily hotspot data (hot, index, time), where index is the popularity value, time is the time, and h is the hotspot; by capturing the relationship between the temporal hotspots and the spatial hotspots, the future trend of the hotspots is predicted.

2. The method for recommending academic book publishing topics based on prediction according to claim 1, characterized in that, In step (1), the input features of the model are divided into three main categories: basic user information, user behavior sequence, and candidate items; Process user behavior data to remove special characters and spaces; Extract the two fields, article_title and keywords, from the overall data to form a dictionary. The key in the dictionary is the content of article_title, and the data in the key is a list of the corresponding keywords.

3. The method for recommending academic book publishing topics based on prediction according to claim 1, characterized in that, In step (3), the generated hot data is constructed into a sequence of triples, where the triples are the topic selection direction, the hot value, and the time. Using the Transformer model for time series forecasting, we can predict the popularity changes of each topic over a future period. The forecast time granularity can be divided into days, weeks, and months.

Citation Information

Patent Citations

  • Book publishing and topic selecting system based on academic big data

    CN114647675A

  • Academic accurate recommendation-oriented heterogeneous scientific research information integration method and system

    WO2023272748A1