Smart reading recommendation and cooperation analysis method

By constructing research interest vectors and cooperation fit, combining academic social networks and citation relationships, an intelligent cooperation framework is generated, and the problem of being unable to capture the dynamic changes in scholars' research interests in the existing technology is solved, and efficient interdisciplinary cooperation matching and resource optimization are achieved.

CN120296248AInactive Publication Date: 2025-07-11BEIJING DUSHANGGAOLOU CULTURAL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510359557.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing academic cooperation recommendation system cannot accurately capture the dynamic changes in scholars' research interests, and relies too much on static field labels and keywords, resulting in the inability to maximize cooperation efficiency, and lack real-time feedback mechanisms, so it cannot be adjusted based on scholars' latest research results and academic dynamics.

Method used

By collecting scholars' historical research data, destop words and word segmentation process and converting them into vector representations, constructing research interest vectors, using long and short-term memory networks to predict interest, combining academic social networks and citation relationships to calculate cooperation fit, generating intelligent cooperation frameworks, and personalized reading recommendations and interdisciplinary cooperation matching.

Benefits of technology

It has achieved accurate identification of scholars with high fit, optimized resource allocation, improved cooperation efficiency, provided efficient cooperation suggestions and task division, adapted to the dynamic changes of scholars' research, and improved the execution efficiency of cooperative projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296248A_ABST
    Figure CN120296248A_ABST
Patent Text Reader

Abstract

The invention relates to the field of academic cooperation recommendation and task allocation, and discloses a smart reading recommendation and cooperation analysis method, which comprises the following steps: academic data collection and preprocessing: collecting historical research data of scholars, the research data comprising papers, academic conference records, scientific research fund application forms and academic social platform data, performing stop word removal and word segmentation processing on the research data, and converting text content into vector representation; and research interest modeling: based on the historical research data of the scholars, extracting research themes, constructing scholar research interest vectors, and carrying out weighted calculation on the research interest vectors according to semantic features and introduced conditions of papers published by the scholars. According to the invention, based on scholar research interest vector, reference relationship and social network analysis, in combination with an intelligent task allocation mechanism, high-integrating-degree cooperative partners are accurately matched, task division is optimized, and the precision and efficiency of interdisciplinary cooperation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of academic cooperation recommendation and task assignment, and specifically to a method for intelligent reading recommendation and cooperation analysis. Background Art

[0002] In the process of academic research and technological innovation, interdisciplinary cooperation has become increasingly important. As research problems become more complex, the knowledge of a single discipline is often insufficient to solve practical problems. Therefore, cooperation among scholars is particularly crucial. Through precise cooperation matching, scholars can quickly find suitable partners based on their research directions, expertise, and cooperation potential. However, most existing academic cooperation recommendation systems rely on superficial research fields or keyword matching and cannot deeply explore the true interests and cooperation potential of scholars.

[0003] In the prior art, academic cooperation mostly matches scholars through simple keyword matching or domain label matching. By analyzing the research fields and interests of scholars, the recommendation system can provide cooperation suggestions for scholars. The advantage of such technical solutions is that they can quickly match scholars in similar fields, help scholars find potential partners, and thus improve the efficiency of cooperation. In addition, some existing systems also consider the citation relationships of scholars, evaluate the academic influence of scholars and their connections with other scholars, and thereby enhance the reliability of cooperation recommendations.

[0004] However, the prior art has obvious deficiencies in several aspects. First, traditional cooperation matching methods rely too much on static domain labels or keywords, which results in their inability to capture the dynamic changes in scholars' interests. Scholars' research interests are constantly evolving, and the prior art cannot flexibly respond to this change, often missing potential interdisciplinary cooperation opportunities. Second, the existing systems are too rough in the allocation of cooperation tasks. They often simply divide the work based on the research backgrounds of scholars, lacking in-depth analysis of scholars' interests and abilities, and thus unable to maximize the cooperation efficiency. In addition, the recommendation systems in the prior art mostly rely on fixed models and lack a real-time feedback mechanism, making it difficult to adjust according to the latest research results and academic trends of scholars, resulting in the lack of adaptability and flexibility of the recommendation results. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method for intelligent reading recommendation and cooperation analysis, which solves the problems that the academic cooperation recommendation system in the prior art cannot accurately capture the dynamic changes in scholars' research interests and that cooperation matching relies too much on static keywords and domain labels.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for intelligent reading recommendation and cooperation analysis, comprising the following steps: Academic data collection and preprocessing: Collect historical research data of scholars, where the research data includes papers, academic conference records, scientific research fund applications, and academic social platform data. Perform stop word removal and word segmentation on the research data, and convert the text content into vector representation; Research interest modeling: Based on the historical research data of scholars, extract research topics and construct scholars' research interest vectors, where the research interest vectors are weighted and calculated according to the semantic features and citation situations of the papers published by scholars; Research interest evolution prediction: Based on time series modeling, analyze the historical change trends of scholars' research data interests, and use long short-term memory networks for prediction to determine the future research directions that scholars will focus on; Personalized reading recommendation: Based on the research data interest evolution prediction results, calculate the similarity between the scholars' research data interest vectors and the literature vectors in the academic literature database, and screen and recommend the most relevant literature; Interdisciplinary cooperation matching: Based on the scholars' research data interest vectors, citation relationships, and academic social network relationships, calculate the cooperation fit between different scholars, and screen potential cooperation partners according to the fit; Intelligent cooperation framework generation: After determining high-fit cooperation partners, generate cooperation suggestions, recommend research directions, and perform task division according to the research expertise of scholars to form a preliminary cooperation framework Preferably, the research topics include: Text feature extraction: Based on the historical research data of scholars, use natural language processing methods for text feature extraction to identify paper titles, abstracts, keywords, and main research contents; Theme clustering analysis: Use clustering algorithms to perform clustering analysis on text features, generate multiple research theme categories, and calculate the proportion of each research theme in the scholars' research interests; Weight calculation: Assign weights to each research theme, where the weights are calculated based on the citation times, publication time, and research field influence of the papers to highlight the core research directions.

[0007] Preferably, the time series modeling includes: Time window division: According to the time when scholars publish papers, group the research interest vectors in chronological order to form a time series data set; Feature extraction: Extract features from the research interest vectors of different time windows to identify the change trends of research topics over time; Sequence modeling: Use recurrent neural networks or variational autoencoders for sequence modeling to establish a prediction model for the change of research interests over time.

[0008] Preferably, the historical change trends include: Research topic changes, analyze the changes in scholars' research topics over different time periods, and identify the evolution trends of research focuses; Interdisciplinary trends, detect whether scholars' research directions gradually expand to other disciplinary fields, and analyze the degree of interdisciplinary research; Research intensity fluctuations, count the number and influence of published papers in different time periods, and judge the changes in research activity.

[0009] Preferably, the long short-term memory network includes: Input layer construction, take the research interest vector as the input and perform normalization processing; Hidden layer calculation, use long short-term memory units to process time series data and remember the evolution trends of research interests over a long time span; Output layer optimization, perform non-linear transformation on the prediction results to ensure the accuracy of the research interest prediction results, and optimize the prediction model through the loss function.

[0010] Preferably, the prediction results include: Future research topics, based on scholars' historical research data, predict the research topics that will be focused on in the future; Potential interdisciplinary fields, identify the interdisciplinary research directions involved by scholars and provide corresponding research suggestions; Research development trends, combined with the global academic development dynamics, predict the matching degree between scholars' research directions and international frontier trends.

[0011] Preferably, the research data interest vector includes: Topic vector, according to the modeling results of scholars' research interests, transform the research topics into vector representations; Time decay factor, assign different weights to research data in different time periods, with higher weights for recent research interest vectors to reflect the latest research interests; Multimodal fusion, combine various research forms such as text, charts, and experimental data to construct a comprehensive interest vector.

[0012] Preferably, the literature vector includes: Text feature vector, based on the title, abstract, keywords, and full text content of academic papers, construct a text feature vector; Citation network vector, analyze the citation relationships of papers, construct a citation network, and calculate the influence of papers in the network; Author influence vector, combine the historical publication records and academic influence of paper authors, calculate the contribution degree of authors in specific research fields to optimize the recommendation results.

[0013] Preferably, the interdisciplinary cooperation matching further includes: Cooperation fitness calculation: Based on the research interest vectors, citation relationships, and academic social networks of scholars, calculate the cooperation fitness between scholars; Potential cooperation partner screening: Sort according to the cooperation fitness and recommend the scholar with the highest matching degree as a potential cooperation partner; Cooperation mode suggestion: Combine the research fields, technical expertise, and past cooperation experiences of scholars to provide suitable cooperation mode suggestions.

[0014] Preferably, the generation of the intelligent cooperation framework further includes: Research direction recommendation: According to the cooperation matching results, recommend research directions that both parties to the cooperation are interested in, and optimize the recommendation results in combination with the academic frontier dynamics; Task division optimization: Based on the research expertise of both parties to the cooperation, divide the cooperation tasks into different modules and automatically allocate them to the most suitable scholars; Cooperation plan generation: Based on the research progress of both parties to the cooperation, automatically generate a cooperation schedule, optimize the research plan, and improve the cooperation efficiency.

[0015] The present invention provides a wisdom reading recommendation and cooperation analysis method. It has the following beneficial effects: 1. The present invention adopts an interdisciplinary cooperation matching technical solution based on scholars' research interest vectors and cooperation fitness, achieving the technical effect of accurately identifying scholars with high fitness. Compared with the traditional keyword matching or domain-based matching methods in the prior art, the present invention can analyze the multi-dimensional interests and potentials of scholars more meticulously, avoiding the problem of over-reliance on surface similarity in the traditional methods.

[0016] 2. By combining citation relationships and academic social network information, the present invention constructs a more comprehensive cooperation fitness evaluation system, achieving the effect of comprehensively evaluating the cooperation potential of scholars. Compared with the technical solutions in the prior art that only rely on a single data source for cooperation recommendation, the present invention effectively enhances the reliability of the recommendation through multi-dimensional data fusion.

[0017] 3. The present invention adopts an intelligent task allocation mechanism, reasonably divides tasks according to the research expertise and cooperation fitness of scholars, achieving the technical effects of optimizing resource allocation and improving cooperation efficiency. Compared with the simple task division methods in the prior art, the present invention avoids the problems of resource waste and uneven task allocation through refined division of labor, and greatly improves the execution efficiency of cooperation projects.

[0018] 4. Through the generation of the intelligent cooperation framework, the present invention provides an automated recommendation mechanism based on scholars' interest vectors and research directions, achieving the technical effect of efficiently generating cooperation suggestions. Compared with the practices in the prior art that rely on manual intervention or fixed recommendation models, the automated framework generation of the present invention can be adjusted in real time according to scholars' interests and research dynamics. Brief Description of the Drawings

[0019] Figure 1 It is a schematic flowchart of the method of the present invention. Detailed Embodiments

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] Please refer to the attached Figure 1 , the embodiment of the present invention provides an intelligent reading recommendation and cooperation analysis method, including the following steps: S1. Academic data collection and preprocessing, collecting historical research data of scholars, where the research data includes papers, academic conference records, scientific research fund applications, and academic social platform data, performing stop word removal and word segmentation processing on the research data, and converting the text content into a vector representation; In the intelligent reading recommendation and cooperation analysis method, the collection and preprocessing of academic data is the starting point of the entire system, directly affecting the accuracy of subsequent research interest modeling, interest evolution prediction, and personalized recommendation. Therefore, in the data acquisition stage, it is necessary to ensure the diversity and integrity of the data source, and at the same time minimize the interference of irrelevant information in the preprocessing to construct a high-quality research interest vector. Generally, the research data of scholars is often scattered in different platforms and formats, so appropriate technical means need to be taken to uniformly format the data and convert it into a computable vector representation.

[0022] First, academic data is collected. The collected academic data mainly includes published papers, academic conference records, scientific research fund applications, and academic social platform data. Among them, the paper data can be sourced from journal databases (such as IEEE Xplore, ACM Digital Library, Springer, etc.), the conference data requires extracting relevant conference proceedings, and the scientific research fund applications can be obtained from the public data of national scientific research funding agencies, and the academic social platform data comes from platforms such as ResearchGate and Google Scholar. For data from different sources, structured processing is required to ensure the unity of the data.

[0023] In a possible implementation, for academic paper data, the paper title, abstract, keywords, body content, and its reference list are mainly extracted. Among them, the body content is usually quite large. To improve processing efficiency, generally, only parts such as the introduction, methods, and experimental analysis can be selected for text processing. At the same time, the reference list can build a paper citation network to support subsequent interest evolution prediction and collaboration matching.

[0024] For academic conference records, since some conferences may not publicly disclose detailed papers, the conference topics, speakers, abstracts, and related keywords can be extracted to supplement the information on scholars' research directions. For research fund applications, the research project name, abstract, funding amount, research objectives, and project members are mainly concerned to analyze scholars' research interest directions and their academic collaboration situations.

[0025] In some embodiments, to enrich scholars' research data, academic social platform data is further introduced, such as information on scholars' research interest tags, publication dynamics, and being followed on ResearchGate. As an option, these data can be parsed through natural language processing techniques to extract keywords related to scholars' research interests and fuse them with structured data.

[0026] After data collection is completed, data preprocessing is required to improve the stability and accuracy of subsequent calculations.

[0027] In this embodiment, text cleaning is first performed. Specifically, irrelevant content such as HTML tags, special characters, and redundant spaces is removed. Then, stop word removal processing is carried out, and a stop word list is used to filter out common but meaningless words, such as "the", "is", etc., to reduce noise interference.

[0028] In some embodiments, the text data is subjected to word segmentation processing for subsequent feature extraction. Generally, for English text, standard space word segmentation or dictionary-based word segmentation methods are adopted. For Chinese text, word segmentation algorithms based on statistical learning, such as jieba segmentation or THULAC, can be used. As an option, named entity recognition (NER) can also be combined to identify specific terms and discipline nouns in the research field to improve the accuracy of research topic extraction.

[0029] After data cleaning and word segmentation, the text data needs to be converted into a vector representation for subsequent calculations.

[0030] In this embodiment, multiple text representation methods are adopted to improve the vector expression ability. First, the TF-IDF (term frequency-inverse document frequency) method is used to calculate the importance of text keywords. The TF-IDF calculation formula is as follows: TF-IDF(t, d) = TF(t, d) × IDF(t); Among them, d represents a specific document; t represents a word or term in the text; IDF is the inverse document frequency, indicating the importance of term t; TF represents the frequency of term t appearing in document d, and the calculation formula is: Among them, TF represents the frequency of term t appearing in document d; f t is the number of occurrences of term t in document d; ∑ t′ f t′ is the total number of occurrences of all terms in the document.

[0031] IDF is used to measure the importance of a term in the entire document collection, and the calculation formula is as follows: Among them, IDF is the inverse document frequency, indicating the importance of term t; t represents a word or term in the text; N is the total number of documents; DF(t) represents the number of documents containing term t. To avoid a zero denominator, a smoothing process is performed by adding 1.

[0032] In some embodiments, the Word2Vec method can also be used to map the text to a low-dimensional continuous vector space. As an option, the CBOW (Continuous Bag-of-Words) or Skip-gram model can be used to train semantic vectors, so that the distances between words with similar concepts in the high-dimensional space are closer. For short texts, Doc2Vec can be further used to convert the whole paper or conference abstract into a vector with a fixed dimension to capture richer semantic information.

[0033] In addition, in some embodiments, BERT can also be combined to obtain more contextually relevant text vectors through a pre-trained deep language model. BERT adopts a masked language model and a next sentence prediction mechanism, and is pre-trained on a large-scale corpus, which can effectively capture the bidirectional semantic relationships of the text.

[0034] During the data preprocessing process, vector normalization also needs to be considered. Generally, the value range of text vectors is large, and directly using them for calculation may lead to numerical instability. Therefore, min-max normalization or Z-score normalization can be adopted to map all vectors to the same scale and improve the robustness of the calculation.

[0035] In a possible implementation manner, the following normalization formula can be adopted: where min(X) is the minimum value in the dataset; max(X) is the maximum value in the dataset; X represents the original vector value; X ′ is the value after normalization. This method can ensure that the data range is normalized to [0,1], thereby reducing the calculation deviation caused by different feature scales.

[0036] S2. Research interest modeling: Based on the scholar's historical research data, extract research topics and construct a scholar's research interest vector, where the research interest vector is calculated with weights according to the semantic features and citation situation of the scholar's published papers. Research interest modeling is one of the core links. Its main goal is to extract research topics from the scholar's historical research data and calculate the research interest vector based on semantic features and citation situations. This research interest vector is not only used to represent the scholar's current research direction but also provides a basis for subsequent prediction of interest evolution and personalized recommendation. Therefore, it is necessary to fully consider the text semantic information and, at the same time, combine the citation situation to weight different research topics to enhance the accuracy and interpretability of vector representation. Generally, research interests are not static but change dynamically over time. Therefore, in the modeling process, it is necessary to comprehensively consider time factors and the changing trends of research directions to ensure the timeliness and accuracy of research interests.

[0037] In this embodiment, first, based on the scholar's historical research data, including published papers, academic conference records, research fund applications, and academic social platform data, research topic extraction is performed. In a possible implementation, a topic model (TopicModel) is used to identify the core topics of the scholar's research interests. Among them, Latent Dirichlet Allocation (LDA) is a commonly used modeling method. LDA infers the topic distribution in the document through statistical methods, and its mathematical expression is as follows: where w is the word; z is the topic; d is the document; k is the total number of topics; K is the total number of topics; P(w|z) represents the probability distribution of the word under a certain topic; P(z|d) represents the probability of the document belonging to each topic. LDA obtains the topic distribution of each paper through Bayesian inference, thereby identifying the scholar's research interests.

[0038] As an option, for short text data (such as conference abstracts, research fund applications), the TF-IDF combined with the K-Means clustering method can be used for topic clustering. In this method, first, calculate the TF-IDF feature vector of the document, and then use the K-Means algorithm for text clustering to obtain different research topics.

[0039] In some embodiments, to further improve the accuracy of topic extraction, a deep learning model can also be combined, such as the BERT topic embedding model based on Transformer. This model can obtain more accurate semantic representations through a pre-trained language model, thereby improving the accuracy of research topic classification.

[0040] In this embodiment, after extracting the research topics, it is necessary to construct a scholar's research interest vector. Specifically, the research interest vector of each scholar can be represented as a weighted topic vector, defined as follows: where, V s is the research interest vector of scholar s; T i represents the semantic vector of the i-th research topic; w i is the corresponding weight; N is the total number of documents; the weight w i is determined by multiple factors, including semantic features and citation situation.

[0041] In a possible implementation, the topic weight w i is calculated as follows: where, w i is the corresponding weight; P(T i |D) represents the average topic probability of this research topic in all the scholar's papers; C i is the total number of citations of the papers related to this topic; t i is the publication time of the papers related to this topic (represented by a timestamp); α, β, γ are hyperparameters used to control the influence weights of semantic features, citation situation, and time factors. Generally, the optimal values of the hyperparameters can be determined by methods such as grid search or Bayesian optimization to improve the accuracy of interest modeling.

[0042] As an option, if a scholar's research topics are relatively extensive, the hierarchical clustering (HierarchicalClustering) method can be used to hierarchically organize the research topics, thereby representing the scholar's research interests at different granularities. For example, all the research topics can be clustered first to obtain several high-level topics, and then refined into multiple sub-topics to provide a more refined interest modeling.

[0043] In some embodiments, to improve the timeliness of the research interest vector, a time decay factor can be introduced. Specifically, an exponential decay function is used to calculate the research interest weight within a certain time period, defined as follows: where, w i (t) is at time t; w iis the corresponding weight; T current represents the current time; λ is the time decay coefficient, usually determined by experiments; t i is the last occurrence time of word i. This method can ensure that the influence weight of recently published papers on research interests is greater than that of earlier published papers, thus dynamically reflecting the change of scholars' research directions.

[0044] In this embodiment, after the research interest vector is constructed, it can be further used for predicting the evolution of research interests and personalized recommendation. In the scenario of personalized recommendation, the cosine similarity between the scholar's interest vector and the literature vector can be calculated, and its calculation formula is as follows: where V s is the scholar's research interest vector; V d is the topic vector of the literature to be recommended; the higher the value of cos(θ), the more the literature matches the scholar's research interests, and the higher the recommendation priority.

[0045] In a possible implementation, the research interest vector can also be used for academic cooperation matching. By calculating the Euclidean distance or Mahalanobis distance between the interest vectors of different scholars, potential cooperation partners can be identified. For example, if the research interest vectors of two scholars are close in the high-dimensional space, it indicates that their research directions have strong similarity, and thus they are recommended as potential cooperation partners.

[0046] S3. Prediction of the evolution of research interests. Based on time series modeling, analyze the historical change trend of the scholar's research data interests, and use the long short-term memory network for prediction to determine the research directions that the scholar will focus on in the future; In the research interest modeling method, research interests are not static but evolve continuously with the academic activities of scholars. Therefore, in order to more accurately depict the change trend of scholars' research interests, it is necessary to analyze historical data based on time series modeling and use deep learning methods for prediction to determine the research directions that scholars may focus on in the future. Generally, the change of research interests is affected by various factors, including the evolution of research hotspots, the change of scholars' cooperation relationships, and the expansion of personal research directions. Therefore, in the modeling process, it is necessary to fully consider the time factor and introduce modeling methods that can capture long-term dependence relationships to improve the accuracy and stability of prediction.

[0047] In this embodiment, first, a research interest time series is constructed. Specifically, based on the historical research data of scholars, their research interest vectors are divided according to a time window to form a time series dataset. The time window can be divided according to annual, quarterly, or other appropriate time intervals, and the specific choice depends on the rate of change of scholars' research interests. For example, for rapidly developing fields such as artificial intelligence, it is recommended to use a shorter time window, while for relatively stable research fields such as mathematics or physics, a longer time window can be used.

[0048] In a possible implementation, the time series data is constructed using a sliding window method, that is, for any time point t, its research interest vector is composed of the interest vectors of the past m time points. Let denote the research interest vector of scholar s at time t, then the time series data can be expressed as: where m is the length of the time window; X s is the similarity of vocabulary S; V s is the research interest vector of the scholar; and t is the current time.

[0049] In some embodiments, in order to enhance the stability of the time series data, time-weighted smoothing can be performed on the interest vectors, that is, lower weights are assigned to the data at earlier times, while higher weights are assigned to the recent data to emphasize the impact of changes in recent research directions on future trends.

[0050] In this embodiment, in order to model the evolution trend of research interests, a long short-term memory network (LSTM) is used for prediction. As a type of recurrent neural network, LSTM can effectively capture long-term dependencies and is suitable for time series prediction tasks.

[0051] Generally, the input of the LSTM network is a sequence of historical research interest vectors, and the output is the predicted value of the research interest vector at future time points. The core computational units of LSTM include an input gate, a forget gate, and an output gate, which are used to control the transmission of information and memory update respectively.

[0052] Specifically, during the LSTM calculation process, first, the activation value of the forget gate is calculated to determine whether the information from the previous time step needs to be retained at the current time step. Then, the activation value of the input gate is calculated to decide how much of the new input information should be stored in the memory unit. Finally, through the output gate, the hidden state information of the current time step is determined and used for the calculation of subsequent time steps.

[0053] In a possible implementation, the dimension of the hidden layer of the LSTM network is determined experimentally and is usually set to 1.5 to 2 times the dimension of the research interest vector to ensure that the network has sufficient expressive power while avoiding overfitting. In addition, to improve the stability of the prediction, a bidirectional LSTM (Bi-LSTM) is adopted, that is, on the basis of the standard LSTM structure, a backpropagation path is introduced to enhance the model's ability to understand historical data.

[0054] In some embodiments, to further improve the accuracy of the prediction, a variational autoencoder (VAE) can be combined to learn the potential change patterns of the research interest vector. The VAE maps the research interest vector to a latent space through an encoder-decoder structure and generates possible future interest vectors by sampling, thereby improving the diversity and stability of the prediction.

[0055] In this embodiment, when training the LSTM prediction model, the cosine similarity loss is used as the loss function to ensure the similarity in direction between the predicted interest vector and the true interest vector while avoiding the influence of the numerical scale on the optimization process.

[0056] In addition, during the prediction process, the problem of topic drift (TopicDrift) of the research interest needs to be considered, that is, the research direction of scholars may change suddenly over time. Generally, the occurrence of topic drift may be affected by external factors (such as the rise of new technologies and changes in research cooperation). Therefore, during the prediction process, external academic dynamic data (such as hot papers and research hotspots) can be combined for supplementation to improve the prediction ability of scholars' interest changes.

[0057] In a possible implementation, to improve the detection ability of topic drift, the change rate of the research interest vector at consecutive time steps can be calculated and a drift threshold can be set. When the change rate exceeds the threshold, the LSTM prediction strategy is adjusted. For example, if a scholar's research direction has remained stable in the past five years but has changed significantly in the most recent year, the attention weight of the model to the most recent time step can be increased to ensure that the prediction results can reflect the new research trends in a timely manner.

[0058] In some embodiments, an attention mechanism (AttentionMechanism) can also be combined to enhance the model's attention to key time steps. Specifically, during the LSTM calculation process, attention weights are introduced, enabling the model to adaptively allocate importance weights for different time steps, thereby improving the flexibility and accuracy of the prediction of the evolution of research interests.

[0059] S4, Personalized reading recommendation: Based on the prediction results of the evolution of research data interests, calculate the similarity between the research data interest vector of the scholar and the literature vectors in the academic literature database, and screen and recommend the most relevant literature; Under the framework of research interest modeling and evolution prediction, personalized reading recommendation aims to automatically match highly relevant literature based on the scholar's interest vector to improve the scholar's academic reading efficiency. Research interests are not static but dynamically evolving. Therefore, during the recommendation process, it is necessary to comprehensively consider the scholar's historical research interests and their possible future research directions to provide more accurate recommendation results. Generally, traditional keyword-matching-based recommendation methods are difficult to effectively capture the semantic information of research interests. Therefore, in the present invention, a vectorization modeling method is adopted to calculate the high-dimensional similarity between the scholar's interest vector and the literature vector to screen the most relevant academic literature for recommendation.

[0060] In this embodiment, first, based on the research data interest evolution prediction result, the scholar's research interest vector is obtained, and the research interest representation of the scholar at the current moment is constructed. Specifically, the research interest vector is obtained by weighted calculation of the scholar's research interests in different time windows, and combined with a time decay mechanism, so that the recent research interests have a greater impact on the recommendation result. In a possible implementation, the time decay weight adopts an exponential decay method to ensure that the interest vector closer to the current time has a greater recommendation weight.

[0061] As an option, to improve the stability of the research interest vector, the interest vectors of multiple time steps can be processed by moving average, so that the volatility of the research interest vector is reduced and the extreme impact of a single paper on the recommendation result is reduced. For example, a weighted moving average method can be used to smooth the interest vectors of the past n time steps to improve the stability of the vector.

[0062] After obtaining the scholar's research interest vector, it is necessary to calculate its similarity with each piece of literature in the academic literature database. Specifically, each piece of literature to be recommended has a literature vector, which is composed of factors such as the literature's topic distribution, keyword features, and citation relationships. Generally, the literature vector can be calculated by text embedding methods such as TF-IDF, Word2Vec, BERT, etc., and dimension reduction is performed by principal component analysis (PCA) to reduce the computational complexity.

[0063] In this embodiment, the cosine similarity method is used for similarity calculation, which can effectively measure the directional similarity of two vectors in a high-dimensional space. Specifically, the scholar's interest vector V s and the literature vector V d The similarity S between them is calculated as follows: where, V s is the scholar's research interest vector; V d is the topic vector of the literature to be recommended; S is the similarity.

[0064] As an option, to further improve the diversity of recommendation results, literature screening can be carried out by combining the Mahalanobis Distance on the basis of cosine similarity. The Mahalanobis Distance can take into account the correlation between features, thus effectively avoiding the problem of topic overfitting caused by only relying on word vector similarity. In some embodiments, the calculation formula of the Mahalanobis Distance is as follows: where s is the similarity; D M calculates the normalized distance between two vectors. The smaller the Mahalanobis Distance, the closer the research interest is to the distribution of the literature in the high-dimensional space; V s is the scholar's research interest vector; V d is the topic vector of the literature to be recommended; T represents the transpose operation of a matrix or vector, which is used to interchange the rows and columns of a vector or matrix.

[0065] In this embodiment, after calculating the similarity, it is necessary to screen out the literature with the highest correlation for recommendation. Generally, the number of recommended literatures can be dynamically adjusted according to the scholar's reading habits. For example, for active scholars, more literatures can be recommended, while for scholars with a more focused research direction, the number of recommendations can be reduced to improve the recommendation accuracy. In one possible implementation, the number N of recommended literatures r adopts a dynamic adjustment strategy, and its calculation formula is as follows: N r = N base + λ × log(C s + 1); where N base is the basic number of recommendations; C s is the literature reading volume of the scholar in the recent period; λ is the time decay coefficient, which is usually determined by experiments; this strategy ensures that the number of recommendations is adjusted according to the scholar's academic activity to meet the needs of different types of users.

[0066] In some embodiments, to improve the diversity of recommendation results, an exploratory recommendation mechanism can be introduced, that is, a certain proportion of low-correlation but academically potential literatures are randomly added to the screened high-similarity literatures. As an option, based on the citation network of academic literatures, some literatures that are not directly matched but have potential connections can be calculated and added to the recommendation list with a certain probability to increase the scholar's attention to new research directions.

[0067] S5. Cross-disciplinary cooperation matching, based on the scholar's research data interest vector, citation relationship, and academic social network relationship, calculate the cooperation fit degree between different scholars, and screen potential cooperation partners according to the fit degree; In academia, interdisciplinary collaboration can promote innovation and drive scientific progress. To efficiently identify scholars with potential for collaboration, this invention proposes an interdisciplinary collaboration matching method based on scholars' research data interest vectors, citation relationships, and academic social network relationships. By comparing the research interest vectors of different scholars and combining their influence in the academic field, citation relationships, and interaction information in the academic social network, the collaboration compatibility between scholars can be effectively evaluated, thus screening out potential collaboration partners. This method can not only recommend collaboration partners based on scholars' academic interests but also further optimize the matching effect through citation and social network information.

[0068] In this embodiment, first, based on the research interest vectors of scholars, the interest similarity between scholars is calculated. As mentioned before, the research interest vector is predicted through the evolution of scholars' historical research data and represents the research direction of a scholar at a certain moment. Specifically, the interest similarity S int between scholars can be calculated using the cosine similarity, and the formula is as follows: where V i and V j are the research interest vectors of scholars i and j respectively; · represents the dot product of vectors, and ∥V i ∥ and ∥V j ∥ are the Euclidean norms of the interest vectors respectively. This formula measures the similarity of the research directions between scholars i and j, and the larger the value, the closer their research directions are.

[0069] As an option, to avoid the limitations of a single similarity measure, other similarity measure methods such as Jaccard similarity or Manhattan distance can be combined to further enrich the dimensions of the matching.

[0070] Secondly, the citation relationship, as an important manifestation of the academic influence between scholars, can provide important information for collaboration matching. In this embodiment, the citation relationship is evaluated by calculating the citation overlap between scholars to assess the intersection of their academic influences. Specifically, the citation overlap S cite between scholar i and scholar j can be calculated as the ratio of the size of the intersection of the papers they cited to the size of the union: where C i and C j represent the citation document sets of scholars i and j respectively; |C i ∩C j | is the size of the intersection of their cited documents; |C i ∪C j| is the union size of the cited references of both. This formula can reflect the intersection degree of cited references among scholars, thus indirectly evaluating the tightness of academic connections among scholars.

[0071] Specifically, the analysis of citation relationships can be combined with the number of times a scholar's papers are cited, and higher weights can be given to scholars with higher citation counts to reflect their influence in the academic field. In some embodiments, the influence of citation relationships can be combined with the interest vectors and weighted to improve the matching degree between highly cited scholars and scholars with matching interests.

[0072] In addition, the academic social network relationships of scholars, such as co-authored papers, jointly participated academic activities, jointly funded research projects, etc., also provide strong evidence for interdisciplinary cooperation. In this embodiment, the social network relationships are represented by constructing a cooperation graph among scholars, where the nodes in the graph represent scholars and the edges represent the academic cooperation relationships among scholars. The weight of each edge can be determined by the cooperation frequency among scholars or the influence of cooperation results. The construction and analysis methods of the cooperation graph are similar to the community discovery algorithms in social network analysis, such as the Louvain algorithm or K-core analysis, for discovering cooperation communities and cooperation potential among scholars.

[0073] In some embodiments, the social network relationships of scholars not only rely on co-authored academic papers but can also be combined with academic conferences, lectures, and other academic interaction events. By integrating this social information, the calculation of the cooperation fitness among scholars can be further optimized, so that the recommended cooperation partners are not limited to scholars with similar academic fields but can also include those scholars with potential cooperation opportunities.

[0074] In this embodiment, the cooperation fitness S among scholars coop combines the research interest similarity, citation relationships, and academic social network relationships and is comprehensively calculated as follows: S coop (i,j) = αS int (i,j) + βS cite (i,j) + γS net (i,j); where α, β, γ are hyperparameters used to control the influence weights of semantic features, citation situations, and time factors; S coop (i,j) represents the cooperation tightness between scholar i and scholar j in the academic social network and can be calculated through graph theory algorithms; S cite represents the citation overlap degree between scholar i and scholar j; S int represents the interest similarity between scholars; S net represents the cooperation degree between scholar i and scholar j.

[0075] In a possible implementation, the screening method based on cooperation fit is as follows: First, calculate the cooperation fit between each scholar and all other scholars, sort them according to the fit values, and select the top N scholars with the highest fit as potential cooperation partners. To avoid over-recommending scholars in the same field, a field diversity constraint can be set to ensure that the recommended results include scholars from different disciplinary fields to promote true interdisciplinary cooperation.

[0076] S6. Generation of an intelligent cooperation framework. After determining high-fit cooperation partners, generate cooperation suggestions, recommend research directions, and divide tasks according to the research expertise of scholars to form a preliminary cooperation framework. In the process of academic cooperation, in addition to accurately identifying high-fit cooperation partners, it is equally crucial to further generate an effective cooperation framework. Through multi-dimensional information such as research interests, citation relationships, and academic social networks among scholars, after identifying high-fit cooperation partners, the present invention further generates cooperation suggestions, recommends appropriate research directions, and at the same time combines the expertise of scholars to carry out reasonable task division to form a preliminary cooperation framework. This process of generating an intelligent cooperation framework aims to ensure the smooth progress of cooperation and maximize the cooperation benefits through effective resource allocation.

[0077] In this embodiment, first, based on the previously calculated results of cooperation fit, identify cooperation scholars with relatively high fit. On this basis, generate cooperation suggestions and recommend appropriate research directions. Specifically, the recommendation of research directions takes into account the current research interest vector of scholars and the prediction of future research trends. These research directions usually involve the forefront issues in the current field of scholars and combine interdisciplinary fields that scholars may be interested in in the future to promote more innovative cooperation. For example, if the research interests of scholar A and scholar B are highly consistent in certain fields, and scholar A has the potential to expand into bioinformatics in the future, while scholar B has rich experience in data analysis, then it can be recommended that the two jointly carry out research in the field of bioinformatics.

[0078] As an option, when generating cooperation research directions, academic hotspots, that is, currently popular or urgent research issues in the academic community, can be combined. In this process, external data sources, such as hot papers in academic journals and key funding directions of scientific research funding agencies, can be introduced as a reference basis for the selection of research directions. In this way, the recommended research directions not only meet the interests of scholars but also ensure their high academic value and practical application prospects.

[0079] Specifically, the process of generating cooperation suggestions includes the following steps: First, calculate and generate the potential cooperation research fields of each scholar through the scholar's research interest vector, academic social network analysis, and citation relationships; then, propose specific research questions or topics by combining the scholar's academic development direction and cutting-edge hotspots; finally, ensure the practical feasibility of the recommended research direction in cooperation by combining the scholar's expertise.

[0080] In some embodiments, the recommended research directions are automatically generated by an algorithm model and accompanied by predictions of possible research results, so that scholars can make effective decisions based on the expected results.

[0081] In this embodiment, after determining the cooperation research direction, the next step is to make a reasonable task division according to the research expertise of the scholars. The principle of task division is to assign suitable scholars to each task in the cooperation project according to the scholar's professional background, research ability, and historical research results. For example, in a cooperation involving data processing and experimental design, Scholar A is good at data analysis, so he is assigned to the data processing module; while Scholar B has deep accumulation in experimental design and model verification, so he is assigned to the experimental design module. During the task division process, the cooperation history, research output, and academic achievements of the scholars can be combined for evaluation to ensure the scientificity and rationality of the division.

[0082] In a possible implementation, the specific steps of task division include: First, classify all research tasks based on the research expertise of the scholars, such as data analysis, model design, experimental design, etc.; then, match each task according to the scholar's interest vector and past research results; finally, reasonably allocate tasks according to the cooperation fit between scholars and task requirements to form a preliminary cooperation framework. During the task allocation process, optimization algorithms, such as linear programming or genetic algorithms, can be used to find the optimal task allocation plan under various constraints.

[0083] In some embodiments, task division and cooperation framework generation can be combined with an intelligent recommendation system to dynamically adjust the task allocation of scholars through historical data and real-time feedback. For example, during the cooperation process, if a scholar has a low completion rate for a certain task, the system can automatically adjust the task allocation to ensure that the progress of the cooperation project is not affected.

[0084] The generated cooperation framework not only includes research directions and task division, but also can provide scholars with cooperation progress tracking and collaboration platform support. Specifically, the cooperation framework should include the execution schedule, resource requirements, expected results, etc. of each task, and be monitored and fed back through a project management system. Scholars can view the completion status of tasks, research progress, and the contributions of partners in real time on this platform, thereby improving cooperation efficiency.

[0085] As an option, in order to improve the adaptability of the cooperation framework, an adaptive adjustment mechanism can be added to automatically adjust the task division and research direction according to the performance of scholars in cooperation. For example, if a scholar shows new research interests or acquires new skills during cooperation, the system can dynamically update the cooperation framework according to the new changes.

[0086] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent reading recommendation and cooperation analysis, characterized in that It includes the following steps: Academic data collection and preprocessing: Collect the historical research data of scholars, where the research data includes papers, academic conference records, scientific research fund applications, and academic social platform data. Perform stop-word removal and word segmentation on the research data, and convert the text content into vector representation; Research interest modeling: Based on the historical research data of scholars, extract research topics and construct scholars' research interest vectors, where the research interest vectors are weighted and calculated according to the semantic features and citation situations of the papers published by scholars; Prediction of research interest evolution: Based on time series modeling, analyze the historical change trends of scholars' research data interests, and use long short-term memory networks for prediction to determine the future research directions that scholars will focus on; Personalized reading recommendation: Based on the prediction results of research data interest evolution, calculate the similarity between the research data interest vectors of scholars and the literature vectors in the academic literature database, and screen and recommend the most relevant literature; Interdisciplinary cooperation matching: Based on the research data interest vectors, citation relationships, and academic social network relationships of scholars, calculate the cooperation fit between different scholars, and screen potential cooperation partners according to the fit; Intelligent cooperation framework generation: After determining high-fit cooperation partners, generate cooperation suggestions, recommend research directions, and perform task division according to the research expertise of scholars to form a preliminary cooperation framework.

2. The intelligent reading recommendation and cooperation analysis method according to claim 1, wherein The research topics include: Text feature extraction: Based on the historical research data of scholars, use natural language processing methods to extract text features and identify the paper titles, abstracts, keywords, and main research contents; Thematic clustering analysis: Use clustering algorithms to perform clustering analysis on text features, generate multiple research topic categories, and calculate the proportion of each research topic in scholars' research interests; Weight calculation: Assign weights to each research topic, where the weights are calculated based on the citation times, publication time, and influence in the research field of the papers to highlight the core research directions.

3. A method for intelligent reading recommendation and cooperation analysis according to claim 1, characterized in that, The time series modeling includes: Time window division: According to the time when scholars publish papers, group the research interest vectors in chronological order to form a time series data set; Feature extraction: Extract features from the research interest vectors of different time windows to identify the change trends of research topics over time; Sequence modeling: Use recurrent neural networks or variational autoencoders for sequence modeling to establish a prediction model for the change of research interests over time.

4. A method for intelligent reading recommendation and cooperation analysis according to claim 1, characterized in that The historical change trends include: Research topic change: Analyze the changes in scholars' research topics in different time periods and identify the evolution trends of research focuses; Interdisciplinary trend: Detect whether scholars' research directions gradually expand to other disciplinary fields and analyze the degree of interdisciplinary research; Research intensity fluctuation: Count the number and influence of papers published in different time periods to judge the changes in research activity.

5. The intelligent reading recommendation and cooperation analysis method according to claim 1, characterized in that The long short-term memory network includes: Input layer construction: Take the research interest vector as the input and perform normalization processing; Hidden layer calculation: Use long short-term memory units to process time series data and remember the evolution trends of research interests over a long time span; The output layer is optimized to perform a non-linear transformation on the prediction results, ensuring the accuracy of the research interest prediction results, and optimizing the prediction model through the loss function.

6. The intelligent reading recommendation and cooperation analysis method according to claim 1, wherein, The prediction results include: Future research topics, predicting the research topics that will be focused on in the future based on the historical research data of scholars; Potential interdisciplinary fields, identifying the interdisciplinary research directions that scholars are involved in and providing corresponding research suggestions; Research development trends, predicting the matching degree between the research directions of scholars and the international frontiers by combining the global academic development trends.

7. A method for intelligent reading recommendation and cooperation analysis according to claim 1, characterized in that The research data interest vector includes: Topic vector, converting the research topic into a vector representation according to the modeling results of scholars' research interests; Time decay factor, assigning different weights to the research data in different time periods, with a higher weight for the recent research interest vector to reflect the latest research interests; Multi-modal fusion, constructing a comprehensive interest vector by combining various research forms such as text, charts, and experimental data.

8. A method for intelligent reading recommendation and cooperation analysis according to claim 1, characterized in that The literature vector includes: Text feature vector, constructing a text feature vector based on the title, abstract, keywords, and full text content of academic papers; Citation network vector, analyzing the citation relationships of papers, constructing a citation network, and calculating the influence of papers in the network; Author influence vector, calculating the contribution degree of authors in a specific research field by combining the historical publication records and academic influence of paper authors to optimize the recommendation results.

9. A method for intelligent reading recommendation and cooperation analysis according to claim 1, characterized in that The interdisciplinary cooperation matching further includes: Calculation of cooperation fitness, calculating the cooperation fitness between scholars based on their research interest vectors, citation relationships, and academic social networks; Screening of potential cooperation partners, recommending the scholars with the highest matching degree as potential cooperation partners according to the ranking of cooperation fitness; Suggestions on cooperation modes, providing suitable cooperation mode suggestions by combining the research fields, technical expertise, and past cooperation experiences of scholars.

10. A method for intelligent reading recommendation and cooperation analysis according to claim 1, characterized in that, The generation of the intelligent cooperation framework further includes: Recommendation of research directions, recommending the research directions that both parties of the cooperation are interested in according to the cooperation matching results and optimizing the recommendation results by combining the academic frontier trends; Optimization of task division, dividing the cooperation tasks into different modules according to the research expertise of both parties of the cooperation and automatically assigning them to the most suitable scholars; Generation of cooperation plans, automatically generating a cooperation schedule based on the research progress of both parties of the cooperation, optimizing the research plan, and improving the cooperation efficiency.

Citation Information

Patent Citations

  • Scholar cooperation relationship prediction method based on interest evolution

    CN111325390A

  • Scholar recommendation method and system based on scholar research interest knowledge graph, and medium

    CN114547275A

  • Smart reading recommendation and cooperation analysis method

    CN117312676A

  • Data mining method and device, equipment and storage medium

    CN119474171A