Scientific and technological information automatic extraction and recommendation method and system
By analyzing the user's historical browsing and interactive behavior in real time, combining the correlation index between scientific and technological articles and topics, a recommendation index is generated, which solves the problems of lag and mismatch in the existing technology, and efficient and personalized scientific and technological information recommendations are achieved.
Patent Information
- Application Number
- CN202510055114.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The prior art does not fully consider the dynamic technological interest changes of users to be recommended at different time periods, and ignores the analysis of interactive behavior, resulting in the recommendation results lag or do not match user needs.
By obtaining text information of scientific and technological articles in real time and pre-processing, extracting and analyzing it with pre-trained classification models, obtaining user's historical browsing and interactive behavior timing data, analyzing the user's interest preference index, and comprehensively analyzing the correlation index between the article and the topic, generating a recommendation index for pushing.
It realizes the accuracy and personalization of recommendations, ensures that the recommended content is highly consistent with the user's latest interests, and improves user stickiness and reading experience.
Smart Images

Figure CN120030229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information data analysis, and in particular to a method and system for automatically extracting and recommending scientific and technological information. Background Art
[0002] With the continuous development of information technology, the Internet and various information platforms have generated a large amount of scientific and technological information. At present, traditional information retrieval methods mostly rely on keyword searches. Although they can provide certain relevant information, when faced with massive amounts of information, the retrieval efficiency is low and it is difficult to meet personalized needs. Automated information extraction and recommendation technology, through natural language processing, machine learning, data mining and other technical means, can quickly extract key information from a huge information base and make personalized recommendations based on user needs. With the continuous development of big data and artificial intelligence technologies, automatic information extraction and recommendation systems are evolving rapidly.
[0003] Prior art, such as a method and system for automatically extracting and recommending science and technology policy information disclosed in a patent application with announcement number: CN117743564B, collects data from a target website based on a preset crawler strategy to obtain science and technology policy source data; extracts keywords from the science and technology policy source data to form feature word data; performs text conversion and semantic analysis based on a science and technology policy database to extract entity, attribute, and relationship data in science and technology policies, and constructs a knowledge graph; based on the feature word data and user feature keywords, based on a collaborative recommendation algorithm, retrieves recommended data from the knowledge graph to obtain first recommended policy data; based on the user's real-time website browsing data, determines whether the recommended data needs to be updated through the knowledge graph, and if so, performs a secondary data search based on the user's real-time data and the knowledge graph to obtain second recommended policy data. Through the present invention, the update time of the recommended data can be effectively located and the corresponding recommended data can be effectively inferred, thereby improving the efficiency of users in analyzing science and technology policies.
[0004] Based on the above solution, it is found that the limitations of the existing technology include at least the following problems. First, the existing technology does not fully consider the dynamic changes in the technological interests of the recommended users in different time periods, which may easily lead to delayed recommendation results or mismatch with the user's current actual needs, and thus it is difficult to reflect the user's interests in real time. Secondly, the analysis of interactive behaviors is ignored, which may easily lead to ignoring the user's deep-seated needs and interests, thereby affecting the personalization and accuracy of the recommendation results, and failing to provide the technological information that best meets the user's needs. Summary of the invention
[0005] In view of the deficiencies of the prior art, the present invention provides a method and system for automatically extracting and recommending scientific and technological information, which solves the problem that the prior art does not fully consider the dynamic changes in scientific and technological interests of users to be recommended in different time periods and ignores the analysis of interactive behaviors.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for automatic extraction and recommendation of scientific and technological information, comprising the following steps: real-time acquisition of text information of several scientific and technological articles to be extracted, and preprocessing; inputting the preprocessed text information of each scientific and technological article to be extracted into a pre-trained classification model for extraction and analysis, and obtaining a correlation index between each scientific and technological article to be extracted and each scientific and technological topic; at the same time, obtaining the historical browsing behavior time series data of each scientific and technological topic of the user to be recommended, the historical reading behavior time series data includes historical reading time series data and historical interaction time series data, and performing data analysis respectively to obtain each historical browsing behavior of each scientific and technological topic of the user to be recommended. The historical reading preference index and historical interaction preference index of each period are calculated and comprehensively analyzed to obtain the comprehensive interest preference index of each historical period of each scientific and technological theme of the recommended user; the comprehensive interest preference index of each historical period of each scientific and technological theme of the recommended user is processed by moving weighted average to obtain the comprehensive interest preference change index of each scientific and technological theme of the recommended user, and a comprehensive analysis is performed in combination with the correlation index of each scientific article to be extracted and each scientific theme to obtain the recommendation index of each scientific article to be extracted; the recommendation index of each scientific article to be extracted is judged and analyzed with the preset recommendation index threshold, and the recommended user is pushed based on the judgment and analysis results.
[0007] Furthermore, the historical reading time series data includes the page scroll depth value, page focus stay time value, page scroll rate value, page pause time value, reading time value, reading number value, and reading response time value of each historical time period, and the historical interaction time series data includes the interaction conversion rate value, interaction frequency value, and interaction reaction speed value.
[0008] Furthermore, the classification model is specifically a bidirectional encoder representation model, which includes an input embedding layer, an encoder, a classification layer, and an output layer, and the specific process of obtaining the correlation index between each scientific article to be extracted and each scientific topic is as follows: in the input embedding layer of the bidirectional encoder representation model, the text information of each scientific article to be extracted is converted and processed to obtain the initial feature representation vector of each scientific article to be extracted; in the encoder of the bidirectional encoder representation model, the initial feature representation vector of each scientific article to be extracted is deeply processed to obtain the representation vector set of each scientific article to be extracted; in the classification layer of the bidirectional encoder representation model, the context-related representation vector set of each scientific article to be extracted is classified to obtain the correlation probability of each scientific article to be extracted and each scientific topic; in the output layer of the bidirectional encoder representation model, the correlation probability of each scientific article to be extracted and each scientific topic is judged and analyzed with the preset correlation probability threshold respectively to obtain the correlation index of each scientific article to be extracted and each scientific topic.
[0009] Furthermore, the pre-training process of the classification model is: obtain several groups of training text information, and divide them into a text information training set and a text information verification set; divide the text information training set into several batch training sets, and each batch training set contains several text information; perform forward propagation processing on each batch training set to obtain a representation vector set of each text information in each batch training set, and calculate the text loss function; perform iterative training at the same time, evaluate each training result of the classification model based on the back propagation algorithm and the text information verification set, and adjust the model parameters according to the verification results until the model meets the expected standards.
[0010] Furthermore, the specific steps for obtaining the historical reading preference index for each historical period of each scientific and technological theme of the user to be recommended are as follows: normalize the page scroll depth value, page focus dwell time value, page scroll rate value, page pause duration value, reading duration value, number of readings value, and reading response duration value for each historical period of each scientific and technological theme of the user to be recommended; comprehensively analyze the normalized page scroll depth value, page focus dwell time value, page scroll rate value, and page pause duration value for each historical period of each scientific and technological theme of the user to be recommended to obtain the historical attraction index for each historical period of each scientific and technological theme of the user to be recommended; and comprehensively analyze the normalized reading duration value, number of readings value, and reading response duration value for each historical period of each scientific and technological theme of the user to be recommended to obtain the historical reading investment index for each historical period of each scientific and technological theme of the user to be recommended; and comprehensively analyze the historical attraction index and historical reading investment index for each historical period of each scientific and technological theme of the user to be recommended to obtain the historical reading preference index for each historical period of each scientific and technological theme of the user to be recommended.
[0011] Furthermore, the specific formula for calculating the historical attraction index, historical reading investment index, and historical reading preference index of each historical period of each science and technology theme of the recommended user is as follows:
[0012] The historical attraction index of the jth historical period of the i-th technology theme recommended to the user, YgS′ ij is the page scroll depth value of the jth historical period of the i-th technology topic of the recommended user after normalization, α 1 is the depth coefficient stored in the database, YmS′ ij is the page focus dwell time value of the i-th technology topic of the recommended user in the j-th historical period after normalization, α 2 is the focal coefficient stored in the database, YdS′ ij is the page scrolling rate value of the jth historical period of the i-th technology topic of the recommended user after normalization, α 3 is the rate coefficient stored in the database, YtS′ ij is the normalized page pause duration of the jth historical period of the i-th technology topic of the recommended user, α 4 is the pause coefficient stored in the database, α 1 +α 2 +α 3 +α 4 =1, e is a natural constant, TyD ij YdS′ is the historical reading investment index of the jth historical period of the i-th scientific and technological topic of the recommended user,ij is the normalized reading time value of the jth historical period of the i-th science and technology topic of the recommended user, β 1 is the reading time coefficient stored in the database, YdS′ ij is the normalized reading count of the i-th technology topic in the j-th historical period of the recommended user, β 2 is the reading frequency coefficient stored in the database, YdS′ ij is the normalized reading response time value of the jth historical period of the i-th scientific and technological topic of the recommended user, β 3 is the response time coefficient stored in the database, β 1 +β 2 +β 3 =1,YdX ij is the historical reading preference index of the jth historical period of the i-th scientific and technological topic of the recommended user, δ 1 is the attraction coefficient stored in the database, δ 2 is the input coefficient stored in the database, δ 1 +δ 2 =1,i=1,2,3,…,i 0 ,i 0 is the number of scientific and technological subject categories, j = 1, 2, 3, ..., j 0 , j 0 is the number of historical time periods.
[0013] Furthermore, the specific steps for obtaining the historical interaction preference index for each historical period of each technological theme of the user to be recommended are as follows: standardize the interaction conversion rate value, interaction frequency value, and interaction reaction speed value for each historical period of each technological theme of the user to be recommended; comprehensively analyze the standardized interaction conversion rate value, interaction frequency value, and interaction reaction speed value for each historical period of each technological theme of the user to be recommended to obtain the historical interaction preference index for each historical period of each technological theme of the user to be recommended.
[0014] Furthermore, the formula for calculating the comprehensive interest preference index of each historical period of each technology theme of the recommended user is as follows: Among them, ZhQ ij YdX is the comprehensive interest preference index of the i-th technology topic in the j-th historical period of the recommended user, ij is the historical reading preference index of the jth historical period of the i-th science and technology topic of the user to be recommended, is the reading coefficient stored in the database, LhD ij is the historical interaction preference index of the i-th technology topic of the recommended user in the j-th historical period, is the interaction coefficient stored in the database, i=1,2,3,…,i 0 ,i 0 is the number of scientific and technological subject categories, j = 1, 2, 3, ..., j 0 , j 0 is the number of historical periods, and e is a natural constant.
[0015] Furthermore, the specific formula for calculating the recommendation index of each scientific article to be extracted is as follows: Among them, TqZ a is the recommendation index of the a-th scientific article to be extracted, BhZ i is the comprehensive interest preference change index of the i-th technology topic of the recommended user, μ 1 is the comprehensive interest coefficient stored in the database, ZtG ai is the correlation index between the a-th scientific article to be extracted and the i-th scientific topic, μ 2 is the correlation coefficient stored in the database, μ 1 +μ 2 =1,a=1,2,3,…,a 0 , a 0 is the number of scientific articles, i = 1, 2, 3, ..., i 0 ,i 0 is the number of science and technology theme types.
[0016] A system for automatic extraction and recommendation of scientific and technological information, comprising: a data acquisition module, a data extraction module, a data analysis module, a recommendation analysis module, and a user recommendation module; the data acquisition module is used to acquire the text information of a number of scientific and technological articles to be extracted in real time and perform preprocessing; the data extraction module is used to input the preprocessed text information of each scientific and technological article to be extracted into a pre-trained classification model for extraction and analysis, and obtain the correlation index between each scientific and technological article to be extracted and each scientific and technological theme; the data analysis module is used to simultaneously acquire the historical browsing behavior time series data of each scientific and technological theme of the user to be recommended, the historical reading behavior time series data includes historical reading time series data and historical interaction time series data, and perform data analysis respectively to obtain the correlation index of each scientific and technological theme of the user to be recommended. The historical reading preference index and historical interaction preference index of each historical period of the science and technology theme are comprehensively analyzed to obtain the comprehensive interest preference index of each historical period of each science and technology theme of the recommended user; the recommendation analysis module is used to perform moving weighted average processing on the comprehensive interest preference index of each historical period of each science and technology theme of the recommended user to obtain the comprehensive interest preference change index of each science and technology theme of the recommended user, and conduct a comprehensive analysis in combination with the correlation index of each science and technology article to be extracted and each science and technology theme to obtain the recommendation index of each science and technology article to be extracted; the user recommendation module is used to judge and analyze the recommendation index of each science and technology article to be extracted with a preset recommendation index threshold, and push the recommended user based on the judgment and analysis results.
[0017] The present invention has the following beneficial effects:
[0018] (1) This method of automatic extraction and recommendation of scientific information extracts the relevance of scientific articles to each scientific topic in real time, and combines it with the changes in user historical preferences to make recommendations more accurate. It uses the recommendation index as a filtering criterion to effectively select scientific articles that best meet user needs for push, thereby helping users quickly obtain scientific information of interest, and increasing the exposure and reading rate of articles, thereby enhancing the dissemination effect and influence of information.
[0019] (2) This method of automatic extraction and recommendation of scientific and technological information analyzes the user's browsing behavior time series data to accurately analyze the user's interest changes in each scientific and technological topic, and uses the comprehensive interest preference index to dynamically adjust the recommendation strategy, thereby ensuring that the recommended content is highly consistent with the user's latest interests and enhancing user stickiness, thereby enabling the recommendation system to continuously provide high-quality scientific and technological content that meets user needs and improves the user's reading experience.
[0020] (3) The automatic extraction and recommendation system of scientific and technological information realizes efficient information extraction and intelligent recommendation through the collaborative work of various modules. The data acquisition and extraction module ensures real-time and accuracy, while the data analysis module ensures that the recommended content keeps pace with changes in user needs by analyzing the user's historical behavior time series data. The weighted average processing of the recommendation analysis module and the comprehensive analysis of the correlation index make the recommendation not only accurate, but also dynamically adapt to changes in user interests, thereby significantly improving the efficiency and intelligence of the recommendation, and being able to adjust the recommendation strategy in real time, thereby providing a continuously optimized recommendation service experience.
[0021] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a flow chart of a method for automatically extracting and recommending scientific and technological information according to the present invention.
[0023] Figure 2 This is a block diagram of a system for automatic extraction and recommendation of scientific and technological information according to the present invention. DETAILED DESCRIPTION
[0024] The overall idea of the problem in the embodiment of this application is as follows:
[0025] First, the text information of several scientific articles to be extracted is obtained in real time. Then, it is input into the pre-trained classification model for extraction and analysis to obtain the correlation index between each scientific article to be extracted and each scientific topic. At the same time, the user's historical behavior time series data is obtained and analyzed to obtain the user's historical reading preference index and historical interaction preference index for each scientific topic in each historical period, and a comprehensive analysis is performed to obtain the comprehensive interest preference index of each scientific topic. Then, the comprehensive interest preference index of each historical period of the recommended user is weighted averaged to obtain the comprehensive interest preference change index. Finally, it is combined with the correlation index of each scientific article for comprehensive analysis to obtain the recommendation index, and the set recommendation index threshold is used to determine which scientific articles to push to the user.
[0026] See also Figure 1, an embodiment of the present invention provides a technical solution: a method for automatic extraction and recommendation of scientific and technological information, comprising the following steps: obtaining text information of several scientific and technological articles to be extracted in real time and performing preprocessing; inputting the preprocessed text information of each scientific and technological article to be extracted into a pre-trained classification model for extraction and analysis, and obtaining a correlation index between each scientific and technological article to be extracted and each scientific and technological topic; simultaneously obtaining the historical browsing behavior time series data of each scientific and technological topic of the user to be recommended, the historical reading behavior time series data including historical reading time series data and historical interaction time series data, and performing data analysis respectively to obtain the historical reading preference index and historical interaction preference index of each historical period of each scientific and technological topic of the user to be recommended, and performing comprehensive analysis to obtain the historical reading preference index of each scientific and technological topic of the user to be recommended. The comprehensive interest preference index of each historical period is calculated; the comprehensive interest preference index of each historical period of each scientific theme of the recommended user is processed by moving weighted average to obtain the comprehensive interest preference change index of each scientific theme of the recommended user, and a comprehensive analysis is performed in combination with the correlation index of each scientific article to be extracted and each scientific theme to obtain the recommendation index of each scientific article to be extracted; the recommendation index of each scientific article to be extracted is judged and analyzed with the preset recommendation index threshold, and the recommended user is pushed based on the judgment and analysis results (if the recommendation index of the scientific article to be extracted is lower than or equal to the preset recommendation index threshold, it will not be pushed to the recommended user; if the recommendation index of the scientific article to be extracted is higher than the preset recommendation index threshold, it will be pushed to the recommended user).
[0027] Historical reading time series data include the page scroll depth value, page focus dwell time value, page scroll rate value, page pause time value, reading time value, reading count value, and reading response time value for each historical period. Historical interaction time series data include interaction conversion rate value, interaction frequency value, and interaction response speed value.
[0028] The page scroll depth value is the average value of the proportion of the page content scrolling from the top to the bottom when the user browses each article, expressed as a percentage of the page length, and can be obtained through user behavior analysis tools (such as Hotjar).
[0029] The page focus dwell time value is the average time that users stay when browsing each article. It focuses on the dwell time in specific areas, such as article titles, pictures, paragraphs, etc., and can be obtained through user behavior analysis tools (such as Hotjar).
[0030] The page scrolling rate value is the average scrolling speed of users when browsing each article. A slower scrolling speed means that the user reads carefully, and a faster scrolling speed means that the user browses the content quickly. It can be obtained through user behavior analysis tools (such as Hotjar).
[0031] The page pause duration value is the average length of time that users stay on each article without scrolling or interacting. It indicates that users may be thinking, reading, or viewing content in a certain area. It can be obtained through user behavior analysis tools (such as Hotjar).
[0032] The reading response time is the average time it takes for the page to fully respond after the user browses each article, clicks or scrolls the page. It can be obtained through user behavior analysis tools (such as Hotjar).
[0033] The interactive conversion rate is the proportion of users who perform certain interactive behaviors (such as likes, comments, shares, click links, etc.) after browsing the content within a time period, that is, the interactive conversion rate = actual number of interactive behaviors / number of page views, and the actual number of interactive behaviors and page views can be obtained through user behavior analysis tools (such as: Hotjar).
[0034] The interaction frequency value is the sum of the number of times a user interacts with a topic or content within a time period (such as collecting, liking, and sharing), and the number of each type of interaction can be obtained through user behavior analysis tools (such as Hotjar).
[0035] The interactive response speed value is the average time interval from when a user initiates an interactive behavior (such as clicking, sliding, scrolling, inputting, etc.) to when the interactive behavior generates a response or feedback within a time period. It can be obtained through user behavior analysis tools (such as Hotjar).
[0036] Specifically, the classification model is a bidirectional encoder representation model (BERT), which includes an input embedding layer, an encoder, a classification layer, and an output layer, and the specific process of obtaining the correlation index between each scientific article to be extracted and each scientific topic is as follows: in the input embedding layer of the bidirectional encoder representation model, the text information of each scientific article to be extracted is converted (that is, the input text is converted into a vector representation that the model can understand), and the initial feature representation vector of each scientific article to be extracted (including word semantic vectors, position representation vectors, and sentence distinction vectors) is obtained; in the encoder of the bidirectional encoder representation model, the initial feature representation vector of each scientific article to be extracted is deeply processed (that is, a multi-head self-attention mechanism is used to calculate the dependency relationship between each word and all other words in the article, thereby capturing global semantics and context information; a feedforward neural network is used to perform nonlinear characterization of each word. Transformation to enhance its feature expression ability; residual connection to retain input feature information and alleviate the gradient disappearance problem; layer normalization to standardize the output of each layer to ensure the stability of training) to obtain the (context-related) representation vector set of each scientific article to be extracted (that is, the semantic and contextual information of the words in the initial feature representation vector in the context of the article); in the classification layer of the bidirectional encoder representation model, the context-related representation vector set of each scientific article to be extracted is classified to obtain the correlation probability of each scientific article to be extracted and each scientific topic; in the output layer of the bidirectional encoder representation model, the correlation probability of each scientific article to be extracted and each scientific topic is judged and analyzed with the preset correlation probability threshold (that is, if it is lower than or equal to the preset correlation probability threshold, it is discarded, otherwise it is retained and marked as a correlation index) to obtain the correlation index of each scientific article to be extracted and each scientific topic.
[0037] The specific steps of the classification operation are as follows: the context-related representation vector set is pooled (such as average pooling or maximum pooling) to generate an article-level representation (a high-dimensional vector), and a fully connected layer is used to reduce the dimension of the article-level representation to the feature space required for the classification task, and then an activation function is applied to convert the score into an interpretable probability value, namely the relevance index.
[0038] In this implementation, the use of the bidirectional encoder representation model (BERT) can fully mine the deep semantic information in the article text, and through multiple processing of the input embedding layer, encoder and classification layer, BERT can capture the dependency between each word and other words through the multi-head self-attention mechanism, so as to better understand the contextual information in the article, thereby comprehensively understanding the semantics and context of the text, thereby improving the in-depth analysis and accurate recognition of the article content, and combining the context-related representation vector of the BERT model and the correlation probability output by the classification layer, so as to achieve highly accurate matching of scientific articles with various scientific and technological topics, and at the same time generate article-level high-dimensional vectors through pooling operations, and further use the fully connected layer and activation function to convert them into correlation indexes, so as to ensure that the recommendation of each article is highly consistent with user interests.
[0039] Specifically, the pre-training process of the classification model is as follows: obtain several groups of training text information and divide them into text information training sets and text information verification sets; divide the text information training sets into several batch training sets, and each batch training set contains several text information; perform forward propagation processing on each batch training set to obtain the representation vector set of each text information in each batch training set, and calculate the text loss function; perform iterative training at the same time, evaluate each training result of the classification model based on the back propagation algorithm and the text information verification set, and adjust the model parameters according to the verification results until the model meets the expected standards.
[0040] In this implementation scheme, by dividing the training data into multiple batch training sets for forward propagation processing, the model can gradually optimize its learning process, thereby avoiding the overfitting problem that may occur in a single training process, and the representation vector set of each batch training set can effectively transmit feature information, so that the classification model can better understand the semantic features of different texts. In addition, batch training can also effectively improve training efficiency and reduce the consumption of computing resources. At the same time, by using the back propagation algorithm and the text information verification set for iterative training, the classification model can adjust parameters based on the verification results after each training, so that the model can adaptively optimize its parameters and continuously improve the classification performance, thereby ensuring that the final model's performance in actual applications meets the expected standards, thereby enabling the classification model to continuously adapt to new data features and enhance its prediction ability for unknown data.
[0041] Specifically, the specific steps for obtaining the historical reading preference index of each historical period of each scientific and technological theme of the user to be recommended are as follows: normalize (i.e., remove the unit) the page scroll depth value, page focus dwell time value, page scroll rate value, page pause duration value, reading time value, number of readings value, and reading response duration value of each historical period of each scientific and technological theme of the user to be recommended; comprehensively analyze the normalized page scroll depth value, page focus dwell time value, page scroll rate value, and page pause duration value of each historical period of each scientific and technological theme of the user to be recommended to obtain the historical attraction index of each historical period of each scientific and technological theme of the user to be recommended; and comprehensively analyze the normalized reading time value, number of readings value, and reading response duration value of each historical period of each scientific and technological theme of the user to be recommended to obtain the historical reading investment index of each historical period of each scientific and technological theme of the user to be recommended; and comprehensively analyze the historical attraction index and historical reading investment index of each historical period of each scientific and technological theme of the user to be recommended to obtain the historical reading preference index of each historical period of each scientific and technological theme of the user to be recommended.
[0042] The specific formula for calculating the historical attraction index, historical reading investment index, and historical reading preference index for each historical period of each science and technology topic for the recommended user is as follows: Among them, XyZ ij YgS′ is the historical attraction index of the jth historical period of the i-th technology theme of the recommended user, ij is the page scroll depth value of the jth historical period of the i-th technology topic of the recommended user after normalization, α 1 is the depth coefficient stored in the database, YmS′ ij is the page focus dwell time value of the i-th technology topic of the recommended user in the j-th historical period after normalization, α 2 is the focal coefficient stored in the database, YdS′ ij is the page scrolling rate value of the jth historical period of the i-th technology topic of the recommended user after normalization, α 3 is the rate coefficient stored in the database, YtS′ ij is the normalized page pause duration of the jth historical period of the i-th technology topic of the recommended user, α 4 is the pause coefficient stored in the database, α 1 +α 2 +α 3 +α 4 =1, e is a natural constant, and in this embodiment, its value is 2.71, TyD ijYdS′ is the historical reading investment index of the jth historical period of the i-th scientific and technological topic of the recommended user, ij is the normalized reading time value of the jth historical period of the i-th science and technology topic of the recommended user, β 1 is the reading time coefficient stored in the database, YdS′ ij is the normalized reading count of the i-th technology topic in the j-th historical period of the recommended user, β 2 is the reading frequency coefficient stored in the database, YdS′ ij is the normalized reading response time value of the jth historical period of the i-th scientific and technological topic of the recommended user, β 3 is the response time coefficient stored in the database, β 1 +β 2 +β 3 =1,YdX ij is the historical reading preference index of the jth historical period of the i-th scientific and technological topic of the recommended user, δ 1 is the attraction coefficient stored in the database, δ 2 is the input coefficient stored in the database, δ 1 +δ 2 =1,i=1,2,3,…,i 0 ,i 0 is the number of scientific and technological subject categories, j = 1, 2, 3, ..., j 0 , j 0 is the number of historical time periods.
[0043] It needs to be explained that α 1 , α 2 , α 3 , α 4 The specific acquisition process is: read the page scroll depth value, page focus dwell time value, page scroll rate value, page pause duration value of each historical period of each technology theme to be recommended to the user after normalization, and perform mean analysis, and perform sum analysis based on the mean analysis results to obtain the attraction sum value, and perform proportion analysis on the mean analysis results and the attraction sum value respectively, and use the proportion analysis results as the corresponding coefficient.
[0044] β 1 , β 2 , β 3 The specific acquisition process is: read the normalized reading time value, reading frequency value, and reading response time value of each historical period of each science and technology topic for the recommended user, perform mean analysis, and perform sum analysis based on the mean analysis results to obtain the input and value, and perform proportion analysis on the mean analysis results and the input and value respectively, and use the proportion analysis results as the corresponding coefficients.
[0045] δ 1 ,δ 2 The specific acquisition process is as follows: perform mean analysis on the historical attraction index and historical reading investment index (both the historical attraction index and the historical reading investment index are dimensionless indexes and can be calculated) of each historical period of each scientific and technological theme for the recommended user, and perform sum analysis based on the mean analysis results to obtain the preference and value, and perform proportion analysis on the mean analysis results and the preference and value respectively, and use the proportion analysis results as the corresponding coefficients.
[0046] In this implementation scheme, by analyzing the user's behavioral data in each historical period of each science and technology theme, the user's interests and preferences are accurately captured, which helps to make more personalized recommendations on science and technology theme content that suits the user's taste, thereby improving the accuracy of the recommendation. Secondly, the historical attraction index and historical reading investment index, which are comprehensively evaluated by multiple indicators, avoid the limitation of relying on a single data point, thereby comprehensively reflecting the user's attention and participation in different science and technology topics, and through normalization and coefficient analysis, different behavioral indicators can be converted into a unified dimensionless index, which helps to make unified comparisons and optimizations in the recommendation process. In addition, coefficient ratio analysis can provide a more detailed preference index, further improving the accuracy of recommendations. Finally, the comprehensive analysis of the historical reading preference index can recommend content that is more in line with their interests and needs to users, improve user satisfaction and participation, and reduce the interference of irrelevant content.
[0047] Specifically, the specific steps for obtaining the historical interaction preference index of each historical period of each scientific and technological theme of the user to be recommended are as follows: the interaction conversion rate value, interaction frequency value, and interaction reaction speed value of each historical period of each scientific and technological theme of the user to be recommended are standardized (i.e., unit removal); the interaction conversion rate value, interaction frequency value, and interaction reaction speed value of each historical period of each scientific and technological theme of the user to be recommended after the standardized processing are comprehensively analyzed to obtain the historical interaction preference index of each historical period of each scientific and technological theme of the user to be recommended; wherein, the specific formula for calculating the historical interaction preference index of each historical period of each scientific and technological theme of the user to be recommended is as follows: Among them, LhD ij HzP is the historical interaction preference index of the jth historical period of the i-th technology topic of the recommended user, ij ′ is the interaction conversion rate value of the i-th technology topic in the j-th historical period of the recommended user after standardized processing, φ 1 is the conversion factor stored in the database, HdP ij ′ is the interaction frequency value of the i-th technology topic in the j-th historical period of the recommended user after standardized processing, φ 2is the interaction frequency coefficient stored in the database, HsD′ ij is the interaction response speed value of the i-th technology topic of the recommended user in the j-th historical period after standardized processing, φ 3 is the reaction speed coefficient stored in the database, φ 1 +φ 2 +φ 3 =1,i=1,2,3,…,i 0 ,i 0 is the number of scientific and technological subject categories, j = 1, 2, 3, ..., j 0 , j 0 is the number of historical periods, e is a natural constant and its value is 2.71 in this implementation example.
[0048] It needs to be explained that φ 1 ,φ 2 ,φ 3 The specific acquisition process is: read the interaction conversion rate value, interaction frequency value, and interaction reaction speed value of each historical period of each technology theme of the recommended user after standardized processing, and perform mean analysis, and perform sum analysis based on the mean analysis results to obtain the interaction preference and value, and perform proportion analysis on the mean analysis results and the interaction preference and value respectively, and use the proportion analysis results as the corresponding coefficient.
[0049] In this implementation scheme, by standardizing indicators such as interaction conversion rate, interaction frequency, and interaction response speed, different interaction behaviors can be compared on the same scale, which helps to accurately measure users' interaction preferences in different historical time periods and technology topics, thereby improving the personalized accuracy of recommendations. Secondly, multiple interaction indicators are comprehensively analyzed to understand users' interaction tendencies from multiple angles, which helps to reflect users' interest and depth of interaction in specific technology topics. Finally, based on the interaction data of historical time periods and over time, the recommendation strategy is dynamically adjusted based on users' interaction feedback, thereby enhancing the adaptability of recommendations.
[0050] Specifically, the formula for calculating the comprehensive interest preference index of each historical period for each technology topic of the recommended user is as follows: Among them, ZhQ ij YdX is the comprehensive interest preference index of the i-th technology topic of the recommended user in the j-th historical period, ij is the historical reading preference index of the jth historical period of the i-th science and technology topic of the user to be recommended, is the reading coefficient stored in the database, LhD ij is the historical interaction preference index of the i-th technology topic of the recommended user in the j-th historical period, is the interaction coefficient stored in the database, i=1,2,3,…,i 0 ,i 0 is the number of scientific and technological subject categories, j = 1, 2, 3, ..., j 0 , j 0 is the number of historical periods, e is a natural constant and its value is 2.71 in this implementation example.
[0051] It needs to be explained that The specific acquisition process is: read the historical reading preference index and historical interaction preference index of each historical period of each science and technology topic for the recommended user, and perform mean analysis, and perform sum analysis based on the mean analysis results to obtain the comprehensive interest preference and value, and perform proportion analysis on the mean analysis results and the comprehensive interest preference, and use the proportion analysis results as the corresponding coefficient.
[0052] In this implementation, by combining the user's historical reading behavior and historical interactive behavior, the comprehensive interest preference index can more comprehensively and accurately reflect the user's interests and preferences. Secondly, through mean analysis and proportion analysis, the interest preference index is adjusted according to the user's historical behavior and different weights, thereby avoiding the limitations of a single indicator, and flexibly adapting to changes in user behavior for recommendations. Mean analysis and proportion analysis enable multi-dimensional data to be reasonably integrated, thereby evaluating the user's interests from multiple angles and multiple behavioral levels, further improving the recommendation effect.
[0053] Specifically, the specific formula for calculating the recommendation index of each scientific article to be extracted is as follows: Among them, TqZ a is the recommendation index of the a-th scientific article to be extracted, BhZ i is the comprehensive interest preference change index of the i-th technology topic of the recommended user, μ 1 is the comprehensive interest coefficient stored in the database, ZtG ai is the correlation index between the a-th scientific article to be extracted and the i-th scientific topic, μ 2 is the correlation coefficient stored in the database, μ 1 +μ 2 =1,a=1,2,3,…,a 0 , a 0 is the number of scientific articles, i = 1, 2, 3, ..., i 0 ,i 0 is the number of science and technology theme types.
[0054] It needs to be explained that μ 1 , μ 2The specific acquisition process is: read the comprehensive interest preference change index of each scientific and technological topic of the user to be recommended and the correlation index of each scientific and technological article to be extracted and each scientific and technological topic, and perform mean analysis, and perform sum analysis based on the mean analysis results to obtain the recommended sum value, and perform proportion analysis on the mean analysis results and the comprehensive interest preferences respectively, and use the proportion analysis results as the corresponding coefficients.
[0055] In this implementation, the user's interest preference changes are combined with the relevance of the article to ensure that the recommended scientific articles are in line with the user's interest change trend and are highly relevant to the scientific topics that the user is concerned about, which helps to improve the accuracy of the recommendation and make the recommendation more in line with the user's dynamic needs. Secondly, by introducing the "comprehensive interest preference change index", the user's current interests are taken into account, and the changing trends of the user's interests are also reflected, so that the recommendation can identify and adapt to the user's interest changes and make more flexible adjustments. Then, by performing summation and proportion analysis based on mean analysis, different weights are given to the user's interest changes and the relevance of the article, so that the recommendation strategy can be adjusted at any time according to data changes, thereby improving the timeliness and accuracy of the recommendation.
[0056] See also Figure 2 The embodiment of the present invention provides a technical solution: a system for automatic extraction and recommendation of scientific and technological information, including a data acquisition module, a data extraction module, a data analysis module, a recommendation analysis module, and a user recommendation module; the data acquisition module is used to obtain the text information of several scientific and technological articles to be extracted in real time and perform preprocessing; the data extraction module is used to input the preprocessed text information of each scientific and technological article to be extracted into a pre-trained classification model for extraction and analysis, and obtain the correlation index between each scientific and technological article to be extracted and each scientific and technological topic; the data analysis module is used to simultaneously obtain the historical browsing behavior time series data of each scientific and technological topic of the user to be recommended, the historical reading behavior time series data includes historical reading time series data and historical interaction time series data, and perform data analysis separately to obtain the scientific and technological topics to be recommended. The historical reading preference index and historical interaction preference index of each historical period of each scientific theme of the user are comprehensively analyzed to obtain the comprehensive interest preference index of each historical period of each scientific theme of the user to be recommended; the recommendation analysis module is used to perform moving weighted average processing on the comprehensive interest preference index of each historical period of each scientific theme of the user to be recommended to obtain the comprehensive interest preference change index of each scientific theme of the user to be recommended, and conduct a comprehensive analysis in combination with the correlation index between each scientific article to be extracted and each scientific theme to obtain the recommendation index of each scientific article to be extracted; the user recommendation module is used to judge and analyze the recommendation index of each scientific article to be extracted with the preset recommendation index threshold, and push the recommended users based on the judgment and analysis results.
[0057] In summary, this application has at least the following effects:
[0058] By extracting the relevance of scientific articles to each scientific topic in real time and combining it with changes in user historical preferences, recommendations can be made more accurate. The recommendation index is used as a filtering criterion to effectively screen out the scientific articles that best meet user needs for push, thereby helping users quickly obtain scientific information of interest, and increasing the exposure and reading rate of articles, thereby enhancing the dissemination effect and influence of information.
[0059] By analyzing the time series data of users' browsing behavior, we can accurately analyze the changes in users' interests in each technology topic, and use the comprehensive interest preference index to dynamically adjust the recommendation strategy, thereby ensuring that the recommended content is highly consistent with the user's latest interests and enhancing user stickiness. This enables the recommendation system to continue to provide high-quality technology content that meets user needs, thereby improving the user's reading experience.
[0060] Through the collaborative work of various modules, efficient information extraction and intelligent recommendation are achieved. The data acquisition and extraction modules ensure real-time and accuracy, while the data analysis module ensures that the recommended content keeps pace with changes in user needs by analyzing the user's historical behavior time series data. The weighted average processing of the recommendation analysis module and the comprehensive analysis of the correlation index make the recommendation not only accurate, but also dynamically adapt to changes in user interests, thereby significantly improving the efficiency and intelligence of the recommendation, and being able to adjust the recommendation strategy in real time, thereby providing a continuously optimized recommendation service experience.
[0061] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0062] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for automatically extracting and recommending scientific and technological information, characterized in that: The following steps are involved: Real-time acquisition of text information of several scientific articles to be extracted and preprocessing; Input the preprocessed text information of each scientific article to be extracted into a pre-trained classification model for extraction and analysis, and obtain the correlation index between each scientific article to be extracted and each scientific topic; At the same time, the historical browsing behavior time series data of each technology theme of the user to be recommended is obtained, and the historical reading behavior time series data includes historical reading time series data and historical interaction time series data, and data analysis is performed separately to obtain the historical reading preference index and historical interaction preference index of each historical period of each technology theme of the user to be recommended, and a comprehensive analysis is performed to obtain the comprehensive interest preference index of each historical period of each technology theme of the user to be recommended; Perform moving weighted average processing on the comprehensive interest preference index of each scientific and technological theme of each historical period of the recommended user to obtain the comprehensive interest preference change index of each scientific and technological theme of the recommended user, and perform comprehensive analysis on the correlation index between each scientific and technological article to be extracted and each scientific and technological theme to obtain the recommendation index of each scientific and technological article to be extracted; The recommendation index of each scientific article to be extracted is compared with the preset recommendation index threshold for judgment and analysis, and the recommended users are pushed based on the judgment and analysis results.
2. The method for automatic extraction and recommendation of scientific and technological information according to claim 1, characterized in that: The historical reading time series data includes the page scroll depth value, page focus dwell time value, page scroll rate value, page pause time value, reading time value, reading count value, and reading response time value of each historical time period; the historical interaction time series data includes the interaction conversion rate value, interaction frequency value, and interaction reaction speed value.
3. The method for automatic extraction and recommendation of scientific and technological information according to claim 1, characterized in that: The classification model is specifically a bidirectional encoder representation model, which includes an input embedding layer, an encoder, a classification layer, and an output layer, and the specific process of obtaining the relevance index of each scientific article to be extracted and each scientific topic is: In the input embedding layer of the bidirectional encoder representation model, the text information of each scientific article to be extracted is converted to obtain the initial feature representation vector of each scientific article to be extracted; In the encoder of the bidirectional encoder representation model, the initial feature representation vector of each scientific article to be extracted is deeply processed to obtain a representation vector set of each scientific article to be extracted; In the classification layer of the bidirectional encoder representation model, the context-related representation vector set of each scientific article to be extracted is classified to obtain the probability of relevance between each scientific article to be extracted and each scientific topic; In the output layer of the bidirectional encoder representation model, the correlation probability between each scientific article to be extracted and each scientific topic is judged and analyzed with the preset correlation probability threshold to obtain the correlation index between each scientific article to be extracted and each scientific topic.
4. The method for automatic extraction and recommendation of scientific and technological information according to claim 1, characterized in that: The pre-training process of the classification model is: Obtain several sets of training text information and divide them into a text information training set and a text information verification set; The text information training set is divided into a number of batch training sets, and each batch training set contains a number of text information; Perform forward propagation processing on each batch of training sets to obtain the representation vector set of each text information in each batch of training sets, and calculate the text loss function; At the same time, iterative training is carried out to evaluate the training results of each classification model based on the back propagation algorithm and the text information verification set, and the model parameters are adjusted according to the verification results until the model meets the expected standards.
5. The method for automatic extraction and recommendation of scientific and technological information according to claim 2, characterized in that: The specific steps of obtaining the historical reading preference index of each historical period of each science and technology topic for the user to be recommended are as follows: Normalize the page scroll depth value, page focus dwell time value, page scroll rate value, page pause time value, reading time value, reading times value, and reading response time value of each historical period of each technology topic for the recommended user; Comprehensively analyze the page scroll depth value, page focus dwell time value, page scroll rate value, and page pause time value of each historical period of each technology theme of the recommended user after normalization, and obtain the historical attraction index of each historical period of each technology theme of the recommended user; A comprehensive analysis is performed on the normalized reading time value, reading frequency value, and reading response time value of each historical period of each scientific and technological topic of the user to be recommended, to obtain the historical reading investment index of each historical period of each scientific and technological topic of the user to be recommended; A comprehensive analysis is then conducted on the historical attraction index and historical reading investment index for each historical period of each scientific and technological theme for the recommended user, to obtain the historical reading preference index for each historical period of each scientific and technological theme for the recommended user.
6. The method for automatic extraction and recommendation of scientific and technological information according to claim 5, characterized in that: The specific formula for calculating the historical attraction index, historical reading investment index, and historical reading preference index for each historical period of each science and technology topic for the recommended user is as follows: Among them, XyZ ij YgS′ is the historical attraction index of the jth historical period of the i-th technology theme of the recommended user, ij is the page scroll depth value of the jth historical period of the i-th technology topic of the recommended user after normalization, α1 is the depth coefficient stored in the database, YmS′ ij is the page focus dwell time value of the i-th technology topic of the recommended user in the j-th historical period after normalization, α2 is the focus coefficient stored in the database, and YdS′ ij is the page scrolling rate value of the i-th technology topic of the recommended user in the j-th historical period after normalization, α3 is the rate coefficient stored in the database, YtS′ ij is the page pause duration value of the jth historical period of the i-th science and technology topic to be recommended to the user after normalization, α4 is the pause coefficient stored in the database, α1+α2+α3+α4=1, e is a natural constant, TyD ij YdS′ is the historical reading investment index of the jth historical period of the i-th scientific and technological topic of the recommended user, ij is the reading time value of the jth historical period of the i-th science and technology topic of the recommended user after normalization, β1 is the reading time coefficient stored in the database, and YdS′ ij is the reading count value of the i-th technology topic of the recommended user in the j-th historical period after normalization, β2 is the reading count coefficient stored in the database, and YdS′ ij is the reading response time value of the jth historical period of the i-th scientific and technological topic of the recommended user after normalization, β3 is the response time coefficient stored in the database, β1+β2+β3=1, YdX ij is the historical reading preference index of the j-th historical period of the ith science and technology theme of the user to be recommended, δ1 is the attraction coefficient stored in the database, δ2 is the investment coefficient stored in the database, δ1+δ2=1, i=1, 2, 3, …, i0, i0 is the number of science and technology theme types, j=1, 2, 3, …, j0, j0 is the number of historical periods.
7. The method for automatic extraction and recommendation of scientific and technological information according to claim 2, characterized in that: The specific steps for obtaining the historical interaction preference index of each historical period for each technology topic of the user to be recommended are as follows: Standardize the interactive conversion rate, interactive frequency, and interactive response speed values of each historical period for each technology theme of the recommended user; The standardized interaction conversion rate values, interaction frequency values, and interaction response speed values of each historical period of each technological theme of the recommended user are comprehensively analyzed to obtain the historical interaction preference index of each historical period of each technological theme of the recommended user.
8. The method for automatic extraction and recommendation of scientific and technological information according to claim 1, characterized in that: The formula for calculating the comprehensive interest preference index of each historical period for each technology topic of the recommended user is as follows: Among them, ZhQ ij YdX is the comprehensive interest preference index of the i-th technology topic in the j-th historical period of the recommended user, ij is the historical reading preference index of the jth historical period of the i-th science and technology topic of the user to be recommended, is the reading coefficient stored in the database, LhD ij is the historical interaction preference index of the i-th technology topic of the recommended user in the j-th historical period, is the interaction coefficient stored in the database, i=1, 2, 3, …, i0, i0 is the number of scientific and technological subject types, j=1, 2, 3, …, j0, j0 is the number of historical periods, and e is a natural constant.
9. The method for automatic extraction and recommendation of scientific and technological information according to claim 1, characterized in that: The specific formula for calculating the recommendation index of each scientific article to be extracted is as follows: Among them, TqZ a is the recommendation index of the a-th scientific article to be extracted, BhZ i is the comprehensive interest preference change index of the i-th technology topic of the recommended user, μ1 is the comprehensive interest coefficient stored in the database, ZtG ai is the correlation index between the a-th scientific article to be extracted and the i-th scientific topic, μ2 is the correlation coefficient stored in the database, μ1+μ2=1, a=1, 2, 3, …, a0, a0 is the number of scientific articles, i=1, 2, 3, …, i0, i0 is the number of scientific topic types.
10. A system for automatic extraction and recommendation of scientific and technological information, using the method for automatic extraction and recommendation of scientific and technological information according to any one of claims 1 to 9, characterized in that: include: Data acquisition module, data extraction module, data analysis module, recommendation analysis module, user recommendation module; The data acquisition module is used to acquire text information of several scientific articles to be extracted in real time and perform preprocessing; The data extraction module is used to input the pre-processed text information of each scientific article to be extracted into a pre-trained classification model for extraction and analysis, and obtain the correlation index between each scientific article to be extracted and each scientific topic; The data analysis module is used to simultaneously obtain the historical browsing behavior time series data of each science and technology theme of the user to be recommended, the historical reading behavior time series data includes historical reading time series data and historical interaction time series data, and perform data analysis respectively to obtain the historical reading preference index and historical interaction preference index of each historical period of each science and technology theme of the user to be recommended, and perform comprehensive analysis to obtain the comprehensive interest preference index of each historical period of each science and technology theme of the user to be recommended; The recommendation analysis module is used to perform moving weighted average processing on the comprehensive interest preference index of each scientific and technological theme of each historical period of the recommended user to obtain the comprehensive interest preference change index of each scientific and technological theme of the recommended user, and to perform comprehensive analysis on the correlation index between each scientific and technological article to be extracted and each scientific and technological theme to obtain the recommendation index of each scientific and technological article to be extracted; The user recommendation module is used to judge and analyze the recommendation index of each scientific article to be extracted and the preset recommendation index threshold, and push the recommended articles to the recommended users based on the judgment and analysis results.
Citation Information
Patent Citations
A method and system for automatically extracting and recommending science and technology policy information
CN117743564B
Eye movement tracking-based text recommendation method
CN106897363A
Project recommendation method and system based on multiple preference degrees
CN115936939A
Scientific and technological policy information automatic extraction and recommendation method and system
CN117743564A
Method and device for predicting point of interest, electronic equipment, computer readable storage medium and computer program product
CN118296237A