Scientific and technological information online pushing system based on big data
By using big data technology and gradient-improving decision tree algorithms and other technical means in the online push system, the problem of information overload and push accuracy is solved, and personalized and accurate push of scientific and technological information is achieved.
Patent Information
- Application Number
- CN202510055121.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-23
AI Technical Summary
Existing online push systems may cause problems that overload information and push accuracy, especially when user interest preferences change over time.
A technology information online push system based on big data is adopted to generate a personalized push list through user data collection, processing and recommendation algorithms, and a gradient is used to improve the accuracy of push.
It effectively reduces information overload, improves the accuracy and personalization of push, and can update push content in real time to adapt to changes in user behavior data.
Smart Images

Figure CN120030231A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an online scientific and technological information push system based on big data. Background Art
[0002] The online technology information push system based on big data collects various technology information through the Internet, including academic papers, news reports, patent information, technology trends, etc. The system uses big data technology and machine learning algorithms to analyze and mine this information to understand the user's interest preferences and behavior patterns. Then, the system pushes relevant technology information to users based on their personalized needs, helping them quickly obtain the information they need.
[0003] The existing online push system has the following problems: 1. A large amount of scientific and technological information may be pushed to users, and users may face the problem of information overload and find it difficult to filter out truly valuable information; 2. The push system can perform personalized push based on the user's interest preferences, but since the user's interests may change over time, the accuracy of the system's push may be affected.
[0004] To sum up, the present invention relates to an online scientific and technological information push system based on big data. Summary of the invention
[0005] In order to overcome the above-mentioned shortcomings, the present invention provides an online scientific and technological information push system based on big data.
[0006] The present invention achieves the above-mentioned purpose through the following technical solutions:
[0007] A scientific and technological information online push system based on big data, comprising the following steps:
[0008] Step 1: User data collection, which is used to collect user behavior data, interests and hobbies, and technological information data;
[0009] Step 2: User data processing: cleaning, integrating and analyzing the data collected in step 1;
[0010] Step 3: Data recommendation processing: after the data processed in step 2 is processed, a recommendation algorithm is used to generate a personalized list of scientific and technological information push for the user;
[0011] Step 4: User push interaction: display the push information in the push list of step 3 to the user;
[0012] Among them, in the data recommendation processing step, the gradient boosting decision tree algorithm is used, and the gradient boosting decision tree algorithm includes the following steps:
[0013] S31, data division, dividing the data in the push list in step 3 into a training set and a validation set;
[0014] S32, initial model training, using the training set to train a basic decision tree model as the initial model;
[0015] S33, residual calculation, calculate the residual on the training set, that is, the difference between the true value and the initial model prediction value;
[0016] S34, new model training, using the residual as the new target variable to train a new decision tree model;
[0017] S35, model fusion, weighted fusion of the new model and the initial model to obtain a new prediction result;
[0018] S36, iterative training, repeating the above steps until the predetermined number of iterations is reached or the performance of the model on the validation set begins to decline;
[0019] S37, generating a push list.
[0020] Preferably, the step 1 comprises the following steps:
[0021] S11, distributed data collection, is used to collect distributed information of users. Kafka is used to collect distributed information of users. Multiple data producers send data to the Kafka cluster, and multiple data consumers are allowed to subscribe to and obtain data of interest from the Kafka cluster. Data producers include web crawlers, API interfaces, and IoT sensors, and data consumers include Sqoop and Canal tools.
[0022] S12, establishing a relational database, for performing data association collection based on the distributed data collected in S11, and establishing a relational database, wherein Sqoop performs one-to-many collection of the distributed data in S11 and transmits it to the relational database;
[0023] S13, database information change tracking, is used to track relational database change information in real time, import real-time data increments through Canal, synchronize relational database change information in real time, and record the type of change information. For example, if the user's behavior data increases by more than the set value, it will be marked in step three to re-update the user's recommendation data.
[0024] Preferably, the step 2 comprises the following steps:
[0025] S21, data cleaning, used to clean the data in step 1;
[0026] S22, data extraction, extracting the required data from the database in step 1;
[0027] S23, data conversion, converting data into a format that can be stored and used in a data warehouse;
[0028] S24: Data loading: loading the data converted in step S23 into a new database.
[0029] Preferably, the step S21 includes removing duplicate data, processing missing values and correcting erroneous data, and using a data cleaning tool to identify errors, missing values and outliers in the data of step one.
[0030] Preferably, in step S12, the relational database includes a direct relational database and an indirect relational database. The data imported into the direct relational database is obtained by summarizing the data collected in step one, and the indirect relational database is obtained by directly accessing the database information of the cooperation platform and retrieving the relevant data information of the user.
[0031] Preferably, in step three, the recommendation algorithm also includes a content recommendation algorithm and a knowledge recommendation algorithm. The data after processing in step two is first subjected to priority judgment, and then the data is accurately processed by a gradient boosting decision tree algorithm, that is, the priority is selected according to the amount of user behavior data and interest data collected, and the content recommendation algorithm and the knowledge recommendation algorithm are used to recommend content related to the user's historical behavior and interests. It usually generates recommendation results by analyzing user behavior data (such as browsing, clicking, purchasing) and content features (such as titles, descriptions, and tags). Knowledge recommendation uses technologies such as knowledge graphs to derive appropriate recommendations through reasoning in complex scenarios. The knowledge graph is a structured knowledge base that contains a large amount of information such as entities, attributes, and relationships. By mining the information in the knowledge graph, the recommendation system can recommend more accurate and personalized content to users.
[0032] Preferably, the step 4 further includes collecting user feedback data, the feedback data including the number of user clicks, click frequency and browsing time, so as to optimize and improve the system, and the data is incorporated into the verification set in step S31.
[0033] The beneficial effects of the present invention are as follows: In the online scientific and technological information push system based on big data:
[0034] 1. In step 1, distributed data collection is first performed on user data, and then relational data collection is performed based on the distributed data to increase the breadth of user data collection. Finally, database information change tracking is added to track changes in user behavior data in real time, and push content can be updated in real time based on user behavior data;
[0035] 2. In step 3, the content recommendation algorithm and the knowledge recommendation algorithm are added with priority judgment to improve the accuracy of push notifications. Combined with the gradient boosting decision tree algorithm, user data can be calculated for accurate push notifications.
[0036] 3. In step 4, user feedback data collection is added and then fed back to step 3 to improve the accuracy of the gradient boosting decision tree algorithm and the accuracy of user push. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The present invention will now be described by way of example with reference to the accompanying drawings, in which:
[0038] Figure 1 It is a system step diagram of the present invention;
[0039] Figure 2 is a step diagram of step 1 of the present invention;
[0040] Figure 3 is a step diagram of step 2 of the present invention;
[0041] Figure 4 It is a step diagram of step three of the present invention. DETAILED DESCRIPTION
[0042] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0043] like Figure 1 and Figure 4 As shown, a scientific and technological information online push system based on big data includes the following steps:
[0044] Step 1: User data collection, which is used to collect user behavior data, interests and hobbies, and technological information data;
[0045] Step 2: User data processing: cleaning, integrating and analyzing the data collected in step 1;
[0046] Step three: data recommendation processing. For the data processed in step two, a recommendation algorithm is used to generate a personalized science and technology information push list for the user.
[0047] In step three, the recommendation algorithm also includes a content recommendation algorithm and a knowledge recommendation algorithm. The data processed in step two is first subjected to priority judgment, and then the data is accurately processed by the gradient boosting decision tree algorithm, that is, the priority is selected according to the collected user behavior data and interest data. The content recommendation algorithm and the knowledge recommendation algorithm are used to recommend content based on the user's historical behavior and interests, and recommend related content. It usually generates recommendation results by analyzing user behavior data (such as browsing, clicking, purchasing) and content features (such as title, description, and tags). Knowledge recommendation uses technologies such as knowledge graphs to derive appropriate recommendations through reasoning in complex scenarios. The knowledge graph is a structured knowledge base that contains a large amount of information such as entities, attributes, and relationships. By mining the information in the knowledge graph, the recommendation system can recommend more accurate and personalized content to users.
[0048] Step 4, user push interaction, displays the push information in the push list of step 3 to the user, and also includes collecting user feedback data, which includes the number of user clicks, click frequency and browsing time, so as to optimize and improve the system. The data is integrated into the verification set in step S31.
[0049] Among them, in the data recommendation processing step, the gradient boosting decision tree algorithm is used, and the gradient boosting decision tree algorithm includes the following steps:
[0050] S31, data division, dividing the data in the push list in step 3 into a training set and a validation set, and the validation set also includes the feedback data in step 4;
[0051] S32, initial model training, using the training set to train a basic decision tree model as the initial model;
[0052] S33, residual calculation, calculate the residual on the training set, that is, the difference between the true value and the initial model prediction value;
[0053] S34, new model training, using the residual as the new target variable to train a new decision tree model;
[0054] S35, model fusion, weighted fusion of the new model and the initial model to obtain a new prediction result;
[0055] S36, iterative training, repeating the above steps until the predetermined number of iterations is reached or the performance of the model on the validation set begins to decline;
[0056] S37, generating a push list.
[0057] like Figure 2 As shown, as a specific embodiment, the step one includes the following steps:
[0058] S11, distributed data collection, is used to collect distributed information of users. Kafka is used to collect distributed information of users. Multiple data producers send data to the Kafka cluster, and multiple data consumers are allowed to subscribe to and obtain data of interest from the Kafka cluster. Data producers include web crawlers, API interfaces, and IoT sensors, and data consumers include Sqoop and Canal tools.
[0059] S12, establishing a relational database, for performing data association collection based on the distributed data collected in S11, and establishing a relational database, wherein Sqoop performs one-to-many collection of the distributed data in S11 and transmits it to the relational database;
[0060] S13, database information change tracking, is used to track the relational database change information in real time, import the real-time data increment through Canal, synchronize the relational database change information in real time, and record the type of change information. For example, if the user's behavior data increases by more than the set value, it will be marked in step 3 to re-update the user's recommendation data.
[0061] In the step S12, the relational database includes a direct relational database and an indirect relational database. The data imported into the direct relational database is obtained by summarizing the data collected in step 1, and the indirect relational database is obtained by directly accessing the database information of the cooperation platform and retrieving the relevant data information of the user.
[0062] like Figure 3 As shown, as a specific embodiment, the step 2 includes the following steps:
[0063] S21, data cleaning, is used to clean the data in step 1. Step S21 includes removing duplicate data, processing missing values and correcting erroneous data. The data cleaning tool can be used to identify errors, missing values and abnormal values in the data in step 1.
[0064] S22, data extraction, extracting the required data from the database in step 1;
[0065] S23, data conversion, converting data into a format that can be stored and used in a data warehouse;
[0066] S24: Data loading: loading the data converted in step S23 into a new database.
[0067] The above is based on the present invention as an inspiration. Through the above description, relevant staff can make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A scientific and technological information online push system based on big data, characterized by: The following steps are involved: Step 1: User data collection, which is used to collect user behavior data, interests and hobbies, and technological information data; Step 2: User data processing: cleaning, integrating and analyzing the data collected in step 1; Step 3: Data recommendation processing: after the data processed in step 2 is processed, a recommendation algorithm is used to generate a personalized list of scientific and technological information push for the user; Step 4: User push interaction: display the push information in the push list of step 3 to the user; Among them, in the data recommendation processing step, the gradient boosting decision tree algorithm is used, and the gradient boosting decision tree algorithm includes the following steps: S31, data division, dividing the data in the push list in step 3 into a training set and a validation set; S32, initial model training, using the training set to train a basic decision tree model as the initial model; S33, residual calculation, calculate the residual on the training set, that is, the difference between the true value and the initial model prediction value; S34, new model training, using the residual as the new target variable to train a new decision tree model; S35, model fusion, weighted fusion of the new model and the initial model to obtain a new prediction result; S36, iterative training, repeating the above steps until the predetermined number of iterations is reached or the performance of the model on the validation set begins to decline; S37, generating a push list.
2. According to the big data-based online scientific and technological information push system of claim 1, it is characterized by: The step 1 includes the following steps: S11, distributed data collection, used to collect distributed information of users (using Kafka to collect distributed information of users, multiple data producers send data to the Kafka cluster, and allow multiple data consumers to subscribe and obtain data of interest from the Kafka cluster. Data producers include web crawlers, API interfaces, and IoT sensors, and data consumers include Sqoop and Canal tools); S12, establish a relational database, which is used to perform data association collection based on the distributed data collected in S11, and establish a relational database (wherein Sqoop performs one-to-many collection of the distributed data in S11 and transmits it to the relational database; S13, database information change tracking, used to track relational database change information in real time (import real-time data increments through Canal, synchronize relational database change information in real time, and record the type of change information. For example, if the user's behavior data increases by more than the set value, it will be marked in step three to re-update the user's recommendation data).
3. According to the big data-based online scientific and technological information push system of claim 1, it is characterized by: The step 2 includes the following steps: S21, data cleaning, used to clean the data in step 1; S22, data extraction, extracting the required data from the database in step 1; S23, data conversion, converting data into a format that can be stored and used in a data warehouse; S24: Data loading: loading the data converted in step S23 into a new database.
4. According to the big data-based online scientific and technological information push system of claim 3, it is characterized by: The step S21 includes removing duplicate data, processing missing values and correcting erroneous data (using data cleaning tools to identify errors, missing values and outliers in the data of step one).
5. According to the big data-based online scientific and technological information push system of claim 2, it is characterized by: In the step S12, the relational database includes a direct relational database and an indirect relational database. The data imported into the direct relational database is obtained by summarizing the data collected in step 1, and the indirect relational database is obtained by directly accessing the database information of the cooperation platform and retrieving the relevant data information of the user.
6. The online scientific and technological information push system based on big data according to claim 1 is characterized by: In step three, the recommendation algorithm also includes a content recommendation algorithm and a knowledge recommendation algorithm. The data processed in step two is first subjected to priority judgment, and then the data is accurately processed by a gradient boosting decision tree algorithm, that is, the priority is selected according to the amount of user behavior data and interest data collected, and the content recommendation algorithm and the knowledge recommendation algorithm are used to recommend content related to the user's historical behavior and interests. It usually generates recommendation results by analyzing user behavior data (such as browsing, clicking, purchasing, etc.) and content features (such as titles, descriptions, tags, etc.). Knowledge recommendation uses technologies such as knowledge graphs to make appropriate recommendations through reasoning in complex scenarios. Knowledge graph is a structured knowledge base that contains a large amount of information such as entities, attributes, and relationships. By mining the information in the knowledge graph, the recommendation system can recommend more accurate and personalized content to users.
7. The online push system for scientific and technological information based on big data according to claim 1 is characterized by: The step 4 also includes collecting user feedback data, which includes the number of user clicks, click frequency and browsing time, so as to optimize and improve the system. The data is incorporated into the verification set in step S31.
Citation Information
Patent Citations
Top-N movie recommendation method for performing weighted fusion on selected local models based on random anchor points
CN108763362A
Domain knowledge graph construction method and system based on big data driving
CN109597855A
Model optimization method based on decision tree and recommendation method
CN113326432A
Scientific and technological information push service system based on big data
CN118445479A