Method and system for information notification based on aggregation model
Through an aggregation model-based method, multi-source data is captured and processed in real time, features are extracted using deep learning and natural language processing technology, and an automated notification generation engine is built, which solves the problems of data silos and processing delays in the existing information notification system, and achieves efficient and accurate information notification.
Patent Information
- Application Number
- CN202510222533.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-20
AI Technical Summary
The existing information reporting system faces problems such as data silos, processing delays, and information overload, and lacks the aggregation ability and real-time processing capabilities of multi-source data, resulting in limited time and accuracy of notifications.
Using an aggregation model-based method, multi-source data is captured in real time through network crawlers, features are extracted using deep learning and natural language processing technology, and an automated notification generation engine is built to realize the fusion and real-time processing of multi-source data.
It improves the timeliness and accuracy of information notification, realizes comprehensive integration and automated information notification, and enhances the real-timeness of the system and user feedback capabilities.
Smart Images

Figure CN120179889A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information processing, and particularly relates to a method and system for information notification based on an aggregation model. Background Art
[0002] In the current era of information explosion, quickly and accurately extracting key information from massive data and promptly notifying relevant personnel is the key to improving decision-making efficiency and response speed. However, existing information notification systems often face problems such as data silos, processing delays, and information overload.
[0003] Existing solutions usually rely on a single data source and lack effective information integration capabilities. In addition, the degree of automation of information processing is not high, resulting in limited timeliness and accuracy of notifications. At the same time, existing systems often ignore the importance of user feedback and lack the ability to continuously optimize.
[0004] The existing design solutions have the following disadvantages:
[0005] 1. Single data source: Lack of the ability to aggregate multi-source data, resulting in incomplete information.
[0006] 2. Manual processing: Rely on manual screening and notification of information, with low efficiency.
[0007] 3. Delayed update: The information processing and notification process lacks real-time performance and cannot respond to new situations in a timely manner.
[0008] 4. Lack of user feedback: The system cannot perform statistics and subsequent analysis based on user feedback. Summary of the Invention
[0009] In order to solve the technical problems that the existing technology lacks effective information integration methods and the timeliness and accuracy of notifications are limited, the present invention proposes a method and system for information notification based on an aggregation model. Through the fusion and real-time processing of multi-source data, the timeliness and accuracy of information notification are improved, and information integration is effectively carried out.
[0010] The specific solutions are as follows:
[0011] A method for information notification based on an aggregation model
[0012] S1: Data collection and preprocessing: Use web crawlers or API interfaces to crawl multi-source data from public data sources in real time and clean the multi-source data;
[0013] S2: Feature extraction: Use deep learning to convert the text content of multi-source data into embedding vectors. Use NLP technology to extract and store the metadata of each data item, that is, text features, into the metadata list, and calculate the event occurrence time based on the text content and add it to the metadata list.
[0014] S3: Multi-source data fusion: Construct a multi-source data input set. Based on a large model, create an automated notification generation engine. Input the multi-source data input set into the automated notification generation engine to output an event notification text in a set format. The multi-source data input set includes: a message set constructed based on the text features obtained in step S2, and reconstruct the message set based on the set keywords, message chains, and message weight pairs generated from the message set, and use the reconstructed message set as the multi-source data input set.
[0015] Preferably, the publicly available data sources in S1 include: social media, news sources, and official announcements.
[0016] Preferably, for the method of real-time data scraping in step S1, design a real-time data processing mechanism to reduce information latency and data loss.
[0017] The real-time data processing mechanism integrates a streaming API, a message queue system, and a distributed computing framework.
[0018] For data sources with strong real-time performance such as social media and news sources, use the streaming API to receive data in real time.
[0019] The message queue system buffers the received real-time data stream.
[0020] The distributed computing framework processes the real-time data stream.
[0021] Preferably, the method of scraping data and cleaning data in S1 is as follows:
[0022] S11: Scrape data and label it: The data includes: timestamp, source identifier, and content text.
[0023] S12: Data cleaning: Remove irrelevant information, filter prohibited words, unify case, and perform simplified and traditional font conversion. The irrelevant tags include: HTML tags and advertisements.
[0024] Preferably, the metadata table headers include: source, author, keywords, release time, number of likes, and number of forwards.
[0025] Preferably, the method for calculating the event occurrence time and adding it to the metadata list is as follows: Add "Occurrence Time" to the metadata table header, use NLP technology to extract the event occurrence time mentioned in the data, or calculate the event occurrence time based on the release time and time-related information. If there is no description of the event occurrence time, use the release time to replace the event occurrence time.
[0026] Preferably, the method for constructing the multi-source data input set in step S3 is as follows:
[0027] S31: Construct a message set: Use the text cosine similarity algorithm and text features to calculate the similarity of text information in different data sources, and combine the data from different data sources but with high content similarity to form a message set;
[0028] S32: Construct set keywords: Extract keywords based on the Bert model, and combine the three most frequently occurring keywords in the message set as the set keywords;
[0029] S33: Obtain a message chain: Sort the messages according to the event occurrence time to obtain a message chain sorted by the time line;
[0030] S34: Construct message weights: Assign message weights to each data item according to the source credibility of the data, user interaction, and time freshness, and use the weighted summation algorithm to determine the priority of the information, which will be highlighted in the subsequent notification;
[0031] S35: Reconstruct the message set: Store the set keywords, message chain, and message weights in the message set to form a multi-source data input set.
[0032] Preferably, the method for creating an automated notification generation engine based on a large model and performing information notification in step S3 is as follows:
[0033] A31: Construct the large model input: The large model input is a multi-source data set, that is, the reconstructed message set in the streaming API. The reconstructed message set includes set keywords, and each message in the message set contains the occurrence time, information source, and information weight;
[0034] A32: Set the large model output format: Set a built-in prompt to let the large model transform the message information into a specified output format. The output format includes: sorting by time sequence, outputting the time first and then the message, and bolding the information with high weight;
[0035] A33: Set the large model working mode: When the user inputs an instruction, the large model recognizes the user keywords, matches them with the set keywords, hits the set, and outputs the event message text in the output format set in A32;
[0036] A34: Call the large model interface to access the output platform and notify the user of the event message text.
[0037] Preferably, after the user receives the event message text, the user chooses to provide feedback through the notification path via email. The feedback email will be recorded and submitted to the application large model that generates reports. The large model will merge and optimize the problems in the feedback email with the user feedback record and then output it to the user. The system will automatically generate a feedback daily report for the engineer to optimize the system.
[0038] Preferably, the method includes step S4 of setting security and privacy protection measures during the information notification process;
[0039] A41: During the data cleaning process, desensitize sensitive data;
[0040] A42: Implement a real-time data encryption policy during data transmission and storage;
[0041] A43: Use an open-source and locally deployable large model application development platform.
[0042] A system for information notification based on an aggregation model, including: a data collection and preprocessing module, a feature extraction and representation module, a multi-source data fusion module, and an automated notification engine module;
[0043] The data collection and preprocessing module includes: a multi-source data acquisition sub-module that obtains multi-source data through web crawlers, and a data preprocessing sub-module that removes irrelevant information, filters prohibited words, and processes text formats;
[0044] The feature extraction and representation module includes: a text conversion sub-module that converts the text of the multi-source data into vector representations, a metadata acquisition sub-module that extracts metadata and deletes privacy information, and an occurrence time acquisition sub-module that uses NLP technology to extract the event occurrence time;
[0045] The multi-source data fusion module includes: a message set construction sub-module that calculates the similarity of text information in different data sources based on cosine similarity and constructs a message set, a set keyword acquisition sub-module that selects the three keywords with the largest quantity in the set as set keywords, and a message weight acquisition sub-module that calculates the message weight based on data source, reposts, comments, likes, and time;
[0046] The automated notification engine module includes: a large model input sub-module that inputs the user's question into the large model development platform, inputs the data of the multi-source data fusion module into the large model development platform and sets the message set format, a message set matching sub-module that matches the keywords in the user's question with the keywords in the message set to obtain a matching result, and a large model output sub-module that outputs the message set in a specified format as the event process based on the matching result and synchronously sends it to the email.
[0047] Beneficial effects:
[0048] The present invention proposes an information notification method and system based on an aggregation model. This method collects and cleans information from multiple data sources in real time, extracts features using deep learning and natural language processing technologies, and finally constructs an automated notification generation engine based on a large model to achieve real-time and accurate message notification results. The present invention integrates a real-time data processing mechanism, an automated notification generation engine, security and privacy protection measures, and a user feedback mechanism, thereby realizing automatic message notification. The present invention can process and integrate similar information from different data sources such as social media, news sources, and official announcements, providing a comprehensive perspective for the same event; it can also sort different news of the same event according to information such as timestamps, connecting the timeline of the event; in addition, it can also perform information importance assessment to determine the priority of information.
[0049] First, the present invention constructs a unified data view by integrating multiple data sources, thereby supporting a more comprehensive and accurate analysis and decision-making process.
[0050] Secondly, the present invention proposes a real-time data processing mechanism to ensure the rapid collection, efficient processing, and timely notification of information. This mechanism relies on an efficient data stream processing framework to minimize the latency of information processing. At the same time, in order to cope with the challenges of a large number of data sources, the present invention adopts a message queue system to avoid data loss.
[0051] In addition, the present invention has also developed an automated notification generation engine that can automatically generate clear and accurate notification texts based on the integrated information and preset templates.
[0052] Thirdly, throughout the entire process of information processing and notification, the present invention implements strict security and privacy protection measures to ensure that the security and privacy of user data are fully protected.
[0053] Fourth, in the present invention, the process of converting message data into embedding vectors makes full use of deep learning techniques. This conversion process provides effective feature representations for subsequent tasks (such as text classification, sentiment analysis, dialogue systems, etc.), significantly improves the efficiency of data processing, and enhances the model's ability to capture complex semantic relationships. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 Flowchart of a method for information notification based on an aggregation model in an embodiment.
[0055] Figure 2 Structural diagram of a system for information notification based on an aggregation model in an embodiment.
[0056] Figure 3 Information transmission flowchart of a system for information notification based on an aggregation model in an embodiment.
[0057] Figure 4 Information transmission flowchart from the multi-source data aggregation subsystem to the automated notification engine module in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0059] As Figure 1 shown, a method for information notification based on an aggregation model,
[0060] S1: Data collection and preprocessing: Use web crawlers or API interfaces to real-time crawl multi-source data from public data sources and clean the multi-source data;
[0061] S2: Feature extraction: Use deep learning to convert the text content of multi-source data into embedding vectors, use NLP techniques to extract and store the metadata of each data item, i.e., text features, into a metadata list, and add the event occurrence time calculated based on the text content to the metadata list;
[0062] S3: Multi-source data fusion: Construct a multi-source data input set, create an automated notification generation engine based on a large model, input the multi-source data input set into the automated notification generation engine, and output an event notification text in a set format; the multi-source data input set includes: a message set constructed based on the text features obtained in step S2, and reconstruct the generated message set based on the set keywords, message chains, and message weight pairs of the message set, and use the reconstructed message set as the multi-source data input set.
[0063] Preferably, the public data sources described in S1 include: social media, news sources, official announcements.
[0064] Preferably, for the method of real-time data capture in step S1, a real-time data processing mechanism is designed to reduce information latency and data loss;
[0065] The real-time data processing mechanism integrates a streaming API, a message queue system, and a distributed computing framework;
[0066] For data sources with strong real-time characteristics such as social media and news feeds, the streaming API is used to receive data in real time;
[0067] The message queue system buffers the received real-time data stream;
[0068] The distributed computing framework processes the real-time data stream.
[0069] Preferably, the method of capturing data and cleaning the data in S1 is as follows:
[0070] S11: Capture data and tag it: The data includes: timestamp, source identifier, content text;
[0071] S12: Data cleaning: Remove irrelevant information, filter prohibited words, unify case, and perform simplified and traditional font conversion; The irrelevant tags include: HTML tags, advertisements.
[0072] Preferably, the metadata headers include: source, author, keywords, release time, number of likes, and number of forwards.
[0073] Preferably, the method of adding the calculated event occurrence time to the metadata list is: Add "occurrence time" to the metadata headers, use NLP technology to extract the occurrence time mentioned in the data, or calculate the event occurrence time based on the release time and time-related information. If there is no description of any occurrence time, the release time is used to replace the event occurrence time.
[0074] Preferably, the method of constructing a multi-source data input set in step S3 is as follows:
[0075] S31: Construct a message set: Use the text cosine similarity algorithm and text features to calculate the similarity of text information in different data sources, and combine data from different data sources but with high content similarity to form a message set;
[0076] S32: Construct set keywords: Extract keywords based on the Bert model, and combine the three most frequently occurring keywords in the message set as set keywords;
[0077] S33: Obtain a message chain: Sort the messages according to the event occurrence time to obtain a message chain sorted by the time line;
[0078] S34: Construct message weights: Assign message weights to each data item based on the source credibility of the data, user interaction, and time freshness, and use the weighted summation algorithm to determine the priority of the information, which will be highlighted in subsequent notifications.
[0079] S35: Reconstruct the message set: Store the set keywords, message chains, and message weights in the message set to form a multi-source data input set.
[0080] Preferably, the method of creating an automated notification generation engine based on a large model and performing information notifications in step S3 is as follows:
[0081] A31: Construct the large model input: The large model input is a multi-source data set, that is, the reconstructed message set in the streaming API. The reconstructed message set includes set keywords, and each message in the message set contains the occurrence time, information source, and information weight.
[0082] A32: Set the large model output format: Set the built-in prompt to enable the large model to transform the message information into a specified output format. The output format includes: sorting by time sequence, outputting the time first and then the message, and bolding the messages with high information weights.
[0083] A33: Set the large model working mode: When the user inputs an instruction, the large model recognizes the user keywords, matches them with the set keywords, hits the set, and outputs the event message text in the output format set in A32.
[0084] A34: Call the large model interface to access the output platform and notify the user of the event message text.
[0085] Preferably, after the user receives the event message text, the user chooses to provide feedback through the notification path via email. The feedback email will be recorded and submitted to the large model for generating reports. The large model will merge and optimize the problems in the feedback email with the user feedback records and then output them to the user. The system will automatically generate a feedback daily report for the engineers to optimize the system.
[0086] Preferably, the method includes step S4 of setting security and privacy protection measures during information notification.
[0087] A41: During data cleaning, perform desensitization processing on sensitive data.
[0088] A42: Implement real-time data encryption policies during data transmission and storage.
[0089] A43: Use an open-source, locally deployable large model application development platform.
[0090] Such as Figure 2 And3 As shown in the figure, a system for information notification based on an aggregation model includes: a data collection and preprocessing module, a feature extraction and representation module, a multi-source data fusion module, and an automated notification engine module;
[0091] The data collection and preprocessing module includes: a multi-source data acquisition sub-module that obtains multi-source data through a web crawler, and a data preprocessing sub-module that removes irrelevant information, filters prohibited words, and processes text formats;
[0092] The feature extraction and representation module includes: a text conversion sub-module that converts the text of the multi-source data into a vector representation, a metadata acquisition sub-module that extracts metadata and deletes privacy information, and an occurrence time acquisition sub-module that uses NLP technology to extract the event occurrence time;
[0093] The multi-source data fusion module includes: a message set construction sub-module that calculates the similarity of text information in different data sources based on cosine similarity and constructs a message set, a set keyword acquisition sub-module that selects the three keywords with the largest number in the set as set keywords, and a message weight acquisition sub-module that calculates the message weight based on data source, reposts, comments, likes, and time;
[0094] The automated notification engine module includes: a large model input sub-module that inputs the user's question into the large model development platform, inputs the data of the multi-source data fusion module into the large model development platform, and sets the message set format, a message set matching sub-module that matches the keywords in the user's question with the keywords in the message set to obtain a matching result, and a large model output sub-module that outputs the message set in a specified format as the event process and synchronously sends it to the email.
[0095] As Figure 4 shown in the figure, the multi-source data aggregation subsystem consists of a data collection and preprocessing module, a feature extraction and representation module, and a multi-source data fusion module. In the multi-source data aggregation subsystem, first, multi-source data is obtained through a web crawler, and then a message set of a certain event is extracted based on the multi-source data fusion algorithm. The message set includes: message keywords, message priority sorting, message occurrence time, etc.
[0096] The multi-source data aggregation subsystem is connected to the input end of the automated notification engine module, and the message set is connected to the large model development platform. The large model converts the message set into a specified format, extracts user keywords based on the user's question, and matches them with the keywords in the message set. Finally, the large model outputs the event process in a specified format and synchronously sends it to the email.
[0097] Example 2:
[0098] 1. Real-time data processing mechanism:
[0099] Design a real-time data processing mechanism to ensure that information can be quickly collected, processed, and reported. This mechanism uses an efficient data stream processing framework to reduce the latency of information processing. At the same time, for a large number of data sources, a message queue system is used to prevent data loss.
[0100] In the real-time data processing mechanism, it includes:
[0101] Use streaming APIs: For data sources with strong real-time characteristics such as social media and news feeds, use the streaming APIs provided by them (such as Twitter Streaming API, RSS PubSubHubbub, etc.) to receive data in real time.
[0102] Introduce a message queue system (such as RabbitMQ, Kafka, etc.) to buffer the real-time data stream. In this way, even if the data processing module is temporarily unable to keep up with the data reception speed, the data will not be lost.
[0103] Use a distributed computing framework (such as Apache Spark Streaming, Apache Flink, etc.) to process the real-time data stream. These frameworks can efficiently process large-scale data streams and provide fault tolerance and scalability.
[0104] 2. Multi-source data aggregation
[0105] 1-1. Data cleaning and preprocessing: Before aggregating the data, it is necessary to clean and preprocess the data. Specifically, ① remove irrelevant information such as HTML tags and advertisements; ② filter prohibited words; ③ unify the case and perform simplified and traditional Chinese conversions to ensure the quality and accuracy of the data.
[0106] 1-2. Data fusion technology: The specific path of the data fusion technology is as follows: For each piece of data crawled by the crawler, use the cosine similarity algorithm to calculate the similarity of these data, combine the data with a similarity exceeding the threshold to form a message set. And sort the events in chronological order according to the data timestamp.
[0107] 2. Automatic processing: Use natural language processing technology to automatically extract key information to improve efficiency.
[0108] 2-1. Text summarization technology
[0109] The present invention will use the Bert model to implement the text summarization technology. Specifically, extract the keywords of each piece of data in the message set, and the three keywords with the most occurrences are the keywords of the event.
[0110] 2-2. Information extraction technology
[0111] Information extraction techniques can be divided into two methods: rule-based and machine learning-based. The latter especially benefits from the development of deep learning, such as the application of CNN and RNN in information extraction tasks. This technology is used to extract message metadata during the data preprocessing process.
[0112] 2-3. Deep Learning
[0113] Through a series of steps, including data preprocessing, applying a deep learning model, generating embedding vectors, aggregating the embeddings of messages, and feature extraction and dimensionality reduction, the present invention successfully converts message data into embedding vectors with deep semantic features.
[0114] In the present invention, first, the input message data needs to be preprocessed to better convert it into embedding vectors. The specific implementation methods are as follows:
[0115] (1) Word Segmentation
[0116] For the input message data, the word segmentation technology is used to split it into words or sub-words. For example, for a sentence: "Deep learning is the future direction", it can be split into: ["Deep", "learning", "is", "future", "the", "direction"] through a word segmenter.
[0117] Method Selection: Common word segmentation tools can be used, such as Jieba Segmentation, HanLP, or a custom word segmentation algorithm.
[0118] (2) Stop Word Removal
[0119] After word segmentation is completed, stop words in the message need to be removed. Stop words refer to those words that do not make a substantial contribution to the semantic analysis of the message, such as "the", "is", "in", etc. This step helps to reduce the interference of meaningless information and improve the efficiency of the model.
[0120] (3) Vocabulary Building
[0121] Based on the preprocessed data, build a vocabulary that contains all possible words. The vocabulary will include all the words or sub-words that appear in the training data and assign a unique index to each word. This vocabulary will be used to convert the words in the input data into a numerical representation that the model can process. Traverse the training data, record the frequency of each word, retain the high-frequency words, and assign a unique integer ID to each word.
[0122] 2. Using a Deep Learning Model
[0123] After the data preprocessing is completed, a deep learning model is then used to learn the semantic features of the message data and generate an embedding vector. The specific implementation is as follows:
[0124] (1) Recurrent Neural Network (RNN)
[0125] In the present invention, a recurrent neural network (such as LSTM or GRU) can be used to learn the sequential information in the message data. Through the RNN, each input word is mapped to a vector, and then the model gradually learns the context information through time steps.
[0126] (2) Transformer Model
[0127] To more precisely model the context information of the message data, a Transformer-based model (such as BERT or GPT) is used. The Transformer model captures the global dependencies between words through the self-attention mechanism, thereby better understanding the context information of the input data.
[0128] 3. Generate Embedding Vector
[0129] After being processed by the deep learning model, each word in the input data will be mapped to an embedding vector. The following are the specific steps for generating the embedding vector:
[0130] (1) Embedding Layer
[0131] In the deep learning model, the Embedding Layer is used to convert the input discrete words (for example, the IDs in the vocabulary) into dense vector representations. Each word will be mapped to a fixed-length vector in the embedding layer to capture its semantic features.
[0132] (2) Global Context Modeling
[0133] Use the RNN, LSTM, GRU or Transformer model to generate context-related embedding vectors for each word in combination with the surrounding context information. For example, when using the BERT model, each input word will generate a word vector containing context information through a multi-layer Transformer structure.
[0134] 4. Aggregate the Embedding of the Message
[0135] In the message data, usually multiple words form a complete message or sentence. To aggregate the embedding vectors of these words into a vector representing the entire message:
[0136] (1) Average Pooling
[0137] Average the embedding vectors of all words to generate a message vector of a fixed length. This method is simple and effective, and can retain the global information of the message to a certain extent.
[0138] (2) Weighted Pooling
[0139] According to the importance of the words in the message, assign different weights to different words and perform weighted averaging according to the weights. This can enhance the attention to important words.
[0140] (3) Using RNN / LSTM / Transformer
[0141] Models such as RNN, LSTM, or Transformer can be directly used to process the entire message. The model will process each word step by step and generate an embedding vector representing the entire message in the last step.
[0142] 5. Feature Extraction and Dimensionality Reduction
[0143] For some application scenarios, the dimension of the embedding vector may be high, or further compression is required for storage and calculation. To reduce the computational cost or improve the efficiency of the model, feature extraction and dimensionality reduction techniques can be used.
[0144] (1) Principal Component Analysis (PCA)
[0145] Use PCA to reduce the dimension of the embedding vector to reduce the data dimension while retaining important semantic information as much as possible.
[0146] (2) t-SNE Dimensionality Reduction
[0147] t-SNE (t-Distributed Stochastic Neighbor Embedding) is a non-linear dimensionality reduction technique suitable for the visualization of high-dimensional data. For more complex applications, t-SNE can be used to reduce the dimension of the embedding vector and perform analysis.
[0148] 3. Real-time Update: Respond quickly to new information to ensure the timeliness of the notification.
[0149] 4. Continuous Optimization: Continuously optimize the notification model based on user feedback.
[0150] 2. Automated Notification Generation Engine:
[0151] After obtaining a message chain from different data sources that describes the same event and is sorted chronologically, since it is stored in a streaming database, we can obtain data in real time through the stream processing API.
[0152] Use large model application development platforms such as Dify and FastGPT to create an automated notification generation engine. Specifically, build a large model application process: the data source is the message collection in the API. The message collection has its own keywords, and each message in the collection contains information such as the occurrence time, information source, and information weight; set relevant built-in prompts to let the large model transform the message information into a specified output format, such as sorting by time sequence, outputting time first and then the message, and bolding messages with high information weight, etc.; when the user inputs an instruction, the large model recognizes the user's keywords, matches them with the collection keywords, hits the collection, and outputs the event message text in the previously set format. Finally, call the large model interface and connect it to platforms such as email, so that the text generated by the large model can be directly notified in the form of an email, etc.
[0153] 3. Security and privacy protection measures:
[0154] (1) During the data cleaning process, desensitize sensitive data, such as removing personal information, etc. (for example, removing the social platform spokesperson id from the information obtained from the social platform).
[0155] (2) Implement real-time data encryption policies during data transmission and storage to ensure data security.
[0156] (3) Use open-source and locally deployable large model application development platforms to avoid internal data leakage.
[0157] 4. User feedback mechanism
[0158] After notifying the user of the event message, the user can provide feedback through email via the notification path. The feedback email will be recorded and submitted to the large model of the application that generates the report. This model will record, merge the user feedback, and provide a solution, and will automatically generate a daily feedback report for engineers every day for the engineers to optimize the system.
[0159] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.
Claims
1. A method for information notification based on an aggregation model, characterized in that: S1: Data collection and preprocessing: Based on the real-time data processing mechanism, use web crawlers or API interfaces to crawl multi-source data from public data sources in real time and clean the multi-source data; S2: Feature extraction: Use deep learning to convert the text content of multi-source data into embedding vectors, use NLP technology to extract and store the metadata of each data item, that is, text features, into a metadata list, and calculate the event occurrence time based on the text content and add it to the metadata list; S3: Multi-source data fusion: construct a multi-source data input set, create an automated notification generation engine based on the big model, input the multi-source data input set into the automated notification generation engine, and output an event notification text in a set format; The multi-source data input set includes: a message set constructed based on the text features obtained in step S2, and reconstructing the generated message set based on the message set construction set keywords, message chains and message weights, and using the reconstructed message set as the multi-source data input set.
2. A method for information notification based on an aggregation model according to claim 1, characterized in that: The public data sources mentioned in S1 include: social media, news sources, and official announcements.
3. A method for information notification based on an aggregation model according to claim 1, characterized in that: Step S1 captures data in real time and designs a real-time data processing mechanism to reduce information delay and data loss; The real-time data processing mechanism integrates streaming API, message queue system and distributed computing framework; For data sources with strong real-time characteristics, such as social media and news sources, use the streaming API to receive data in real time; The message queue system buffers the received real-time data stream; The distributed computing framework processes real-time data streams.
4. A method for information notification based on an aggregation model according to claim 1, characterized in that: The method for capturing and cleaning data in S1 is: S11: Capture data and label it: the data includes: timestamp, source identifier, content text; S12: Data cleaning: remove irrelevant information, filter banned words, unify uppercase and lowercase letters, and convert simplified and traditional Chinese fonts; the irrelevant tags include: HTML tags and advertisements.
5. A method for information notification based on an aggregation model according to claim 1, characterized in that: The metadata header includes: source, author, keywords, release time, number of likes and number of reposts.
6. A method for information notification based on an aggregation model according to claim 1, characterized in that: The method for calculating the event occurrence time and adding it to the metadata list is: add "occurrence time" in the metadata header, use NLP technology to extract the occurrence time mentioned in the data, or calculate the event occurrence time based on the release time and time-related information. If there is no description of the occurrence time, use the release time instead of the event occurrence time.
7. A method for information notification based on an aggregation model according to claim 1, characterized in that: The method for constructing a multi-source data input set in step S3 is: S31: Constructing a message set: Using the text cosine similarity algorithm and text features to calculate the similarity of text information in different data sources, combining data from different data sources but with high content similarity to form a message set; S32: constructing a set of keywords: extracting keywords based on the Bert model, and combining the three keywords that appear most frequently in the message set as the set of keywords; S33: Get message chain: sort messages according to the time when the event occurred, and get the message chain sorted according to the timeline; S34: Construct message weights: Assign message weights to each data item based on the credibility of the data source, user interaction, and time freshness, and use a weighted sum algorithm to determine the priority of the information, which will be reported in the subsequent focus; S35: Reconstruct the message set: store the set keywords, message chains and message weights into the message set to form a multi-source data input set.
8. A method for information notification based on an aggregation model according to claim 1, characterized in that: In step S3, the method for creating an automated notification generation engine based on a large model and performing information notification is as follows: A31: Build a large model input: The large model input is a multi-source data set, that is, a reconstructed message set in the streaming API, the reconstructed message set includes a set keyword, and each message in the message set includes an occurrence time, an information source, and an information weight; A32: Set the output format of the large model: Set the built-in prompt to let the large model convert the message information into the specified output format. The output format includes: sorting by time, outputting the time first and then the message, and bolding the information with high weight; A33: Set the big model working mode: When the user inputs a command, the big model identifies the user's keywords, matches them with the set keywords, hits the set, and outputs the event message text according to the output format set in A32; A34: Call the large model interface to access the output platform to notify the user of the event message text.
9. A method for information notification based on an aggregation model according to claim 8, characterized in that: After the user receives the event message text, the user chooses to provide feedback via email through the notification path. The feedback email will be recorded and submitted to the application big model that generates the report; the big model will merge the user feedback records, optimize the problems in the feedback email, and then output it to the user. The system will automatically generate a daily feedback report and provide it to engineers for them to optimize the system.
10. A method for information notification based on an aggregation model according to claim 1, characterized in that: The method comprises step S4, setting security and privacy protection measures during the information notification process; A41: During the data cleaning process, sensitive data is desensitized; A42: Implement real-time data encryption strategies during data transmission and storage; A43: Use an open source, locally deployable large-model application development platform.
11. A system for information notification based on an aggregation model based on the method of claim 1, characterized in that: include: Data collection and preprocessing module, feature extraction and representation module, multi-source data fusion module and automatic notification engine module; The data collection and preprocessing module includes: a multi-source data acquisition submodule for obtaining multi-source data through a web crawler, and a data preprocessing submodule for removing irrelevant information, filtering banned words, and processing text formats; The feature extraction and representation module includes: a text conversion submodule for converting the text of the multi-source data into a vector representation, a metadata acquisition submodule for extracting metadata and deleting privacy information, and an occurrence time acquisition submodule for extracting the occurrence time of an event using NLP technology; The multi-source data fusion module includes: a message set construction submodule that calculates the similarity of text information in different data sources based on cosine similarity and constructs a message set, a set keyword acquisition submodule that selects the three most numerous keywords in the set as set keywords, and a message weight acquisition submodule that calculates message weights based on data source, forwarding, commenting, and likes, and time; The automatic notification engine module includes: a large model input submodule that inputs user questions into a large model development platform, inputs data from a multi-source data fusion module into the large model development platform and sets a message set format, a message set matching submodule that matches keywords in user questions with keywords in a message set to obtain matching results, and a large model output submodule that outputs a message set in a specified format based on the matching results as an event and sends it to an email synchronously.
Citation Information
Patent Citations
Personalized news clue pushing method and system
CN106484733A
Inferring temporal relationships for cybersecurity events
CN113647078A
Emergency time sequence automatic construction method based on abstract generation algorithm
CN114722194A
Case information-based key text element extraction method
CN118607520A
System and method for incident validation and ranking using human and non-human data sources
US20180253813A1