Scientific and technological fast message sensing system based on large language model
By designing a technology-based information awareness system based on large language models, the problem of lack of integrity and systematic coordination in the existing technology is solved, and the rapid and accurate identification and analysis of massive information is achieved, and the efficiency and accuracy of intelligence perception are improved.
Patent Information
- Application Number
- CN202510163963.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-13
AI Technical Summary
When using large language models for scientific and technological information perception, the existing technology lacks overall and systematic coordination, which limits the comprehensive effectiveness of the intelligence perception system.
A technology news intelligence perception system based on large language models was designed, including a technology news monitoring subsystem, identification subsystem and visualization subsystem. Through intelligent means, scientific news is monitored, analyzed and visualized to improve the efficiency and accuracy of intelligence recognition and perception.
It realizes rapid and accurate identification and analysis of massive information, improves the efficiency and accuracy of intelligence perception, enhances users' perception dimensions and intuitiveness of technological news, and supports decision makers to quickly respond to technological changes and market changes.
Smart Images

Figure CN120144845A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligence perception, and more specifically, to a scientific and technological news intelligence perception system based on a large language model. Background Art
[0002] With the accelerating advancement of digitalization and globalization, the volume of various types of data has shown an explosive growth, bringing rich scientific and technological intelligence resources, but also increasing the difficulty of information screening and value identification. Intelligence has become a key resource for researchers, decision-makers, and enterprises to cope with complex challenges and gain a competitive advantage, and the demand for customized solutions from these groups is also increasing continuously. Intelligence perception has always been an important topic in intelligence research, and intelligence perception is a necessary condition for policy-making. In the current big data era, finding valuable scientific and technological information quickly and accurately from massive data for different intelligence requesters is the key issue to meet the needs of scientific and technological intelligence.
[0003] The acquisition and analysis of scientific and technological intelligence also face many challenges. First is the problem of information overload. The noise and redundant information in massive data make intelligence identification more difficult. Second is the issue of information timeliness and accuracy. How to quickly find valuable information from the data has become the core of scientific and technological intelligence work. Zhang Xiaolin et al. emphasized that intelligence work departments should control and reduce the cost of intelligence retrieval as much as possible to fully realize the use value of intelligence. Coincidentally, the rise of large language models represented by ChatGPT has brought new opportunities to intelligence research. Large language models have shown unique advantages in intelligence research. The partial explosion of artificial intelligence technology has allowed intelligence workers to see the open-source intelligence field as an excellent opportunity to play the role of AI on the main battlefield.
[0004] Scientific and technological news intelligence, with its characteristics of fast, concise, and large amounts of information transmission, demonstrates extremely high research value. Its unique timeliness and wide coverage have made it occupy an important position in modern scientific and technological intelligence research. In recent years, scholars have conducted many explorations in scientific and technological news intelligence work and intelligence system construction. However, intelligence perception is a complex process involving multiple sub-links and steps. In existing solutions that apply large language models to the field of intelligence perception, the main mode is human-machine symbiosis, that is, letting machines process problems according to a predetermined program, allowing humans to focus more on complex explorations. These technologies mainly focus on applying the text classification or text generation capabilities of large language models to a specific link of the intelligence situation, and there are very few relevant researches and technical solutions for scientific and technological news intelligence perception truly based on large language models. Current research mostly focuses on the independent exploration of these links, lacking overall and systematic coordination, which limits the comprehensive effectiveness of the intelligence perception system. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a scientific and technological news intelligence perception system based on large language models. By combining large language models with intelligence perception and monitoring and analyzing scientific and technological news through intelligent means, the efficiency and accuracy of intelligence recognition and perception are improved, and significant advantages are achieved in processing massive information and enhancing intelligence perception capabilities.
[0006] The purpose of the present invention is achieved through the following solutions:
[0007] A scientific and technological news intelligence perception system based on large language models, comprising: a scientific and technological news monitoring subsystem, a scientific and technological news recognition subsystem, and a scientific and technological news visualization subsystem;
[0008] The scientific and technological news monitoring subsystem includes an information collection module, a data standardization module, and a database. In the information collection module, it operates according to a predetermined process to obtain multi-modal scientific and technological news data; in the data standardization module, it receives the multi-modal data output by the information collection module and is used to perform data cleaning, structured processing, and semantic standardization operations; after the standardized processing, the data is stored in the database;
[0009] The scientific and technological news recognition subsystem includes an intelligence value screening module and an evaluation system module; the intelligence value screening module combines the evaluation system to identify the intelligence value;
[0010] The scientific and technological news visualization subsystem includes a visualization module; the scientific and technological news visualization subsystem receives the screened information from the database and realizes multi-dimensional visualization presentation through the visualization module.
[0011] Furthermore, the obtaining of multi-modal scientific and technological news data specifically includes: directly sending a request to the target scientific and technological news source through the RSS subscription interface and receiving a structured XML data stream; then parsing the received data, matching predefined topic tags and keywords, extracting content related to the scientific and technological field and storing it in a temporary area; at the same time, the web crawling function is started, and based on the configured target URL list, web page HTML data is obtained through the crawler framework; using DOM tree parsing, the body area is extracted through XPath or CSS selectors, and the text is segmented in combination with the NLP model, and the content that meets the requirements is screened and stored in the temporary area.
[0012] Furthermore, the acquisition of multimodal technology news data specifically includes: setting up a timed task scheduling module to automatically trigger a crawler instance at a set time interval, updating the crawling strategy in real time, dynamically adjusting the weights of target keywords, and preferentially processing websites with high-frequency updates; the processing of audio data uses a speech recognition model to accurately transcribe the audio content into text; the processing of video data obtains key frames through frame extraction technology, generates text information in combination with an image object detection algorithm and a description generation model, and integrates it with the original text and metadata through semantic fusion technology and stores them uniformly in a multimodal database.
[0013] Furthermore, in the data standardization module, the operations of performing data cleaning, structured processing, and semantic standardization specifically include:
[0014] First, run a noise filtering algorithm to analyze the text, eliminate redundant information by calculating term weights, and at the same time call the locality-sensitive hashing algorithm to detect and delete duplicate content; in the formatting and normalization section, use regular expressions to extract key fields, and with the help of a large language model and prompt engineering, convert unstructured text into a standardized JSON data format;
[0015] Subsequently, perform SAO triple extraction operations. Construct a text dependency tree through dependency syntactic analysis to identify the subject, predicate, and object in the sentence; the multi-level semantic relationships in compound sentences are disassembled and triple extractions are performed one by one through a generative language model to ensure content integrity and semantic accuracy; in the fusion process of multimodal data, load a multimodal deep learning model, align the audio-to-text, image description text with the original text through feature vector matching technology, and use cross-attention mechanism to generate a unified semantic representation.
[0016] Furthermore, the database specifically adopts a distributed database architecture to ensure horizontal scalability and disaster tolerance capabilities of the data through sharding and replication mechanisms, and at the same time optimizes the execution efficiency of complex queries in a columnar storage manner.
[0017] Furthermore, the running logic of the database specifically includes data warehousing, index construction, and efficient retrieval;
[0018] In the data warehousing stage, the news data output by the data standardization processing module is first partitioned through the preprocessing interface of the database; based on predefined classification rules, the data is assigned to selected logical partitions, and the data is split into field storage in a columnar storage manner to reduce the I / O burden; each record is assigned a unique identifier, and the operation details are recorded through a transaction log to ensure the atomicity and consistency of the write; for multimodal data, a nested JSON structure is used to uniformly store content such as text, audio-to-text, and image descriptions, and a modal association table is introduced to establish a cross-modal logical mapping relationship;
[0019] During the index construction phase, the database module covers keyword query and semantic retrieval scenarios through a dual-index mechanism of inverted index and semantic index; during the construction of the inverted index, the text content is disassembled into word or phrase units using the tokenization algorithm BERT tokenizer, and its positions in the data set are recorded; the weight of each keyword is calculated by the BM25 algorithm and stored in the inverted list to support fast relevance queries; at the same time, the semantic index module uses the Transformer models BERT and RoBERTa to generate high-dimensional semantic vectors for the text content, and adopts PCA dimensionality reduction processing to reduce the storage overhead; the dimensionality-reduced vectors are stored in the vector database FAISS, and combined with an efficient ANN retrieval algorithm to achieve fast matching of high-dimensional vectors.
[0020] In the data retrieval phase, the user's query request is first analyzed by the query parsing module to understand its intent and parsed into a structured query request; for keyword queries, the inverted index is directly called, and the results are filtered and sorted through boolean logic and relevance scoring; for semantic retrieval, the query text input by the user is used to generate a vector representation by the semantic embedding model, and the similarity is calculated with the vectors already stored in the vector database, and the HNSW graph retrieval algorithm is used to quickly find the most relevant records; the retrieval results are re-ranked by the ranking module, and the ranking criteria are dynamically adjusted according to time priority, relevance priority, or user-defined weight strategies.
[0021] Furthermore, in the cache mechanism design of the database, the results of high-frequency queries are cached in the in-memory database Redis and directly returned in similar requests; at the same time, the database module regularly performs index reconstruction and storage partition merging to ensure the efficiency and stability of the system's long-term operation by optimizing the disk layout and data structure.
[0022] Furthermore, the intelligence value screening module combines an evaluation system to identify the intelligence value, specifically including: the intelligence value screening mechanism achieves precise identification through three-layer filtering. First, the first layer sets an intelligence value threshold filtering module to set a benchmark threshold for preliminary screening of the input information; a feature extraction method based on term frequency-inverse document frequency is used to calculate the weight scores of the text keywords; when the comprehensive score of the information is lower than the set threshold, it is marked as low-value information and filtered.
[0023] The second layer uses a regular matching pattern for filtering, maintaining a regular expression rule library that contains common expression patterns and key phrases in technology bulletins; key information in the text is identified through pattern matching; the higher the matching degree, the higher the professionalism and technical content of the information; and a dynamic rule update mechanism is used to continuously optimize the rule library according to new samples.
[0024] The information value ranking and screening serves as the last layer of filtering, comprehensively evaluating by combining multiple dimensions of the evaluation system module. Among them, the importance dimension is calculated through keyword weights and citation frequencies; the relevance measures the matching degree with the user's field based on the cosine similarity algorithm; the novelty is evaluated through the time decay function and innovation point extraction; the reliability is judged based on the source credibility and content consistency; the trend is identified through time series analysis and hot spot clustering to determine the development direction of technology.
[0025] Secondly, an information recognition weight allocation module is set up. Specifically, the analytic hierarchy process is used to determine the weights of each dimension. First, a judgment matrix is established, and the relative importance between dimensions is determined through expert scoring. After passing the consistency test, the eigenvector is calculated to obtain the weight coefficients of each dimension, and the weight configuration is dynamically adjusted according to different application scenarios.
[0026] Furthermore, the multi-dimensional visualization presentation is realized through the visualization module, which specifically includes:
[0027] Set up a large model information overview module to call the large language model API for semantic understanding and reorganization of high-value bulletins. Construct a standardized prompt template to guide the model to generate structured overview content. Extract key information from the text through the attention mechanism, and conduct correlation analysis on the information combined with knowledge graph technology. Finally, generate a review report containing core viewpoints, technical features, and development trends.
[0028] Set up a heat map module. Based on the geographic information processing engine, map the bulletin data to the world map according to geographical coordinates. Use the kernel density estimation algorithm to calculate the information density, adopt an adaptive color mapping scheme, and present different colors in hot spot areas. Dynamically display the evolution process of the scientific and technological activity levels in each region through the sliding of the time window. At the same time, multi-level drilling is supported, and users can gradually drill down from the global view to specific countries and regions.
[0029] Set up a ranking module to achieve dynamic sorting based on the time dimension. Specifically, through the sliding time window method, aggregate and statistically analyze data at different time granularities. Combine the information value score and the evaluation results of the large model, and use a weighted sorting algorithm to generate a list. The interface adopts a responsive design, supports the switching between list view and card view, and provides trend charts to display historical changes.
[0030] Set up a word cloud module. Use a word segmentation engine to process the text, and calculate the word weights through the TF-IDF algorithm. Adopt a layout algorithm to optimize the placement position of words while ensuring visual beauty. At the same time, multiple shape templates are supported, and the color scheme is matched according to word categories. Optimize the word spacing through the force-directed algorithm to ensure a compact and legible layout.
[0031] Set up a scientific and technological intelligence graph module, and use a deep learning model to achieve SAO triple extraction; use a graph database to store entity relationships, and calculate the node positions through a graph layout algorithm; the visualization layer uses WebGL to achieve smooth rendering of large-scale graphs, realize real-time incremental updates, and new news data can be dynamically incorporated into the existing graph structure.
[0032] The beneficial effects of the present invention include:
[0033] The present invention constructs an intelligence perception system, which improves the recognition and analysis capabilities of scientific and technological news intelligence. The large language model technology is adopted, enabling the system to quickly extract and analyze key information in a vast amount of information. While improving the accuracy and efficiency of intelligence recognition, the visualization method is used to increase the perception dimension of user intelligence and enhance the perception intuitiveness, which helps decision-makers quickly respond to technological changes and market changes, thereby maintaining a competitive advantage and providing solid technical support for scientific and technological intelligence research.
[0034] The system of the present invention can efficiently monitor relevant scientific and technological news according to user needs, accurately identify high-value intelligence, and significantly improve the effect of intelligence perception through visualization. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 It is the structural block diagram of the scientific and technological news intelligence perception model based on the large language model in the embodiment of the present invention;
[0037] Figure 2 It is the flow chart of intelligence value screening in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] All the features disclosed in all the embodiments in this specification, or all the steps in the implicit disclosed methods or processes, except for mutually exclusive features and / or steps, can be combined and / or extended, replaced in any way.
[0039] The specific implementation process of the present invention is as follows:
[0040] In a preferred embodiment, a scientific and technological news intelligence perception system based on the large language model is provided, which specifically includes a scientific and technological news monitoring subsystem, a scientific and technological news recognition subsystem, and a scientific and technological news visualization subsystem.
[0041] Among them, the main objective of the Technology News Monitoring Subsystem is to obtain and perceive the latest technology intelligence information in an efficient manner. This subsystem relies on various technology news bulletins as the main information sources, leveraging their timeliness and objectivity to ensure that the obtained technology intelligence can reflect the latest development trends and objective facts.
[0042] In information collection, the Technology News Monitoring Subsystem collects news bulletins by integrating various data forms such as text, audio, and video, combined with the method of combining Really Simple Syndication (RSS) subscriptions and web crawling with a scheduled task module. The system uses RSS subscriptions to real-time subscribe to various technology news sources to ensure the immediate acquisition and automated update of news bulletin data. At the same time, the scheduled task module and the web crawling program regularly scan multiple technology news websites, and through keyword filtering and topic classification, ensure that the collected information covers diverse data sources such as technology news, industry trends, technical magazines, social media updates, company announcements, and academic research. This multi-level and multi-channel information collection method ensures the comprehensiveness and relevance of the data.
[0043] In information update, due to the large variety and sources of monitored data sources, it often leads to the failure of subscription data and prefabricated crawling mechanisms, requiring manual adjustment. To ensure the timeliness and accuracy of information, the Technology News Monitoring Subsystem of the present invention adopts an intelligent information dynamic management mechanism. The system uses real-time monitoring and intelligent push technologies to ensure that the data monitoring sources always remain up-to-date. The real-time monitoring module continuously monitors the subscribed RSS sources and key technology news websites, checking their connectivity and changes in data streams. When it detects the failure of RSS subscriptions or the loss of monitored data, it automatically alerts system personnel for troubleshooting to ensure the availability of the system.
[0044] In the specific design concept, as Figure 1 shown, the Technology News Monitoring Subsystem includes an information collection module and a data standardization module.
[0045] In the information collection module, the information collection module operates according to a predefined process to ensure the efficient acquisition of multi-modal technology news bulletin data. First, the system directly sends requests to the target technology news sources through the RSS subscription interface and receives structured XML data streams. The system parses the received data, matches predefined topic tags and keywords, extracts content related to the technology field, and stores it in a staging area. At the same time, the web crawling function is started, and based on the configured list of target URLs, the crawler framework is used to obtain the web page HTML data. The system uses DOM tree parsing technology, extracts the body area through XPath or CSS selectors, combines with the NLP model to segment the text, and filters the content that meets the requirements and stores it in the staging area.
[0046] Set up a timed task scheduling module to automatically trigger crawler instances at set time intervals, update the scraping strategy in real time, dynamically adjust the weights of target keywords, and prioritize the processing of websites with high-frequency updates. The processing of audio data uses a speech recognition model to accurately transcribe the audio content into text. The processing of video data obtains key frames through frame extraction technology, combines image object detection algorithms and description generation models to generate text information, and integrates it with the original text and metadata through semantic fusion technology, and stores it uniformly in a multimodal database.
[0047] In the data normalization module, the data normalization module directly receives the multimodal data output by the information collection module and performs high-precision data cleaning, structuring, and semantic normalization operations. First, the system runs a noise filtering algorithm to analyze the text, eliminates redundant information by calculating term weights, and at the same time calls the locality-sensitive hashing algorithm to detect and delete duplicate content. In the formatting and normalization step, regular expressions are used to extract key fields (such as titles, times, sources), and with the help of large language models and prompt engineering, unstructured text is converted into a standardized JSON data format. Subsequently, the system performs SAO triple extraction operations, constructs a text dependency tree through dependency syntactic analysis, and accurately identifies the subject, predicate, and object in the sentence. The multi-level semantic relationships in complex sentences are disassembled and triple extractions are performed one by one through a generative language model to ensure content integrity and semantic accuracy. In the fusion process of multimodal data, the system loads a multimodal deep learning model, aligns the audio-to-text and image description text with the original text through feature vector matching technology, and uses cross-attention mechanisms to generate a unified semantic representation.
[0048] After being standardized, the data is stored in an efficient standardized database, which has a high degree of standardization and semantic consistency, providing a reliable data foundation for subsequent analysis and value assessment. In the technology news monitoring subsystem of the present invention, the entire process realizes the full-automatic cleaning and standardized conversion of multimodal data with clear logic and strict algorithms.
[0049] In the database module, the database module of the present invention is designed according to the requirements of high performance and high reliability, and realizes the storage, management, and efficient retrieval functions of multimodal technology news data. This module adopts a distributed database architecture, ensures the horizontal scalability and disaster tolerance of data through sharding and replication mechanisms, and optimizes the execution efficiency of complex queries in a columnar storage manner. The operation logic of the module includes three major links: data warehousing, index construction, and efficient retrieval, and each link is closely connected through specific algorithms and processing processes.
[0050] In the data storage phase, the flash news data output by the data standardization processing module is first partitioned through the preprocessing interface of the database. Based on predefined classification rules (including topic categories, timestamps, etc.), the data is assigned to specific logical partitions, and the data is split into fields for storage using a columnar storage method, thereby reducing the I / O burden. Each record is assigned a unique identifier (UUID), and the operation details are recorded through a transaction log to ensure the atomicity and consistency of the writes. For multi-modal data, the system uses a nested JSON structure to uniformly store content such as text, audio-to-text, and image descriptions, and introduces a modality association table to establish a cross-modal logical mapping relationship.
[0051] In the index construction phase, the database module covers keyword query and semantic retrieval scenarios through a dual-index mechanism of inverted index and semantic index. During the construction of the inverted index, the text content is disassembled into word or phrase units using the tokenization algorithm BERT tokenizer, and its positions in the data set are recorded. The weight of each keyword is calculated using the BM25 algorithm and stored in the inverted list to support fast relevance queries. At the same time, the semantic index module uses the Transformer models BERT and RoBERTa to generate high-dimensional semantic vectors for the text content, and uses PCA dimensionality reduction processing to reduce the storage overhead. The dimensionality-reduced vectors are stored in the vector database FAISS, and combined with an efficient ANN (Approximate Nearest Neighbor) retrieval algorithm to achieve fast matching of high-dimensional vectors.
[0052] In the data retrieval phase, the user's query request is first analyzed by the query parsing module to parse its intent into a structured query request. For keyword queries, the system directly calls the inverted index and filters and sorts the results through boolean logic and relevance scoring. For semantic retrieval, the query text input by the user is generated into a vector representation through a semantic embedding model, and the similarity is calculated with the vectors stored in the vector database. The HNSW (Hierarchical Navigable SmallWorld) graph retrieval algorithm is used to quickly find the most relevant records. The retrieval results are re-ranked by the ranking module, and the ranking criteria can be dynamically adjusted according to time priority, relevance priority, or user-defined weight strategies.
[0053] The present invention further improves the caching mechanism. The caching mechanism of the database module further enhances the retrieval performance. In the specific concept, the results of high-frequency queries are cached in the in-memory database Redis and directly returned in similar requests, significantly reducing disk I / O and computing resource consumption. At the same time, the database module periodically performs index reconstruction and storage partition merging to ensure the efficiency and stability of the system's long-term operation by optimizing the disk layout and data structure.
[0054] Through the seamless connection of data storage, index construction, and efficient retrieval, the database module realizes the precise management and efficient invocation of multi-modal scientific and technological news data, providing solid basic support for subsequent intelligent analysis and value mining.
[0055] Among them, the core of the scientific and technological news recognition subsystem lies in constructing a multi-dimensional evaluation system, covering importance, relevance, novelty, reliability, and trendiness. These indicators jointly constitute a comprehensive evaluation framework, providing a basis for systematic scientific and technological news evaluation.
[0056] Table 1
[0057]
[0058] In the process of intelligence value evaluation, the present invention adopts a method of combining a large language model with fixed prompts for automated scoring. For each evaluation indicator, a series of precise prompts can be designed to ensure that the model can effectively capture and analyze key information. During the scoring process, the prompts are adjusted and optimized through multiple rounds to ensure that they accurately reflect the evaluation criteria. To improve the applicability and accuracy of the prompts, the prompt system can be systematized into multiple parts such as context, requirements, and demands, forming a standardized paradigm that can adapt to different user needs and tasks. The score of each indicator output by the model is recorded in the intelligence evaluation form according to the established standards.
[0059] Furthermore, based on the evaluation system and the large model scoring framework, the present invention develops an intelligence value screening mechanism. This mechanism adopts a hierarchical screening strategy to ensure that the selected scientific and technological news has the highest intelligence value. In the preliminary screening stage, by setting a basic threshold, the news below the lowest standard is excluded. Then the system excludes the news without intelligence value through the regular matching function, including but not limited to the fields not concerned, directions, specific low-value intelligence features, and other situations that may lead to intelligence irrelevance or invalidity. Finally, the system sorts the remaining news according to the comprehensive score and selects them from high to low.
[0060] In the specific design concept, as Figure 2 shown, the scientific and technological news recognition subsystem includes an intelligence value screening module and an evaluation system module.
[0061] In the intelligence value screening module of the present invention, the intelligence value screening mechanism realizes precise identification through three-layer filtering. First, the intelligence value threshold filtering module sets a benchmark threshold to conduct a preliminary screening of the input information. The system adopts a feature extraction method based on term frequency-inverse document frequency (TF-IDF) to calculate the weight score of text keywords. When the comprehensive score of the information is lower than the set threshold, the system marks it as low-value information and filters it.
[0062] The second layer uses regular matching mode for filtering. The system maintains a regular expression rule library, which contains common expression patterns and key phrases in science and technology news. Key information such as technical indicators, research progress, and application scenarios in the text is identified through pattern matching. The higher the matching degree, the higher the professionalism and technical content of the information. The system uses a dynamic rule update mechanism to continuously optimize the rule library according to new samples.
[0063] As the last layer of filtering, the intelligence value ranking and screening combines the five dimensions of the evaluation system module for comprehensive evaluation. The importance dimension is calculated through keyword weights and citation frequencies; the relevance is measured based on the cosine similarity algorithm to measure the matching degree with the user's field; the novelty is evaluated through the time decay function and innovation point extraction; the reliability is judged based on the source credibility and content consistency; and the trend is identified through time series analysis and hot spot clustering to identify the development direction of technology.
[0064] The intelligence recognition weight distribution module uses the Analytic Hierarchy Process (AHP) to determine the weights of each dimension. First, a judgment matrix is established, and the relative importance between dimensions is determined through expert scoring. After passing the consistency test, the eigenvector is calculated to obtain the weight coefficients of each dimension. The system supports dynamically adjusting the weight configuration according to different application scenarios.
[0065] Among them, the science and technology news visualization subsystem aims to present the information of science and technology news more intuitively and understandably through various visualization technologies. The system adopts five visualization methods, namely the large model intelligence overview, science and technology intelligence heat map, intelligence value ranking list, science and technology hot word cloud map, and science and technology intelligence atlas. Combining with the large language model, it improves the presentation quality and analysis efficiency of science and technology news information, enabling users to intuitively analyze and perceive science and technology news.
[0066] The large model intelligence overview module uses the large language model to generate interpretations of science and technology news. By inputting the news with higher intelligence value within a given time interval into the large language model and analyzing it in combination with fixed prompt words, the system automatically summarizes and refines the science and technology news, identifies the commonalities in the science and technology news, and conducts appropriate interpretations and inferences. The large language model technology can simplify and clarify complex information, helping users quickly grasp the key information.
[0067] The science and technology intelligence heat map module generates a visual representation of the information density by analyzing the number of news in each region and the news with higher intelligence value. It is generated relying on the map resource package and combining the country information of relevant important news. This method can intuitively display the distribution of science and technology news globally or in specific regions, assisting users in quickly identifying information hot spots and trends.
[0068] The Intelligence Value Ranking List is sorted based on the scores of the evaluation system and combines the comprehensive evaluation of the intelligence value by the large language model. It monitors all-weather and screens out the most valuable technology news for display. According to user needs, the ranking list provides summaries of news in different time intervals such as daily, weekly, monthly, quarterly, and annually, helping users quickly obtain the most important technology information in each time period, so as to better grasp the dynamics and trends.
[0069] The Technology Hot Word Cloud Map module generates a word cloud map through the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm and word frequency analysis to display the high-frequency words and key terms in the technology news. The word cloud map visually presents hot topics and core themes through the size and color of the words. Among them, the TF-IDF algorithm is used to highlight important keywords, while word frequency analysis shows the occurrence frequency of the words. This visualization method effectively displays the main trends and key contents in the technology news.
[0070] The Technology Intelligence Map module extracts the SAO (Subject-Action-Object) triples from the news and uses complex networks for visual presentation. This method not only shows the relationships and associated information between each news, but also helps users understand the overall structure and logic of the information. By constructing and analyzing the technology intelligence map, users can intuitively see the connections and interactions between different events and topics, revealing the knowledge network and potential patterns hidden behind the data. This comprehensive perspective helps users understand the technology news at a deeper level and identify key influencing factors and trends.
[0071] In the specific design concept, the technology news visualization subsystem includes a visualization module and a feature data module.
[0072] The technology news visualization subsystem receives the filtered information from the database and realizes multi-dimensional visualization through five core modules in the visualization module. First, the large model intelligence overview module calls the large language model API to perform semantic understanding and reorganization on high-value news. The system uses prompt engineering technology to construct a standardized prompt template to guide the model to generate structured overview content. Key information of the text is extracted through the attention mechanism, and the information is analyzed associatively by combining knowledge graph technology. Finally, a review report containing core viewpoints, technical features, and development trends is generated.
[0073] The heat map module is based on a geographic information processing engine, maps the news data to the world map according to geographical coordinates. The system uses the kernel density estimation algorithm to calculate the information density and adopts an adaptive color mapping scheme, with hot regions showing deeper tones. By sliding the time window, the evolution process of the technological activity in each region can be dynamically displayed. The system supports multi-level drilling, and users can gradually drill down from the global view to specific countries and regions.
[0074] The ranking module implements dynamic sorting based on the time dimension. The system aggregates and statistics data at different time granularities through the sliding time window method. Combining the intelligence value score and the evaluation results of the large model, a weighted sorting algorithm is used to generate the list. The interface adopts a responsive design, supports the switching between list view and card view, and provides a trend chart to display historical changes.
[0075] The word cloud module first processes the text using a word segmentation engine and calculates the word weights through the TF-IDF algorithm. The system adopts a layout algorithm to optimize the placement of words while ensuring visual beauty. It supports multiple shape templates, and the color scheme is automatically matched according to the word category. The word spacing is optimized through the force-directed algorithm to ensure a compact and legible layout.
[0076] The scientific and technological intelligence graph module uses a deep learning model to implement SAO triple extraction. The system uses a graph database to store entity relationships and calculates the node positions through a graph layout algorithm. The visualization layer uses WebGL technology to achieve smooth rendering of large-scale graphs, supporting interactive operations such as zooming, panning, and querying. The system realizes real-time incremental updates, and new news data can be dynamically incorporated into the existing graph structure.
[0077] All visualization modules adopt a component-based design, supporting flexible combination and layout adjustment. The system realizes real-time data push through WebSocket to ensure the timely update of visualization results. It adopts a responsive design to adapt to different terminal devices and realizes a data synchronization mechanism for multi-person collaborative analysis.
[0078] The solution of the present invention has the following technical advantages:
[0079] (1) Cognitive objective consistency: Traditional systems rely on the cognition and perception of intelligence personnel to determine the value of intelligence, and the standards often vary from person to person, prone to subjective biases. The system based on the large language model provided by the present invention can analyze and identify according to clear standards, avoiding the cognitive differences between different intelligence personnel, thus improving the objectivity and consistency of intelligence value recognition.
[0080] (2) Real-time monitoring ability: Traditional systems are difficult to perceive and identify scientific and technological news in real time outside of manual monitoring time, while the large language model system can run continuously all day long for intelligence monitoring and processing. This continuous monitoring ability significantly enhances the timeliness of intelligence recognition, can capture important scientific and technological news more quickly, and ensures that decision-makers can obtain the latest intelligence in a timely manner.
[0081] (3) Information transfer efficiency: In the traditional path, intelligence is usually passed through intelligence workers and then to intelligence requesters, with a rather long process and prone to information distortion. The system based on large language models of the present invention can directly transfer intelligence to intelligence requesters, simplify the path, reduce intermediate links, and improve transfer efficiency and accuracy. This not only shortens the decision-making cycle but also enhances the direct applicability of intelligence, enabling intelligence requesters to respond more quickly.
[0082] The solution of the present invention has the following important significances:
[0083] (1) Improving the accuracy of national science and technology intelligence governance: The research on the identification of the intelligence value of science and technology bulletins can significantly improve the accuracy and efficiency of national science and technology intelligence governance through the application of large language models. In the context of increasingly fierce international science and technology competition, it is particularly important to accurately capture and analyze key information in science and technology bulletins. This research not only helps intelligence workers keep abreast of the trends of scientific and technological development and technological dynamics in a timely manner but also supports the government in making scientific and reasonable decisions when formulating and adjusting science and technology policies, promoting the introduction of relevant policies and the effective allocation of resources.
[0084] (2) Promoting scientific and technological innovation and ensuring scientific and technological security: The research on the identification of the intelligence value of science and technology bulletins is of great significance in promoting scientific and technological innovation and ensuring scientific and technological security. By efficiently identifying and analyzing innovation information, emerging technologies and cutting-edge scientific achievements can be discovered in advance, providing inspiration and direction for national scientific and technological innovation. At the same time, combined with the early warning ability of large language models, potential scientific and technological risks can be identified and prevented, ensuring that the development and application of technologies do not lag behind other countries and enhancing the guarantee of scientific and technological security.
[0085] (3) Improving decision-making efficiency and promoting the development of the intelligence discipline: The timely identification and analysis of the intelligence of science and technology bulletins significantly improve the accuracy and efficiency of decision-making, enabling decision-makers to quickly understand industry dynamics and scientific and technological progress and make timely and accurate decisions. In addition, the research on the identification of the intelligence value of science and technology bulletins also promotes the development of the intelligence discipline, enriches the theories and methods of information science, improves the scientificity and effectiveness of intelligence analysis, and promotes the theoretical innovation and practical progress of the intelligence discipline.
[0086] (4) Optimizing resource allocation and promoting the optimization of scientific and technological resources: Through the application of large language models, the identification of the value of scientific and technological intelligence can achieve the rational allocation and optimal utilization of resources. Scientific intelligence analysis not only helps scientific research institutions and enterprises identify high-value research fields and projects, avoid resource waste, but also can concentrate efforts on overcoming key technical problems, improve resource utilization efficiency, promote the optimal allocation of scientific and technological resources, and promote scientific and technological innovation and development.
[0087] The units involved in the embodiments of the present invention can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the units themselves in some cases.
[0088] According to one aspect of the embodiments of the present invention, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various alternative implementation manners.
[0089] As another aspect, the embodiments of the present invention further provide a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the methods described in the above embodiments.
Claims
1. A scientific and technological news intelligence perception system based on a large language model, characterized in that: include: Science and technology news monitoring subsystem, science and technology news identification subsystem and science and technology news visualization subsystem; The science and technology news monitoring subsystem includes an information collection module, a data standardization module and a database. In the information collection module, the system operates according to a predetermined process to obtain multimodal science and technology news data; in the data standardization module, the multimodal data output by the information collection module is received to perform data cleaning, structured processing and semantic standardization operations; after standardization processing, the data is stored in the database; The science and technology news identification subsystem includes an intelligence value screening module and an evaluation system module; the intelligence value screening module identifies the intelligence value in combination with the evaluation system; The science and technology news visualization subsystem includes a visualization module; the science and technology news visualization subsystem receives filtered information from a database and realizes multi-dimensional visualization presentation through the visualization module.
2. The technology news intelligence perception system based on a large language model according to claim 1 is characterized in that: The method of obtaining multimodal science and technology news data specifically includes: directly sending a request to a target science and technology news source through an RSS subscription interface, and receiving a structured XML data stream; then parsing the received data, matching predefined subject tags and keywords, extracting content related to the science and technology field, and storing it in a temporary storage area; at the same time, the web page crawling function is started, and based on the configured target URL list, the web page HTML data is obtained through a crawler framework; using DOM tree parsing, the body area is extracted through XPath or CSS selectors, and the text is segmented in combination with an NLP model, and the content that meets the requirements is filtered and stored in the temporary storage area.
3. The technology news intelligence perception system based on a large language model according to claim 1 is characterized in that: The acquisition of multimodal technology news data specifically includes: setting a timed task scheduling module to automatically trigger crawler instances at set time intervals, updating crawling strategies in real time, dynamically adjusting target keyword weights, and giving priority to websites with high-frequency updates; the processing of audio data uses a speech recognition model to accurately transcribe the audio content into text; the processing of video data obtains key frames through frame extraction technology, combines image target detection algorithms and description generation models to generate text information, and integrates it with original text and metadata through semantic fusion technology, and uniformly stores it in a multimodal database.
4. The technology news intelligence perception system based on a large language model according to claim 1 is characterized in that: In the data standardization module, the data cleaning, structured processing and semantic standardization operations are performed, specifically including: First, run the noise filtering algorithm to analyze the text, remove redundant information by calculating the word weights, and call the local sensitive hashing algorithm to detect and delete duplicate content; in the formatting and normalization phase, use regular expressions to extract key fields, and use the large language model with the prompt project to convert unstructured text into a standardized JSON data format; Subsequently, the SAO triple extraction operation is performed, and the text dependency tree is constructed through dependency syntactic analysis to identify the subject, predicate and object in the sentence; the multi-level semantic relations in the complex sentence are disassembled and the triples are extracted one by one through the generative language model to ensure the content integrity and semantic accuracy; in the process of multimodal data fusion, the multimodal deep learning model is loaded, and the audio-to-text and image description text are aligned with the original text through feature vector matching technology, and a cross-attention mechanism is used to generate a unified semantic representation.
5. The technology news intelligence perception system based on a large language model according to claim 1 is characterized in that: The database specifically adopts a distributed database architecture, ensures the horizontal scalability and disaster tolerance of data through sharding and replication mechanisms, and optimizes the execution efficiency of complex queries in a columnar storage manner.
6. The technology news intelligence perception system based on a large language model according to claim 5 is characterized in that: The operation logic of the database specifically includes data storage, index construction and efficient retrieval; During the data storage phase, the flash data output by the data standardization processing module is first partitioned through the database preprocessing interface; based on predefined classification rules, the data is assigned to the selected logical partitions, and the data is split into field storage using column storage to reduce the I / O burden; each record is assigned a unique identifier, and the operation details are recorded through transaction logs to ensure the atomicity and consistency of writing; for multimodal data, a nested JSON structure is used to uniformly store text, audio-to-text, image descriptions and other content, and a modal association table is introduced to establish a cross-modal logical mapping relationship; In the index construction phase, the database module covers keyword query and semantic retrieval scenarios through the dual index mechanism of inverted index and semantic index. In the inverted index construction process, the word segmentation algorithm BERT word segmenter is used to break down the text content into word or phrase units and record their positions in the data set. The weight of each keyword is calculated by the BM25 algorithm and stored in the inverted list to support fast relevance query. At the same time, the semantic index module uses the Transformer model BERT and RoBERTa to generate high-dimensional semantic vectors for the text content, and uses PCA dimensionality reduction processing to reduce storage overhead. The vectors after dimensionality reduction are stored in the vector database FAISS, and combined with the efficient ANN retrieval algorithm to achieve fast matching of high-dimensional vectors. In the data retrieval stage, the user query request is first analyzed for its intent by the query parsing module and parsed into a structured query request; for keyword queries, the inverted index is directly called, and the results are filtered and sorted through Boolean logic and relevance scoring; for semantic retrieval, the query text entered by the user is generated into a vector representation through the semantic embedding model, and the similarity is calculated with the vectors already stored in the vector database, and the HNSW graph retrieval algorithm is used to quickly find the most relevant records; the retrieval results are re-sorted by the sorting module, and the sorting criteria are dynamically adjusted according to time priority, relevance priority or user-defined weight strategy.
7. The technology news intelligence perception system based on a large language model according to claim 1 is characterized in that: In the cache mechanism design of the database, the results of high-frequency queries are cached in the memory database Redis and directly returned in similar requests; at the same time, the database module regularly performs index reconstruction and storage partition merging, and ensures the efficiency and stability of the system for long-term operation by optimizing the disk layout and data structure.
8. The technology news intelligence perception system based on a large language model according to claim 1 is characterized in that: The intelligence value screening module identifies intelligence value in combination with the evaluation system, specifically including: the intelligence value screening mechanism realizes accurate identification through three-layer filtering. First, the first layer sets the intelligence value threshold filtering module to set the benchmark threshold and conduct preliminary screening of input information; adopts the feature extraction method based on word frequency-inverse document frequency to calculate the weight score of text keywords; when the comprehensive score of information is lower than the set threshold, it is marked as low-value information and filtered; The second layer uses regular matching pattern filtering to maintain a regular expression rule library that contains common expression patterns and key phrases in science and technology news. It identifies key information in the text through pattern matching. The higher the matching degree, the higher the professionalism and technical content of the information. It also uses a dynamic rule update mechanism to continuously optimize the rule library based on newly added samples. Intelligence value sorting and screening is the last layer of filtering, and a comprehensive evaluation is conducted in combination with multiple dimensions of the evaluation system module. Among them, the importance dimension is calculated by keyword weight and citation frequency; the relevance is measured based on the cosine similarity algorithm to measure the matching degree with the user field; novelty is evaluated by time decay function and innovation point extraction; reliability is judged based on the credibility of the source and the consistency of the content; trend is determined by the development direction of time series analysis and hot spot clustering identification technology; Secondly, an intelligence identification weight allocation module is set up, and the hierarchical analysis method is used to determine the weight of each dimension. First, a judgment matrix is established, and the relative importance of dimensions is determined through expert scoring. After a consistency test, the characteristic vector is calculated to obtain the weight coefficient of each dimension, and the weight configuration is dynamically adjusted according to different application scenarios.
9. The technology news intelligence perception system based on a large language model according to claim 1 is characterized in that: The multi-dimensional visualization presentation is realized by the visualization module, specifically including: Set up a large model intelligence overview module, call the large language model API, and perform semantic understanding and reorganization of high-value news; build a standardized prompt template to guide the model to generate structured overview content; extract key information from the text through the attention mechanism, combine the knowledge graph technology to perform correlation analysis on the information, and finally generate a summary report containing core viewpoints, technical features and development trends; A heat map module is set up to map the news data to the world map according to geographic coordinates based on the geographic information processing engine. The kernel density estimation algorithm is used to calculate the information density, and an adaptive color mapping scheme is adopted to display different colors in hot spots. The evolution of the scientific and technological activity in each region is dynamically displayed by sliding the time window. At the same time, multi-level drilling is supported, and users can gradually drill down from the global view to specific countries and regions. A ranking module is set up to achieve dynamic sorting based on the time dimension. Specifically, the sliding time window method is used to aggregate and count data of different time granularities. A weighted sorting algorithm is used to generate the ranking list by combining the intelligence value score and the large model evaluation results. The interface adopts a responsive design, supports switching between list view and card view, and provides trend charts to display historical changes. Set up the word cloud module, use the word segmentation engine to process the text, and calculate the word weight through the TF-IDF algorithm; use the layout algorithm to optimize the word placement while ensuring visual beauty; support multiple shape templates at the same time, and the color scheme is matched according to the word category; use the force-directed algorithm to optimize the word spacing to ensure a compact and clear layout; A scientific and technological intelligence map module is set up, and a deep learning model is used to realize SAO triple extraction; a graph database is used to store entity relationships, and node positions are calculated through a graph layout algorithm; the visualization layer uses WebGL to achieve smooth rendering of large-scale maps and real-time incremental updates, so that new news data can be dynamically integrated into the existing map structure.
Citation Information
Cited By
Technology development situation awareness system and method
CN120429414A
A technology development situation awareness system and method
CN120429414B
Multi-agent collaborative decision-making intelligence defense method, system and device and storage medium
CN120562406A
Multi-agent collaborative decision-making intelligence defense methods, systems, devices and storage media
CN120562406B
Electric power scientific research information retrieval result optimization method and device
CN120804345A