Agent-based industry brief report generation method and system

Through the Agent-based industry briefing generation method, crawler and big model technology, the problems of low efficiency and unstable content quality of traditional industry briefing generation methods are solved, and the rapid collection and accurate output of industry information is achieved, and user decision-making support capabilities are improved.

CN120068832APending Publication Date: 2025-05-30SHENZHEN VIRTUAL CLUSTERS INFORMATION TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510133760.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional industry briefing generation method relies on manual data collection, is inefficient and easy to miss information, and lacks intelligent classification and analysis methods, resulting in unstable content quality and inability to accurately grasp industry trends and business opportunities.

Method used

Agent-based industry briefing generation method is adopted to collect data from multiple target industry information websites through crawling technology, use large models to extract key information and store them in the database, receive search instructions for matching, and output structured industry briefing documents.

Benefits of technology

It realizes the rapid collection and integration of a large amount of industry information, provides retrieval functions, and users can quickly find and accurately obtain relevant information data, improving the content quality of industry briefings and user decision support capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068832A_ABST
    Figure CN120068832A_ABST
Patent Text Reader

Abstract

The invention discloses an Agent-based industry brief report generation method and system, and the method comprises the steps: collecting information data of a plurality of target industry information websites, and extracting corresponding key information from a plurality of pieces of information data, the method comprises the following steps: acquiring key information, storing the key information and corresponding information data in a preset database in the form of structured data, receiving a retrieval instruction, matching corresponding information data in the database according to the retrieval instruction, extracting the key information of the matched information data, and inputting the key information into a preset report template, according to the method and the system, the industry brief report document is output, rapid collection and integration of a large amount of industry information are realized, and a retrieval function is provided, so that a user can rapidly search and accurately obtain corresponding information data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of retrieval management systems, and particularly to an Agent-based industry briefing generation method and system. Background Art

[0002] In the current era of information explosion, industry briefings, as a tool for quickly transmitting industry information and assisting enterprise decision-making, are becoming increasingly important. However, the traditional methods for generating industry briefings face many dilemmas. On the one hand, data collection mainly relies on manual labor. Staff need to screen and collect industry-related information one by one from a vast amount of network resources, paper materials, etc. This not only consumes a large amount of human and time costs, but also easily leads to information omission or incomplete collection due to human negligence, and fails to cover some fragmented but valuable industry dynamics information on emerging social media platforms.

[0003] On the other hand, most of the existing briefing generation processes lack effective intelligent classification and analysis means. When organizing the collected information, the analysis link often relies on manual experience judgment, making it difficult to deeply explore complex data relationships, accurately grasp industry trends and potential business opportunities, and as a result, the quality of the generated briefing content is unstable. Some briefing contents are empty and lack in-depth analysis, while some are out of touch with user needs and cannot provide accurate and effective decision-making support for users, severely restricting the role of industry briefings in enterprise development. Summary of the Invention

[0004] The object of the present invention is to propose an Agent-based industry briefing generation method, system, device and medium for the technical problems existing in the background art.

[0005] To achieve the above technical object, the technical solution adopted by the present invention is as follows:

[0006] The first implementation manner of the first aspect of the present invention provides an Agent-based industry briefing generation method, which includes:

[0007] Collect information data from multiple target industry information websites;

[0008] Extract corresponding key information from multiple information data, and store these key information and the corresponding information data in a preset database in the form of structured data;

[0009] Receive a retrieval instruction, and match the corresponding information data in the database according to the retrieval instruction;

[0010] Extract the key information of the matched information data, and input these key information into a preset report template to output an industry briefing document.

[0011] Optionally, in the second implementation manner of the first aspect of the present invention, collecting information data of multiple target industry information websites includes:

[0012] Based on web crawler technology, crawl the information data of multiple target industry information websites.

[0013] Optionally, in the third implementation manner of the first aspect of the present invention, extracting corresponding key information from multiple pieces of information data includes:

[0014] Parse multiple target industry information websites, deconstruct the corresponding tag data, and extract text fragments from the tag data as key information.

[0015] Optionally, in the fourth implementation manner of the first aspect of the present invention, storing these key information and the corresponding information data in a preset database in the form of structured data includes:

[0016] Based on a preset design large model, extract these key information, construct a database, and store these key information and the corresponding information data in the database in the form of structured data, where the information data includes the original website, web page content, and publication time.

[0017] Optionally, in the fifth implementation manner of the first aspect of the present invention, receiving a retrieval instruction and matching corresponding information data in the database according to the retrieval instruction includes:

[0018] Construct a retrieval vocabulary library storing multiple industry keywords and provide the retrieval vocabulary library to the user;

[0019] Receive a retrieval instruction and perform a matching search in the database according to the industry keywords in the retrieval instruction to obtain corresponding information data.

[0020] Optionally, in the sixth implementation manner of the first aspect of the present invention, extracting the key information of the matched information data includes:

[0021] Based on the design large model, parse the retrieved and matched information data to obtain multiple key information, where the key information includes the title, time, category, and the section where the information is located.

[0022] The first implementation manner of the second aspect of the present invention provides an Agent-based industry briefing generation system, including an information collection module and a retrieval module. The information collection module includes a crawler unit, a first large model unit, and a data storage unit, and the retrieval module includes a database detection unit and a report generation unit;

[0023] The crawler unit is used to collect information data of multiple target industry information websites;

[0024] The first large model unit is used to respectively extract corresponding key information from multiple information data;

[0025] The data storage unit is used to store these key information and the corresponding information data in the database in the form of structured data;

[0026] The database detection unit is used to receive a retrieval instruction and match corresponding information data in the database according to the retrieval instruction;

[0027] The report generation unit is used to extract the key information of the matched information data and input these key information into a preset report template to output an industry briefing document.

[0028] Optionally, in the second implementation manner of the second aspect of the present invention, the information collection module further includes a web page parsing unit, and the retrieval module further includes a keyword unit and a second large model unit;

[0029] The crawler unit is further used to crawl the information data of multiple target industry information websites based on the crawler technology;

[0030] The web page parsing unit is used to parse multiple target industry information websites, deconstruct the corresponding tag data, and extract text fragments from the tag data as key information;

[0031] The first large model unit is further used to extract these key information based on a preset design large model, construct a database, and store these key information and the corresponding information data in the database in the form of structured data, where the information data includes the original website, web page content, and publication time;

[0032] The keyword unit is used to construct a retrieval vocabulary library storing multiple industry keywords and provide the retrieval vocabulary library to the user;

[0033] The data storage unit is further used to receive a retrieval instruction and perform a matching retrieval in the database according to the industry keyword in the retrieval instruction to obtain corresponding information data;

[0034] The second large model unit is used to parse the information data retrieved and matched based on the design large model to obtain multiple key information, where the key information includes the title, time, category, and information section;

[0035] The present invention has the following beneficial technical effects compared with the prior art: collecting information data from multiple target industry information websites, respectively extracting corresponding key information from the multiple information data, and storing these key information and the corresponding information data in a preset database in the form of structured data, receiving a retrieval instruction and matching the corresponding information data in the database according to the retrieval instruction, extracting the key information of the retrieved information data, and inputting these key information into a preset report template to output an industry briefing document, realizing the rapid collection and integration of a large amount of industry information, and providing a retrieval function to facilitate users to quickly find and accurately obtain the corresponding information data. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic diagram of the method for generating an industry briefing based on an Agent in an embodiment of the present invention;

[0037] Figure 2 It is a schematic diagram of the system for generating an industry briefing based on an Agent in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0039] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0040] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or a specific connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0041] For ease of understanding, the specific process of the embodiments of the present invention in the first aspect will be described below. Please refer to Figure 1 - Figure 2 , in the embodiments of the present invention, the Agent-based industry briefing generation method includes:

[0042] 101. Collect information data from multiple target industry information websites;

[0043] In this embodiment, the Agent-based industry briefing generation method is based on Agent technology, leveraging the autonomy of Agent and the powerful functions of langchain. The overall architecture follows the modular design principle to ensure that each functional module can operate independently and cooperate with each other. By introducing langchain technology, with its rich tool chain and pre-built modules, the information collection and processing processes for various industries are completed.

[0044] The Agent adopts a hierarchical modular design, including an interaction layer, a control layer, and an execution layer. The interaction layer receives instructions such as set industry goals, website lists, and keyword preferences; the control layer formulates detailed information collection, parsing, extraction, and archiving strategies based on the instructions received by the interaction layer, and coordinates the work processes of each module in the execution layer; the execution layer consists of multiple functional modules, including an industry crawler module, a web page parsing module, a large model keyword extraction module, a data storage module, etc., which are responsible for specifically executing various operations of information collection and processing.

[0045] Further, step 101 can specifically also execute:

[0046] Based on crawler technology, crawl the information data from multiple target industry information websites;

[0047] In this embodiment, under the scheduling of the control layer, the industry crawler module starts the crawling task for information websites related to the target industry. In the langchain ecosystem, existing web crawler tools such as BeautifulSoup and Scrapy are integrated. At the same time, using the asynchronous processing ability of langchain, multi-page crawling is realized to improve efficiency and reduce waiting time.

[0048] The control layer parses the keywords in the pre-set industry list, searches for relevant content on mainstream portal websites or search engines, and saves the information of relevant links.

[0049] 102. Extract the corresponding key information from multiple information data respectively, and store these key information and the corresponding information data in a pre-set database in the form of structured data;

[0050] Further, step 102 can specifically also execute:

[0051] Parse multiple target industry information websites, deconstruct the corresponding tag data, and extract text fragments from the tag data as key information;

[0052] Based on a pre-set design large model, extract these key information, build a database, and store these key information and the corresponding information data in the database in the form of structured data, where the information data includes the original website, web page content, and publication time.

[0053] In this embodiment, after obtaining the link information through web crawling, it is quickly transmitted to the basic module based on HTML and XML parsing in the execution layer. The module constructs the DOM tree structure of the web page according to the syntax rules of markup languages such as HTML and XML, and accurately locates to such as <h1>Title label, Common information - carrying tags such as paragraph tags are used to extract preliminary text fragments.

[0054] For unstructured data such as pictures and videos, image recognition and video content analysis technologies (combined with deep - learning models) are used to extract relevant text descriptions or key feature information and convert them into structured data for subsequent processing.

[0055] To ensure the extraction accuracy and consistency of the large model for articles, the Qwen2.5 - 7B - Instruct large model after self - fine - tuning is adopted. During the data production process, the currently best - performing GPT4o is used as the source of data tags to avoid complex and cumbersome manual review. In the fine - tuning stage, by designing prompt words, the large model is guided to focus on the extraction of key information in the industry field, avoiding interference from irrelevant information and improving the extraction accuracy. The production of the fine - tuning dataset is as follows:

[0056] [{"instruction": "The following is the content of a web page. According to this content, extract the keywords related to the industry. Only extract the keywords without modifying the words, and limit them to be highly relevant to the industry. For example, if the input is a blog related to lithium batteries, the extracted output keywords are new energy, ternary lithium, lithium iron phosphate",

[0057] "input": "<Input the web page content, such as an article related to blockchain>",

[0058] "output": "blockchain; smart contract; Ethereum; cryptocurrency"}]

[0059] The fine - tuning of the Qwen2.5 - 7B - Instruct large model adopts the LoRA fine - tuning method. On the one hand, during fine - tuning, the calculation and memory costs are high. LoRA decomposes the part of the weight change into a low - rank representation, without directly calculating the complete weight update matrix, greatly reducing the memory occupation and calculation amount.

[0060] The processed structured data enters the execution layer. The data storage module in the execution layer constructs an object containing fields of "keyword", "link" and "date" according to the JSON format requirements. If there are multiple combinations of keywords and links, they are organized into a JSON array for convenient batch storage. For example:

[0061]

[0062]

[0063] Subsequently, a database connection tool (such as the pymysql library in Python combined with the connection code adapted by the Agent itself) is used to establish a connection with the database, and an insert statement is executed to store the data into the database table. And a secondary verification will also be performed after the storage is completed. By querying the data just stored and comparing it with the original data, it is ensured that there are no problems such as data loss or incorrect conversion during the storage process.

[0064] 103. Receive a retrieval instruction and match the corresponding information data in the database according to the retrieval instruction;

[0065] In this embodiment, the retrieval function is also constructed using Agent technology, which has the same idea as the design information collection. The retrieval-generated Agent realizes the retrieval of industry-related information for the keywords provided by the user, associates and sorts the user's keywords and the keywords in the database, and arranges them in descending order according to the number of matching keywords. The number of matching keywords can generally represent the relevance to a certain extent. Parse the matching web links, and the parsing process is the same as the web page parsing in the information collection Agent. Extract the information therein and briefly summarize the entire content to form structured data for a single link. Finally, integrate these structured data to form a readable industry briefing.

[0066] Furthermore, step 103 can specifically also execute:

[0067] Construct a retrieval vocabulary library storing multiple industry keywords and provide the retrieval vocabulary library to the user;

[0068] Receive a retrieval instruction and perform a matching retrieval in the database according to the industry keywords in the retrieval instruction to obtain the corresponding information data.

[0069] In this embodiment, when the user inputs one or more keywords, each keyword is first preprocessed. Remove irrelevant characters such as spaces and punctuation marks before and after the keyword, and uniformly convert the keyword to lowercase for subsequent exact matching. When the user inputs multiple keywords, such as "artificial intelligence big data", first split the multiple keywords into a single keyword list, and calculate the number of matches for each keyword with the database records respectively.

[0070] Finally, on the basis of counting and sorting the number of keyword matches, further optimize the association sorting. Consider the position factor of the keyword in the database record. If the keyword appears in a key position such as the title or the beginning paragraph, give a certain weight bonus. Finally, arrange them in descending order according to the number of matching keywords and the relevant weights, and return a list according to the Topk threshold input by the user.

[0071] 104. Extract the key information from the matched information data and input these key information into a preset report template to output an industry briefing document.

[0072] Further, step 104 can specifically also execute:

[0073] Based on the design large model, parse the retrieved and matched information data to obtain multiple key information. Among them, the key information includes the title, time, category, and the section where the information is located.

[0074] In this embodiment, the text data processed by the web parsing module enters the large model data structuring module. In order to more accurately achieve the purpose of extracting specific targets, the large model adopted is also based on the Qwen2.5-7B-Instruct large model after its own fine-tuning. Call the API of GPT4o to perform the tagging work on the article. In the fine-tuning stage, design appropriate prompt words to stimulate the ability of the large model to process this task.

[0075] Specifically, it is manifested by collecting the information data of multiple target industry information websites, respectively extracting the corresponding key information from multiple information data, storing these key information and the corresponding information data in a preset database in the form of structured data, receiving a retrieval instruction, matching the corresponding information data in the database according to the retrieval instruction, extracting the key information of the matched information data, and inputting these key information into a preset report template to output an industry briefing document, realizing the rapid collection and integration of a large amount of industry information, and providing a retrieval function to facilitate users to quickly find and accurately obtain the corresponding information data.

[0076] The specific process of the embodiment of the second aspect of the present invention is described. The Agent-based industry briefing generation system includes an information collection module 200 and a retrieval module 300. The information collection module includes a crawler unit 201, a first large model unit 202, and a data storage unit 203. The retrieval module includes a database detection unit 301 and a report generation unit 302;

[0077] The crawler unit 201 is used to collect the information data of multiple target industry information websites;

[0078] The first large model unit 202 is used to respectively extract the corresponding key information from multiple information data;

[0079] The data storage unit 203 is used to store these key information and the corresponding information data in the database in the form of structured data;

[0080] The database detection unit 301 is used to receive a retrieval instruction and match the corresponding information data in the database according to the retrieval instruction;

[0081] The report generation unit 302 is used to extract the key information of the matched information data, and input these key information into a preset report template to output an industry briefing document.

[0082] Furthermore, the information collection module 200 further includes a web page parsing unit 204, and the retrieval module further includes a keyword unit 303 and a second large model unit 304.

[0083] The crawler unit 201 is also used to crawl the information data of multiple target industry information websites based on the crawler technology;

[0084] The web page parsing unit 204 is used to parse multiple target industry information websites, deconstruct the corresponding tag data, and extract text fragments from the tag data as key information;

[0085] The first large model unit 202 is also used to extract these key information based on a preset design large model, construct a database, and store these key information and the corresponding information data in the database in the form of structured data, where the information data includes the original website, web page content, and publication time;

[0086] The keyword unit 303 is used to construct a retrieval vocabulary library storing multiple industry keywords and provide the retrieval vocabulary library to the user;

[0087] The data storage unit 203 is also used to receive a retrieval instruction, and perform a matching retrieval in the database according to the industry keyword in the retrieval instruction to obtain the corresponding information data;

[0088] The second large model unit 304 is used to parse the information data retrieved and matched based on the design large model to obtain multiple key information, where the key information includes the title, time, category, and the section where the information is located.

[0089] Specifically, it is realized by collecting the information data of multiple target industry information websites, respectively extracting the corresponding key information from multiple information data, storing these key information and the corresponding information data in the preset database in the form of structured data, receiving a retrieval instruction and matching the corresponding information data in the database according to the retrieval instruction, extracting the key information of the retrieved and matched information data, and inputting these key information into the preset report template to output an industry briefing document, so as to quickly collect and integrate a large amount of industry information and provide a retrieval function to facilitate the user to quickly find and accurately obtain the corresponding information data.

[0090] The above is a method for generating industry briefings based on Agents or multiple implementation manners provided in combination with specific contents, and it is not considered that the specific implementation of the present invention is only limited to these descriptions. Any approximation, similarity, or several technical deductions or replacements made under the premise of the concept of the present invention should be regarded as the protection scope of the present invention. < / h1>

Claims

1. An agent-based industry briefing generation method, characterized in that: include: Collect information data from multiple target industry information websites; Extracting corresponding key information from the plurality of information data respectively, and storing the key information and the corresponding information data in a preset database in the form of structured data; Receiving a search instruction, and matching corresponding information data in the database according to the search instruction; Extract key information of the matched information data, and input the key information into a preset report template to output an industry briefing document.

2. The method for generating an industry briefing based on Agent according to claim 1, characterized in that: The information data collected from multiple target industry information websites includes: Based on crawler technology, information data of multiple target industry information websites are crawled.

3. The method for generating an industry briefing based on Agent according to claim 2, characterized in that: The extracting corresponding key information from the plurality of information data respectively comprises: A plurality of target industry information websites are parsed to deconstruct corresponding label data, and text fragments are extracted from the label data as key information.

4. The method for generating an industry briefing based on Agent according to claim 2, characterized in that: The storing of the key information and the corresponding information data in the preset database in the form of structured data includes: Based on the preset design model, the key information is extracted and a database is constructed. The key information and the corresponding information data are stored in the database in the form of structured data, wherein the information data includes the original website, web page content and publication time.

5. The method for generating an industry briefing based on Agent according to claim 4, characterized in that: The receiving of the search instruction and matching the corresponding information data in the database according to the search instruction includes: Constructing a search vocabulary library storing a plurality of industry keywords, and providing the search vocabulary library to users; A search instruction is received, and a matching search is performed in the database according to the industry keywords in the search instruction to obtain the corresponding information data.

6. The method for generating an industry briefing based on Agent according to claim 5, characterized in that: The key information of the information data extracted and matched includes: Based on the design macro model, the information data retrieved and matched is parsed to obtain a plurality of key information, wherein the key information includes title, time, category and the section where the information is located.

7. An agent-based industry briefing generation system, characterized in that: It includes an information collection module and a retrieval module, wherein the information collection module includes a crawler unit, a first large model unit and a data storage unit, and the retrieval module includes a database detection unit and a report generation unit; The crawler unit is used to collect information data from multiple target industry information websites; The first large model unit is used to extract corresponding key information from the plurality of information data respectively; The data storage unit is used to store the key information and the corresponding information data in the database in the form of structured data; The database detection unit is used to receive a search instruction and match corresponding information data in the database according to the search instruction; The report generating unit is used to extract key information of the matched information data, and input the key information into a preset report template to output an industry briefing document.

8. The agent-based industry briefing generation system according to claim 7, characterized in that: The information collection module also includes a web page parsing unit, and the search module also includes a keyword unit and a second large model unit; The crawler unit is also used to crawl information data of multiple target industry information websites based on crawler technology; The web page parsing unit is used to parse multiple target industry information websites, deconstruct corresponding tag data, and extract text fragments from the tag data as key information; The first large model unit is further used to extract the key information based on the preset design large model, and construct a database, and store the key information and the corresponding information data in the database in the form of structured data, wherein the information data includes the original website, webpage content and publishing time; The keyword unit is used to construct a search vocabulary library storing multiple industry keywords, and provide the search vocabulary library to the user; The data storage unit is also used to receive a search instruction, and perform a matching search in the database according to the industry keywords in the search instruction to obtain the corresponding information data; The second large model unit is used to parse the information data retrieved and matched based on the design large model to obtain multiple key information, wherein the key information includes title, time, category and the section where the information is located.

Citation Information

Cited By

  • Report generation method and device and electronic equipment

    CN121093925A