LLM-based public source webpage information analysis method and device, medium and product
Through the open source web information analysis method combining LLM and traditional algorithms, the problem of redundant information impact and insufficient flexibility of traditional methods is solved, efficient web information extraction and structured data generation are realized, and multi-web page joint information processing is supported for open source intelligence analysis.
Patent Information
- Application Number
- CN202510447301.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has redundant information in the analysis of public source web page information that affects the acquisition of subject information, traditional methods are insufficient in flexibility, and LLM intelligent web page content extraction capabilities, especially in the extraction of long text and multi-web page joint information.
Using the LLM-based public source web information analysis method, combined with traditional web information extraction algorithms and search enhancement generation technology, efficient extraction of web information and structured data generation through preprocessing, semantic understanding and logical thinking capabilities.
It improves the accuracy and applicability of web page information processing, can generate structured data in long text and multi-web page joint intelligence analysis, and supports hot events and character analysis summary reports for open source intelligence analysis.
Smart Images

Figure CN120296229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of open-source intelligence analysis, and particularly to a method, device, medium, and product for analyzing public-source web page information based on LLM. Background Art
[0002] The statements in this section only provide background information related to the present disclosure and may not constitute prior art.
[0003] The analysis of public-source web page information plays an irreplaceable role in the field of open-source intelligence analysis. Web page information analysis is not only an important way to obtain valuable intelligence, but also a key means to construct a comprehensive situation awareness map, predict trends and patterns, and verify and supplement other types of intelligence. With the continuous development of information technology and the increasing popularity of the Internet, the importance of web page information analysis in the field of open-source intelligence analysis will become more prominent. However, in real public-source web pages, there is an excessive amount of redundant information, such as in-text graphic and text advertisements, web page recommendations arranged in the sidebar, colorful web page animations, and the display of various audio and video information. These redundant information provide certain convenience for people's web browsing, but inevitably have a greater impact on obtaining the main information of the web page, especially when extracting web page information in batches.
[0004] LLM is the abbreviation of Large Language Model. Generally, it has a huge parameter scale (compared with traditional AI neural network models), has certain semantic understanding and logical thinking abilities, and can generate corresponding responses according to the input instructions (prompts) of users. Since 2023, with the significant improvement of LLM capabilities, LLM has been widely used and played an important role in many industries such as customer service, finance, and healthcare. Based on the semantic understanding ability and instruction-following ability of LLM, the flexibility and high customization of public-source web page information mining and analysis can be improved.
[0005] Web page information extraction has always been the core research and application of web crawler systems and is the main way for people to obtain information in large quantities from the Internet. Currently, the display basis of Internet web pages is HTML text, and various browsers render the page according to HTML text to obtain the web page information seen daily. Therefore, HTML text is often used as the source of web page information extraction. Using LLM to read and understand HTML text, identify content, and then implement extraction is beneficial to generating structured text data and storing it in a vector library for answering questions raised by users. The traditional web page information analysis mainly has the following three methods:
[0006] (1) Extract content using the xpath of fixed modules in HTML text: Since HTML is a tagged text with a distinct organizational structure, using certain specific tags such as 、、 and It is possible to organize and display the text structure, and develop a corresponding information extraction method for a fixed template of a certain web page HTML text, which can effectively extract the information therein. However, this method has great limitations. When the template of the website changes greatly, a completely new information extraction method needs to be developed specifically. The form is too rigid, and the writing of the information extraction method is time-consuming. Therefore, it is not suitable for large-scale network information extraction.
[0007] (2) Another relatively common network information extraction method is to use statistical methods to process the HTML text extraction algorithm. Common ones include the Readability algorithm and the web page body extraction method based on text and symbol density. These traditional algorithms mainly target the content in each text tag block of the web page HTML text, calculate the value score of the tag block according to the preset algorithm, and then eliminate the relatively "low-value" tag blocks in the page according to the score, and retain the "high-value" tag blocks as the main content of the web page for extraction. For example, the Readability algorithm has been widely used in the "pure mode" and other settings of various browsers. The advantages of this type of algorithm are wide application, few restrictions, and certain guarantee of the effect; the disadvantages are that the effect of the algorithm needs to be manually fine-tuned, and the extraction effect for websites with inconspicuous text and complex structures is not good.
[0008] (3) With the application and development of LLM, intelligent web page content extraction has also developed vigorously, and a variety of mature web page content extraction frameworks have emerged, such as Scrapegraph-ai, etc. These frameworks often provide an LLM-empowered intelligent crawler framework. The advantages are high versatility and the ability to output diversely according to user needs; the disadvantage is that the long text extraction ability is poor, and it is better at extracting fragmented content. Summary of the Invention
[0009] The purpose of the present invention is to: aiming at the problems existing in the prior art, integrating the accuracy of traditional web page information analysis and the flexibility of LLM intelligent web page content extraction, provide a method, device, medium, and product for analyzing open-source web page information based on LLM, and taking open-source intelligence analysis application as the traction, focusing on the realization of open-source web page information analysis.
[0010] The technical solution of the present invention is as follows:
[0011] A method for analyzing open-source web page information based on LLM, comprising:
[0012] Step S1: Preprocess the web page HTML text and extract the web page information contained in the HTML text;
[0013] Step S2: Perform web information query retrieval and crawling based on the user's question, and use the semantic understanding ability of the LLM to retrieve similar reference information from the web information;
[0014] Step S3: Based on the retrieved similar reference information, use the logical thinking ability of the LLM to reason and analyze the user's question, extract the reference information involved in the question, and give a feedback reply.
[0015] Furthermore, the preprocessing method includes:
[0016] a. Use traditional web extraction algorithms to parse and extract the theme content of the web page;
[0017] b. First, perform cleaning to delete the content unrelated to the main text in the web page HTML text, and then perform segmentation and storage.
[0018] Furthermore, the traditional web extraction algorithms include: Jina Reader web parsing algorithm and Readability algorithm.
[0019] Furthermore, the cleaning process is as follows:
[0020] Remove the format definitions and JS scripts in the HTML text;
[0021] Remove the label modules unrelated to the overall content of the web page;
[0022] Standardize the redundant special symbols in the HTML text.
[0023] Furthermore, the vectorization model used in the segmentation and storage process is the sentense-transformer / all-mpnet-base-v2 model, and the vector library used is Milvus.
[0024] Furthermore, use retrieval-augmented generation technology to perform segmentation and storage processing on the HTML text.
[0025] Furthermore, using retrieval-augmented generation technology to perform segmentation and storage processing on the HTML text includes:
[0026] Adopt the function of semantic similarity query for data recall. During the recall process, select a certain number of similar texts and place them into the prompt words as a reference for the LLM to generate structured data. If the user's question is related to the content of the main text, place the main text information extracted by the traditional web extraction algorithm into the prompt words as another reference material;
[0027] After referring to these materials, the LLM extracts structured data for the user's question.
[0028] The present invention also proposes a computing device, including:
[0029] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a method for analyzing public source web page information based on LLM as described above.
[0030] The present invention also provides a computer terminal storage medium storing computer terminal executable instructions for executing a method for analyzing public source web page information based on LLM as described above.
[0031] The present invention also provides a computer program product, wherein when the computer program is executed by a processor, it implements a method for analyzing public source web page information based on LLM as described above.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0033] 1. The present invention optimizes the traditional web page information crawling process. The present invention crawls web page information based on LLM, which has certain generality and strong flexibility compared with the traditional information crawling process; compared with the existing intelligent AI crawling engine, it integrates the achievements of traditional algorithm web page crawling, has advantages in processing long text content, and improves the accuracy and applicability of web page information processing.
[0034] 2. The present invention improves the method for analyzing public source web page information. By making full use of the semantic understanding ability and instruction following ability of LLM, and integrating common xpath expression extraction technology, statistical algorithm extraction technology, and retrieval augmented generation (RAG) technology based on LLM in the field of web page information crawling, a web page information preprocessing, question and answer query, and application process is designed to achieve effective mining and analysis of public source web page information.
[0035] 3. The present invention explores a new model for open source intelligence analysis applications. Aiming at the application requirements of open source intelligence analysis for multi-web page combination, combining the search engine based on LLM and the web page information extraction ability, according to the hot events and hot figure information provided by users, relevant information is searched, crawled, and queried to generate structured information data for each web page, and following the structure template of common intelligence analysis reports, a summary report of hot event analysis is formed, and factual reference information such as web page links involved in the report is provided. Description of the Drawings
[0036] Figure 1 is a flowchart of an open-source web information analysis method based on LLM;
[0037] Figure 2 is a typical application flowchart of open-source intelligence analysis of the present invention. Detailed implementation manners
[0038] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0039] The features and performance of the present invention will be further described in detail below in conjunction with embodiments.
[0040] Embodiment 1
[0041] In the field of open-source intelligence analysis, open-source web information analysis plays an irreplaceable role. Web information analysis is not only an important way to obtain valuable intelligence, but also a key means to construct a comprehensive situation awareness map, predict trends and patterns, and verify and supplement other types of intelligence. With the continuous development of information technology and the increasing popularity of the Internet, the importance of web information analysis in the field of open-source intelligence analysis will become more prominent. However, there is too much redundant information in real open-source web pages, which inevitably has a great impact on the analysis of the main information of the web pages. The display basis of open-source web information is HTML text, and various browsers render the page according to the HTML text to obtain the web information seen daily. Therefore, HTML text is often used as the source for web information extraction. A large language model (LLM) is an artificial intelligence model trained with a large amount of data, which has certain semantic understanding and logical thinking abilities. It can generate corresponding responses according to the input instructions of users. By using LLM, it is possible to read and understand HTML text, identify content and then extract it, which is beneficial to generating structured text data and storing it in a vector library for answering questions raised by users.
[0042] Based on this, this embodiment proposes a method for analyzing open-source web page information based on LLM, which makes full use of the semantic understanding ability and instruction-following ability of LLM, integrates the common xpath expression extraction technology, statistical algorithm extraction technology and LLM-based retrieval enhancement generation (RAG) technology in the field of web page information crawling, realizes effective open-source web page information mining and analysis, and designs a corresponding open source intelligence analysis application process framework, which can intelligently query and generate briefings and analysis reports on hot events, and provide factual reference information involved in the report, and show users the analysis results and insights of related events, topics or individuals, helping users to better understand the situation, thereby formulating more scientific and reasonable decision-making plans, with good engineering application benefits.
[0043] In this embodiment, please refer to Figure 1 , a method for analyzing public source web page information based on LLM, specifically comprising the following steps:
[0044] Step S1: pre-processing the HTML text of the web page to extract the web page information contained in the HTML text;
[0045] Step S2: perform web page information query, retrieval and crawling according to user questions, and use LLM semantic understanding capabilities to retrieve similar reference information from web page information;
[0046] Step S3: Based on the retrieved similarity reference information, use the LLM logical thinking ability to reason and analyze the user's questions, extract the reference information involved in the questions and provide feedback; that is, based on the LLM model capability, the retrieved and processed web page information is used as an auxiliary decision-making reference in the form of questions and answers.
[0047] In this embodiment, specifically, the preprocessing method includes:
[0048] a. Use traditional web page extraction algorithms to parse and extract the subject content of web pages;
[0049] b. First, clean the HTML text of the web page and delete the content that is irrelevant to the main text, and then segment and store it in the database;
[0050] In this embodiment, it should be noted that the input requirement of this method is the HTML text of the public source web page, and it is assumed that the web page information is already contained in the obtained HTML text. In the HTML text preprocessing stage, there are two ways to process HTML text: one is to use the traditional web page extraction algorithm to parse and extract the subject content of the web page; the other is to first clean it and delete the content irrelevant to the main text in the HTML text to obtain a relatively pure web page HTML text, and then segment and store it; through the above two preprocessing methods, the information contained in the HTML text can be better extracted.
[0051] In this embodiment, specifically, the traditional web page extraction algorithms include: the Jina Reader web page parsing algorithm and the Readability algorithm;
[0052] That is, the open-source Jina Reader web page parsing method is selected as the implementation approach for traditional web page information extraction to convert web page content into MarkDown format text for the convenience of LLM use. Among them, for text web page information, the parsing algorithm used is the Readability algorithm. The processing flow of this algorithm is as follows: First, the HTML text is parsed into a DOM tree, and then each node of the tree is traversed. Different nodes are scored according to a pre-set scoring criterion (the scoring criterion is formulated based on the prior knowledge of the application scenario), and then the high-scoring content of the web page is extracted as the body tag content according to the score.
[0053] In this embodiment, it should be noted that the traditional web page information extraction method is prone to errors when processing web pages with a high information density (such as Wikipedia) because encyclopedia web pages have different tags and the content features of each tag (more links, similar number of texts, etc.) are relatively similar. In practical applications, the processed HTML text is often complex and has various redundant information. Therefore, it is necessary to clean the original HTML text. Referring to the common web page tag design in the industry and the classification definition of tags in traditional algorithms, the following cleaning is performed on the HTML text:
[0054] Remove the format definitions and JS scripts in the HTML text: Delete the tag <script><style><font>内容,删去<svg><button>等模块,删去注释;
[0055] 去掉与网页整体内容无关的标签模块:删去标签<footer><header><iframe><link>内容;
[0056] 规范HTML文本中多余的换行、空格等特殊符号:匹配连续的多个标签,将其转换为一到两个换行符;将连续的多个空白字符替换为单个空格;匹配一个或多个标签及其后的空白字符,并将其移除;去除字符串开头和结尾的空白字符(空格、制表符、换行符等)。
[0057] 在本实施例中,具体的,切分入库处理中采用的向量化模型是sentense-transformer / all-mpnet-base-v2模型,使用的向量库为Milvus;
[0058] 即清洗完成后,根据模型输入参数的长度限制,对HTML文本进行切分入库,若在使用传统网页提取算法时将大部分内容取出,那么在使用切分后的HTML文本进行回复的时候就是作为一个补充,因此不选择太长的长度限制,本实施例根据业务先验经验选择1000作为每一段切分的长度,为了增加准确性,重叠部分设置为200,然后进行向量化入库待查,此处使用的向量化模型是sentense-transformer / all-mpnet-base-v2模型,使用的向量库为Milvus。
[0059] 在本实施例中,需要说明的是,步骤S2和步骤S3的优势在于:
[0060] 通过集成Jina Reader的网页信息提取方法,基于传统网页信息提取算法获取网页信息的初步结果,包括标题和内容。由于传统算法的局限性,对于某些网页提取结果可能会出现遗漏差错,仍需要使用向量相似度查询能力对HTML文本切片进行相似度查询。此外,传统算法只能完成网页信息的提取功能,并不能实现信息的提炼(如摘要生成、实体总结、重要性评级等)与信息的转化(如内容翻译、名词统一等)功能,这些功能实现都在利用LLM进行结构化数据生成的过程中通过提示词指令(prompt)的方式进行。如果用户要求提取网页的主要内容,将会选择传统算法的提取结果直接输出。
[0061] 在本实施例中,具体的,使用检索增强生成技术对HTML文本进行切分入库处理:
[0062] 采用语义相似度查询的功能进行数据召回,召回处理过程中选择一定数量的相似文本,将其置入提示词作为LLM生成结构化数据的参考,如果用户问题与内容正文相关,也会把传统算法提取的正文信息置入提示词作为另外的参考资料。LLM在参考这些资料后,会针对用户问题提取出结构化数据,特别地,如果没有找到相关信息,默认LLM回复没有相关信息。该流程的关键是在LLM生成内容时提供充足的外部参考资料,利用LLM的语义理解与逻辑思维能力进行内容生成,有效降低LLM出现错误、遗漏与编造等情况。
[0063] 进一步地,请参阅图2,通常的公开源网页信息分析大多聚焦于单个网页的信息提取结果,但是在开源情报分析场景中,多网页联合信息提取问答成为重点需求,因此本发明针对开源情报分析应用,结合搜索引擎与网页信息提取,可以对热点事件、热点人物进行挖掘分析,生成分析总结报告,大致流程如下:
[0064] 根据用户提供的热点事件、热点人物信息,在搜索引擎中查询相关信息,并使用本发明的基础网页信息提取能力,生成各个网页的结构化信息数据,再结合这些信息数据,根据常见开源情报分析报告的结构模版,生成热点事件或人物的分析总结报告,并提供报告所参考的网页链接等事实性参考信息。
[0065] 此外,在一些实施例中,还提出了一种计算装置,包括:
[0066] 至少一个处理器;以及与所述至少一个处理器通信连接的存储器;其中,所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够执行如上述的一种基于LLM的公开源网页信息分析方法;计算装置的示例包括PC机、平板电脑、智能手机或PDA等。
[0067] 此外,在一些实施例中,还提出了一种计算机终端存储介质,存储有计算机终端可执行指令,所述计算机终端可执行指令用于执行如上述的一种基于LLM的公开源网页信息分析方法;计算机存储介质的示例包括磁性存储介质(例如,软盘、硬盘等)、光学记录介质(例如,CD-ROM、DVD等)或存储器,如存储卡、ROM或RAM等。计算机存储介质也可以分布在网络连接的计算机系统上,例如是应用程序的商店。
[0068] 此外,在一些实施例中,还提出了一种计算机程序产品,所述计算机程序被处理器执行时实现上述的一种基于LLM的公开源网页信息分析方法。
[0069] 实施例二
[0070] 下面基于QWEN1.5-32B大语言模型进行网页信息分析。
[0071] 1.选择网页目标
[0072] 为了测试系统的多样化能力,选取了多个不同类型、不同结构的网页进行测试,包含下表网站:
[0073] 表1测试网站列表
[0074]
[0075]
[0076] 2.测试问题设计
[0077] 1)LLM存在的必要性
[0078] 根据网页设计xpath提取出网页的主要内容,可以人为定义出更加细致的提取流程,但是LLM以其语义理解、逻辑思考的能力相较于传统的网页信息提取有着重要的优势,比如提取网页中的指定内容并进行翻译、提取出网页中出现的实体与其相关介绍等特定高级需求。因此,针对该场景设计了如下问题:
[0079] 该网页的标题与发布时间是什么,使用json形式回复!
[0080] 列出该网页中可能与中国相关的内容,使用json形式回复!
[0081] 该网站的主要内容是什么?
[0082] 假如我是一名学生,请你使用json形式分条输出该网页中应该关注的内容。
[0083] 2)传统网页提取算法的必要性
[0084] 目前业界拥有较多AI智能网页爬取框架,如ScrapeGraphAI、LangChainRetriever等,这些框架大部分都直接载入网页HTML内容,提取其中的文字,有些会进行切分入库,然后根据用户需求进行召回,这些框架在处理一些小型输出样例的时候效果较优,例如提取网页中的某些关键信息、提取网页中出现的图片链接等要求,但是对于提取网页正文片段、生成正文框架等要求LLM输出较多内容的情况下,这些框架采用的单方式网页信息提取流程就有所缺陷。因此为了体现本发明传统算法提取内容的必要性,设计如下问题:
[0085] 详细描述这个网页主要讲述了什么内容?
[0086] 列出网页的标题、作者、发表时间、摘要、来源、链接等,如果没有就置空,使用json格式输出!
[0087] 假如我是一名情报分析人员,请你分析页面中值得我关注的内容,分条列出。
[0088] 3)测试结果展示
[0089] 在表1所列的10个网站上对设计的7个问题进行了测试,下表是对结果的汇总:
[0090] 表2测试结果表
[0091]
[0092]
[0093] 上表中左栏是针对系统架构设计的7大问题,右栏是各系统对10个网页信息提取问答的效果总结,其中A表示传统网页信息提取功能的结果评分;B表示只利用清洗后的HTML文本切分入库查询的结果评分;C表示本系统的结果评分。评分标准按照结果的完整性进行评分,其中0表示完全没有回答成功,1~3逐级表示回答的完整性提高。这些评分由业务专家随机评估,一定程度上体现本发明方法的能力:
[0094] 观察前4个问题,其中第一与第三两个问题与网页内容相关性较强,因此使用传统算法提取的效果较优,但是在维基百科这样的百科性质网页处理上,由于结构复杂,正文不突出,因此提取效果较差;
[0095] 在第二与第四两个问题的处理上,由于需要对网页内容进行语义上的理解,因此LLM的参与必不可少,但是在多数网页处理方面,由于语义搜索召回的随机性,网页结构不能有效体现,因此本发明参考了传统算法提取,要优于只使用RAG的问答效果;
[0096] 观察后三个问题,对于网页文本的综合理解能力要求更多,因此本发明在多个问题的测试中都显现出优势。
[0097] 综合以上测试结果表明,在处理较为多样复杂的网页信息时,本发明能够较好地对网页信息进行提取,并且有效完成用户指定的任务,具有较好的适应性与可扩展性。
[0098] 4)开源情报应用效果展示
[0099] 在开源情报分析领域进行了应用场景的适配,可以生成相应的事件简报也可以生成相应的深度研判报告。
[0100] 以上所述实施例仅表达了本申请的具体实施方式,其描述较为具体和详细,但并不能因此而理解为对本申请保护范围的限制。应当指出的是,对于本领域的普通技术人员来说,在不脱离本申请技术方案构思的前提下,还可以做出若干变形和改进,这些都属于本申请的保护范围。
[0101] 提供本背景技术部分是为了大体上呈现本发明的上下文,当前所署名的发明人的工作、在本背景技术部分中所描述的程度上的工作以及本部分描述在申请时尚不构成现有技术的方面,既非明示地也非暗示地被承认是本发明的现有技术。< / script>
Claims
1. A method for analyzing public source web page information based on LLM, characterized in that, Including: Step S1: Preprocess the web page HTML text and extract the web page information contained in the HTML text; Step S2: Query, retrieve, and crawl web page information according to the user's question, and use the semantic understanding ability of the LLM to retrieve similar reference information from the web page information; Step S3: Based on the retrieved similar reference information, use the logical thinking ability of the LLM to reason and analyze the user's question, extract the reference information involved in the question, and give a feedback reply.
2. The method for analyzing public source web page information based on LLM according to claim 1, wherein The preprocessing method includes: a. Use traditional web page extraction algorithms to parse and extract the theme content of the web page; b. First, perform cleaning to delete the content irrelevant to the main body in the web page HTML text, and then perform segmentation and storage.
3. The method for analyzing open-source web page information based on LLM according to claim 2, wherein Traditional web page extraction algorithms include: Jina Reader web page parsing algorithm and Readability algorithm.
4. The method for analyzing public source web page information based on LLM according to claim 2, wherein, The cleaning process is as follows: Remove the format definitions and JS scripts in the HTML text; Remove the label modules irrelevant to the overall content of the web page; Standardize the redundant special symbols in the HTML text.
5. The method for analyzing public source web page information based on LLM according to claim 4, wherein, The vectorization model used in the segmentation and storage process is the sentense-transformer / all-mpnet-base-v2 model, and the vector library used is Milvus.
6. The method for analyzing public source web page information based on LLM according to claim 5, characterized in that Use retrieval-augmented generation technology to perform segmentation and storage processing on the HTML text.
7. The method for analyzing public source web page information based on LLM according to claim 6, wherein Using retrieval-augmented generation technology to perform segmentation and storage processing on the HTML text, including: Adopt the function of semantic similarity query to recall data. During the recall process, select a certain number of similar texts and place them in the prompt as a reference for the LLM to generate structured data. If the user's question is related to the content of the main body, place the main body information extracted by the traditional web page extraction algorithm in the prompt as another reference material; After referring to these materials, the LLM extracts structured data for the user's question.
8. A computing device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method for analyzing open-source web page information based on LLM as described in any one of claims 1-7.
9. A computer terminal storage medium stores computer terminal executable instructions, characterized in that, The executable instructions of the computer terminal are used to execute a method for analyzing open-source web page information based on LLM as described in any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program is executed by the processor, it implements a method for analyzing open-source web page information based on LLM as described in any one of claims 1-7.