System and method for intelligently acquiring, extracting and analyzing computing power center information based on large model

The information intelligent acquisition system based on large models solves the problem of difficulty in extracting information from multi-source heterogeneous computing centers, realizes efficient and accurate structured data acquisition, adapts to rapid updates and dynamic expansion, and provides a solid data foundation.

CN122019656APending Publication Date: 2026-05-12SHANGHAI GUOCHUANG YUSUAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI GUOCHUANG YUSUAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2026-01-20
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and accurately extract structured data from multi-source, heterogeneous, and unstructured computing center information, and are costly, failing to meet the demands of rapid updates and dynamic expansion.

Method used

The system employs a large-scale model-based intelligent information acquisition system, which includes a task scheduling module, a multi-source information acquisition module, a content preprocessing module, an AI intelligent analysis module, a rule verification and enhancement module, a multi-round targeted search completion module, and an intelligent merging and deduplication module. Through dynamic search, semantic understanding, domain rule verification, and multi-dimensional feature fusion, it achieves efficient and accurate information acquisition and structured processing.

Benefits of technology

It enables efficient, accurate, and structured collection of information from computing centers, improves data coverage and availability, can automatically process thousands of information sources, adapts to rapid updates and dynamic expansion, and provides a solid data foundation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019656A_ABST
    Figure CN122019656A_ABST
Patent Text Reader

Abstract

The invention discloses a computing power center information intelligent acquisition and extraction analysis system and method based on a large model. Comprising a task scheduling module, a multi-source information acquisition module, a content preprocessing module, an AI intelligent analysis module, a rule verification and enhancement module, a multi-round directional search completion module, an intelligent merging and duplicate removal module and a result processing and output module. The method comprises the steps that a task is automatically started, and multi-source information is retrieved in parallel and preprocessed; key information is extracted through a large language model, and verification and standardization are carried out through domain rules; evaluating information integrity and dynamically initiating multiple rounds of directional retrieval to complement missing data; intelligent merging and duplicate removal are carried out based on text, geography and computing power multi-dimensional feature fusion; and finally, outputting the structured data and visually displaying the structured data. Through AI and rule fusion processing, closed-loop optimization and multi-dimensional intelligent de-duplication, efficient, accurate and full-automatic acquisition and analysis of computing power center information are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing and data mining technology, specifically to a system and method for intelligent acquisition, extraction and analysis of information from computing centers based on large models. Background Technology

[0002] With the rapid development of the digital economy and artificial intelligence technologies, computing centers, as key infrastructure providing centralized computing, storage, and network services, have become crucial sources of information for industrial planning, resource allocation, and market analysis, including their construction scale, regional distribution, and service capabilities. However, information related to computing centers is typically scattered across various sources such as government reports, corporate websites, industry media, and academic papers. This presents challenges due to heterogeneous data sources, unstructured formats, highly specialized content, and frequent updates, resulting in efficient, accurate, and comprehensive information collection and analysis.

[0003] Currently, the collection and processing of information from computing centers mainly relies on the following technical solutions, all of which have significant limitations in practical applications:

[0004] (1) Manual sorting and semi-automated collection method; This method relies on manual retrieval, reading and input, or with the help of simple scripts to assist in the capture. Its disadvantages are mainly reflected in: 1) Low efficiency: It cannot meet the needs of large amount of information and rapid updates in computing centers. Manual processing is time-consuming and labor-intensive, and it is difficult to achieve real-time collection of large-scale information; 2) Poor consistency: Different people have different understandings and input standards for the same information, resulting in chaotic data formats and uneven quality; 3) Weak scalability: It cannot adapt to the dynamic expansion of multiple sources and types of information sources, and the system maintenance cost is high.

[0005] (2) Traditional web crawling and information extraction techniques: This technique automatically retrieves information from specific websites through preset crawling rules and parsing rules based on regular expressions or XPath. However, its limitations are very prominent: First, the workload of writing and maintaining the rules is huge, and it is highly dependent on the structure of the target website. Once the website is redesigned, the rules become invalid, resulting in poor adaptability. Second, traditional methods are difficult to effectively handle unstructured and semi-structured text content, and lack the ability to understand the professional terms, ambiguous expressions, and contextual semantics commonly found in the description of computing centers, leading to low extraction accuracy. Finally, this type of technology usually lacks intelligent cross-source comparison and deduplication mechanisms, which easily generate a large amount of redundant and contradictory data, resulting in poor data availability.

[0006] (3) Direct call to general large language models; Although large language models have strong natural language understanding and generation capabilities, their direct application in information extraction tasks in computing power centers still has obvious limitations: 1) Lack of domain knowledge: General large models have not been specifically trained for the computing power industry, and have insufficient understanding of professional knowledge such as computing power unit conversion, industry terminology, and policy norms; 2) High processing cost: Directly calling commercial large model APIs to process massive amounts of text data is extremely costly and is not suitable for large-scale industrial applications; 3) Limited long text processing capability: The relevant reports or page content of computing power centers are long, exceeding the single processing length limit of the model, resulting in incomplete information extraction.

[0007] (4) Deep research model; This type of model focuses on deep reasoning and multi-step retrieval of complex problems. It is suitable for writing analysis reports or answering comprehensive questions. However, it has the following mismatch in the information collection scenario of computing center: 1) Incompatible task positioning: Information collection in computing center is a "breadth" task, which emphasizes the efficient collection of a large amount of structured atomic information from multiple sources, rather than deep analysis and comprehensive reasoning; 2) Low processing efficiency: Deep research models usually perform multi-round, long-link retrieval and reasoning. The response speed is slow and cannot meet the information collection needs of high throughput. It is not economical and practical.

[0008] Therefore, it can be seen that existing technologies mostly focus on "deep" tasks (such as sentiment analysis, text summarization, and complex question answering) or rely on hard-coded rules to adapt to limited scenarios, failing to form an end-to-end automated solution for "broad" information collection tasks. Therefore, there is an urgent need for a computing center intelligent information collection system that integrates intelligent retrieval, semantic understanding, rule verification, and multi-source fusion to fill this technological gap. Summary of the Invention

[0009] To address the shortcomings of existing technologies, this invention provides a system and method for intelligent acquisition, extraction, and analysis of computing center information based on a large model. This system overcomes the deficiencies of existing technologies, achieves efficient, accurate, and structured acquisition of computing center information, significantly improves data coverage, accuracy, and usability, and provides an effective automated tool for computing resource planning and management.

[0010] To achieve the above objectives, the present invention provides the following technical solution:

[0011] A computing center information intelligent acquisition, extraction, and analysis system based on a large model, comprising:

[0012] The task scheduling module is used to automatically start the computing center information collection task according to the preset time strategy or event triggering conditions, and to monitor and log the execution status of the task in real time.

[0013] The multi-source information acquisition module is connected to the task scheduling module and is used to receive task instructions, query multiple Internet information sources and databases in parallel based on dynamically generated search terms, and obtain initial web page content or data interface return content containing information related to the computing center.

[0014] The content preprocessing module, connected to the multi-source information acquisition module, is used to clean the format of the acquired initial content and apply predefined rules based on keyword matching for fast filtering to remove content that is obviously unrelated to the theme of the computing center and output valid content.

[0015] The AI ​​intelligent analysis module is connected to the content preprocessing module. It is used to receive the preprocessed valid content, construct special prompt words containing system instructions, user instructions and the text to be analyzed, call the large language model for semantic understanding and parsing, and extract predefined structured key information fields from the valid content.

[0016] The rule verification and enhancement module is connected to the AI ​​intelligent analysis module and is used to load a professional knowledge rule base in the field of computing power, and to verify and standardize the key information fields extracted by the AI ​​intelligent analysis module, including at least the unified conversion of computing power units.

[0017] The multi-round targeted search completion module is connected to the output end of the rule verification and enhancement module and the input end of the multi-source information acquisition module, respectively. It is used to evaluate the completeness of the verified information and, when the key information field is found to be missing, dynamically generate supplementary search terms based on the current information and feed them back to the multi-source information acquisition module to start a new round of targeted retrieval and information extraction process, forming a closed-loop feedback.

[0018] The intelligent merging and deduplication module, connected to the rule verification and enhancement module and the multi-round targeted search completion module, is used to compare, merge and deduplicate computing center information records from different retrieval rounds and different information sources. The comparison is based on a multi-dimensional feature strategy that integrates text similarity, geographic information consistency and computing power information matching degree to determine whether multiple records point to the same computing center entity, and to perform information fusion and redundancy elimination on records that point to the same entity.

[0019] The result processing and output module is connected to the intelligent merging and deduplication module. It is used to summarize and format the high-quality structured computing center information data after final merging and deduplication, and to visualize and / or provide structured data file export function through a graphical user interface.

[0020] Preferably, the preset dedicated prompts in the AI ​​intelligent analysis module include at least:

[0021] The system instruction section defines the task role of the large language model as an information extraction expert for the computing center.

[0022] The user instruction section is used to explicitly list the key information fields of the computing center to be extracted from the text and specify the output format;

[0023] The processing rules section provides at least one domain constraint for computing center name normalization, geographic information hierarchy determination, computing data association verification, and data integrity checks.

[0024] Preferably, the predefined unified conversion rule for computing power units in the rule verification and enhancement module specifically performs the following calculations:

[0025] The computational power value of the original text description is converted into a value in PFLOPS using the following formula:

[0026] Standard value (PFLOPS) = Original number × Order of magnitude factor × Unit factor × 10 -15 ;

[0027] The magnitude coefficient is determined based on Chinese magnitude terms, and the unit coefficient is determined based on computing power unit terms.

[0028] The results are then validated for reasonableness, and outliers that exceed the preset reasonable range are filtered out.

[0029] Preferably, in the intelligent merging and deduplication module:

[0030] The text similarity calculation specifically uses the longest common subsequence algorithm to calculate the similarity score between the names of computing centers;

[0031] The geographic information consistency determination requires that the names of provincial-level administrative divisions in the information entries must be completely consistent, and the similarity of the names of municipal-level or more granular administrative divisions must not be less than a first preset threshold.

[0032] The matching degree determination of computing power information requires that, after the computing power information has been uniformly converted into the unit PFLOPS, the numerical difference is within a second preset threshold range;

[0033] The multi-dimensional feature strategy assigns different decision weights to text similarity, geographic information consistency, and computing power information matching degree, with text similarity having the highest weight.

[0034] Preferably, in the multi-round targeted search completion module, the termination condition of the closed-loop feedback is: the fill rate of all required fields reaches a preset completion threshold, or the number of rounds of execution reaches a preset maximum round threshold.

[0035] This invention also discloses a method for intelligent collection, extraction, and analysis of computing center information based on a large model, comprising the following steps:

[0036] S1: Automatically start the information collection task of the computing center through a preset triggering mechanism;

[0037] S2: Dynamically construct initial search terms based on the feature word library of the computing power center, perform multi-source parallel retrieval, and obtain raw content data;

[0038] S3: Perform content preprocessing on the raw content data obtained in step S2, including noise reduction, format standardization, and initial rule screening;

[0039] S4: Using a large language model and based on dedicated prompt words, extract key information fields of the computing center from the preprocessed content in step S3 in a structured manner;

[0040] S5: Apply predefined computing power domain rules to verify and standardize the extracted key information fields. The verification and standardization shall include at least the unified conversion of computing power units.

[0041] S6: Determine the completeness of the information processed in step S5. If key information fields are found to be missing, generate supplementary search terms dynamically based on the missing fields, and return to step S2 to perform a new round of targeted information acquisition and processing, forming a closed-loop optimization process of "retrieval-analysis-re-retrieval".

[0042] S7: For multiple pieces of information obtained through multiple rounds of targeted information acquisition and processing, perform intelligent comparison, merging and deduplication operations based on multi-dimensional feature fusion. The multi-dimensional features include at least text similarity, geographic information consistency and computing power information matching degree.

[0043] S8: Outputs the processed computing center information as a structured data file and supports visualization and human interaction.

[0044] Preferably, the construction of the dedicated prompt word in step S4 includes:

[0045] Provide system instructions for defining large language models as information extraction experts for computing centers;

[0046] Provide user commands that explicitly list the fields to be extracted and the output format;

[0047] Provides constraints that include specific processing rules for computing power domains.

[0048] Preferably, the intelligent merging and deduplication operation of multi-dimensional feature fusion in step S7 specifically includes:

[0049] Calculate the text similarity of computing center names among different information entries;

[0050] Determine the consistency of geographical information among different information items, where provincial information must be completely consistent;

[0051] After unifying the units, compare the computing power values ​​of different information items to see if they are within an acceptable error range;

[0052] Based on the combined results of the above dimensions and a decision made according to the preset weight allocation, it is determined whether the information item describes the same computing center entity.

[0053] Preferably, the termination condition of the closed-loop optimization process in step S6 is: the fill rate of the preset set of required fields meets the requirements, or the number of iterations reaches the upper limit.

[0054] This invention provides a system and method for intelligent information collection, extraction, and analysis from computing centers based on a large-scale model. It offers the following advantages: The task scheduling module automates and enables large-scale completion of the entire process from information discovery to structured extraction; a multi-source information acquisition module performs parallel retrieval based on dynamic keywords; and a multi-round targeted search completion module forms a closed loop for proactive gap filling. This allows the system to run automatically, scanning and processing thousands of dispersed information sources in a single task. The multi-round completion mechanism ensures continuous tracking and acquisition of key information, simultaneously improving the breadth and depth of information collection. It enables timely responses to the establishment, expansion, or information updates of computing centers, providing a solid data foundation for dynamic market analysis.

[0055] By employing a collaborative processing flow of an "AI intelligent analysis module and a rule verification and enhancement module," this approach solves the problems of high processing costs and inaccurate domain knowledge grasp when relying solely on large models, as well as the lack of flexibility and difficulty in understanding natural language descriptions when relying solely on rule systems. First, initial rule screening removes noise at low cost. Second, the powerful semantic understanding capabilities of large language models are utilized to flexibly and accurately extract preliminary structured information from complex text. Finally, the domain rule verification and enhancement module specifically addresses domain knowledge issues with high determinism that large models are not adept at. This ensures that the final result possesses both the flexibility of semantic understanding and the rigor of conforming to domain rules.

[0056] The multi-round targeted search completion module adopts an intelligent judgment algorithm based on the fusion of multi-dimensional features of text, geography, and computing power. It compares and makes decisions on information from different rounds and sources. The system can accurately identify and merge multiple records describing the same entity, thereby effectively solving the problem of redundancy and conflict of multi-source information. Finally, it generates unified, complete, and high-quality structured data assets, which greatly improves the usability and value of the data.

[0057] By constructing a multi-round targeted search and completion closed loop of "evaluation-generation-feedback", the system can automatically identify missing information items and intelligently initiate more targeted secondary searches based on existing results, thereby realizing information self-completion and continuous optimization, significantly enhancing the completeness and intelligence of a single collection task, and enabling the system to have the ability to dynamically adapt and self-improve. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the prior art will be briefly introduced below.

[0059] Figure 1 This is a flowchart illustrating the overall system architecture and data processing of the present invention.

[0060] Figure 2 This is a schematic diagram illustrating the multi-dimensional feature fusion determination of the intelligent merging and deduplication module in this invention;

[0061] Figure 3 This is a flowchart of the closed-loop optimization process for multi-round targeted search completion in this invention;

[0062] Figure 4 This is a flowchart of the content processing workflow for the fusion of AI and rules in this invention;

[0063] Figure 5 This is a flowchart illustrating the overall steps of the intelligent information collection, extraction, and analysis method for computing power centers according to the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0065] Example 1, as Figures 1 to 4 As shown, this invention provides an intelligent data acquisition and extraction analysis system for computing centers based on a large model, consisting of eight core modules connected sequentially and supplemented by feedback loops. The data processing flow begins with the task scheduling module, then flows sequentially through the multi-source information acquisition module, content preprocessing module, AI intelligent analysis module, and rule verification and enhancement module. The multi-round targeted search and completion module sends feedback to the multi-source information acquisition module based on the verified information completeness, forming a closed-loop optimization. Finally, the data is integrated by the intelligent merging and deduplication module and delivered by the result processing and output module. Each module exchanges data based on a clear interface protocol, ensuring fully automated execution of the entire process. The specific implementation methods of each module are as follows:

[0066] The task scheduling module is used to automatically initiate computing center information collection tasks according to preset time strategies or event triggering conditions, and to monitor and log the execution status of the tasks in real time; its core functions include:

[0067] (1) Task Triggering: Two modes are supported. One is a preset time strategy based on cron expressions (such as midnight every day); the other is an event triggering condition in response to external API calls or system events. After triggering, the module instantiates a collection task, assigns a globally unique ID, and initializes the task context.

[0068] (2) Status monitoring and logging: The module has a built-in state machine to track the entire lifecycle of a task from "ready", "running" to "completed" or "failed". Each downstream module reports execution progress, success / failure information and performance indicators through a unified log interface. These are collected and persistently stored in the database by this module, providing real-time monitoring dashboards and historical task auditing functions.

[0069] Taking the execution of a "weekly inspection of the East China computing center" task as an example, the task scheduling module is implemented based on a distributed scheduled task framework such as Quartz or Celery. The system administrator presets a scheduling policy of "execution at 1:00 AM every Monday" through the management backend. When the time is triggered, the module creates a task instance, generates a globally unique task ID (such as Task_20231030_001), records the start timestamp, and sends a start command to the multi-source information acquisition module. At the same time, the module starts a monitoring thread to subscribe in real time to the status logs reported by each subsequent module.

[0070] The multi-source information acquisition module is connected to the task scheduling module and is used to receive task instructions, perform parallel queries on multiple Internet information sources and databases based on dynamically generated search terms, and obtain initial webpage content or data interface return content containing information related to the computing center; specifically including:

[0071] (1) Dynamic search term generation: The module maintains a configurable "seed term library" which includes computing center types (such as "intelligent computing center" and "supercomputing center") and technical terms (such as "GPU cluster" and "AI computing power"). When the task starts, it dynamically combines and generates batch search query terms based on task parameters (such as target region).

[0072] (2) Multi-source parallel query: The module integrates multiple data source adapters, including: 1) public API clients of mainstream search engines; 2) customized crawlers for specific industry websites or government public data platforms; 3) partner data interfaces. These adapters are called concurrently using an asynchronous I / O framework (such as asyncio), submitting search terms and obtaining returned results (including web page links, summaries, and even structured API data).

[0073] (3) Content Acquisition: For web page links, the module uses a headless browser (such as Puppeteer) or an efficient HTTP client (such as aiohttp) to download the complete HTML content in order to handle pages rendered by JavaScript.

[0074] For example, upon receiving the task instruction, the multi-source information acquisition module first dynamically generates a batch of initial search query terms from its built-in "Administrative Division Database" (including "Shanghai," "Jiangsu," and "Zhejiang") and "Core Feature Terminology Database" (including "Intelligent Computing Center," "Artificial Intelligence Computing Center," and "Supercomputing Center"), such as ["Shanghai Intelligent Computing Center," "Jiangsu Supercomputing Center," and "Zhejiang Artificial Intelligence Computing Center"]. Subsequently, the multi-source information acquisition module launches multiple asynchronous crawler threads or calls clients of services such as Google Custom Search API and Bing Web Search API to submit these search terms in parallel. Each thread is responsible for one information source, thereby avoiding the limitations of a single source and improving the collection speed. The module collects the returned search results, including webpage titles, summaries, and links, and passes them down as "initial content." The first round of collection in this task yielded approximately 150 initially relevant webpage links.

[0075] The content preprocessing module is connected to the multi-source information acquisition module. It performs format cleaning on the acquired initial content and applies predefined rules based on keyword matching for rapid filtering to remove content clearly irrelevant to the computing center's theme, outputting valid content. Specifically, this includes:

[0076] (1) Format cleaning: Using HTML parsing libraries such as BeautifulSoup or lxml, irrelevant tags such as scripts, style sheets, navigation bars, footers, and advertisements in the original HTML are removed to accurately extract the main text of the article.

[0077] (2) Keyword-based rule filtering: Establish a "negative list of domain keywords". Quickly scan the extracted text. If any core keyword in the list (e.g., "computing power", "FLOPS", "GPU", "data center", "computing center") does not appear in the text, the page is determined to be obviously irrelevant to the topic and is discarded. This step can filter out a large number of irrelevant pages in advance, greatly reducing the processing load of the downstream AI module.

[0078] (3) Output standardization: Convert the retained text to UTF-8 encoded strings and perform necessary text cleanup (such as removing redundant whitespace characters). Output a uniformly formatted "Valid Content" object with the source information above.

[0079] For example, the content preprocessing module processes the 150 links collected in the first round, first downloading the original HTML of the webpages and removing all... <script>、<style>、<nav>等非主体内容标签,提取核心正文文本。再对正文文本应用快速关键词匹配。预设的关键词列表包括{"算力”,"FLOPS”,"GPU”,"服务器”,"数据中心”}。如果一个页面的正文中不包含列表中的任何一个词(例如,可能是一个仅包含联系方式的页面),则被判定为"无关内容”直接丢弃。将保留下来的有效正文文本统一转换为UTF-8编码的纯文本,并截断至前8000个字符以适应模型输入限制。经过本步骤,150个链接中有40个被过滤掉,剩余110份"有效内容”被送入下一模块。

[0080] AI智能分析模块连接至所述内容预处理模块,用于接收经预处理后的有效内容,构造包含系统指令、用户指令及待分析文本的专用提示词,调用大语言模型进行语义理解与解析,从有效内容中提取预定义的结构化关键信息字段;

[0081] 其中,专用提示词由系统指令部分、用户指令部分和处理规则部分三部分构成,具体地:

[0082] 系统指令部分用于将大语言模型的任务角色定义为算力中心信息抽取专家;例如:"你是一个专注于基础设施领域的专业信息提取助手,任务是精确识别和提取算力中心的属性信息。”

[0083] 用户指令部分用于明确列举待从文本中提取的算力中心关键信息字段清单,并指定输出格式;例如:"请从下文提取:1.name(名称),2.location(省、市), 3.computing_power(数值,单位PFLOPS) ...请以JSON格式输出。”

[0084] 处理规则部分用于提供算力中心名称规范化、地理信息层级判定、算力数据关联验证及数据完整性检查中的至少一种领域约束。如:"注意:‘长三角AI算力池’应规范为‘长三角人工智能计算中心’;算力描述‘千P级’需按1000 PFLOPS处理。”

[0085] 调用大语言模型进行语义理解与解析具体为:模块封装对大语言模型服务(如云端API或本地部署模型)的调用。将构造好的提示词与预处理后的"有效内容”文本组合,发送至模型。接收模型的自然语言或结构化(如JSON)响应,并解析为系统内部的标准数据结构。设有重试、降级和错误处理机制以保证可靠性。

[0086] 例如,AI智能分析模块接收110份有效内容后,对于每一份有效内容,自动构造包含上述三部分指令的专用提示词,再调用大语言模型逐份解析文本,接收返回的JSON数据。本轮处理从110份内容中,初步提取出85条包含部分关键信息的结构化记录。例如,其中一条记录可能为:{"name”: "临港智算中心”, "location”: {"province”: "上海”,"city”: "上海”}, "computing_power”: null, …}。

[0087] 规则校验与增强模块连接至所述AI智能分析模块,用于加载算力领域的专业知识规则库,对所述AI智能分析模块提取出的关键信息字段进行校验与标准化处理,其中至少包括算力单位的统一换算;其中:

[0088] 规则库包含:1)标准化映射表:如运营状态映射{"投产”:"已运营”,"在建”:"在建”,"规划”:"规划中”}。2)逻辑校验规则:如判断location字段是否包含有效的省、市二级信息。3)算力单位统一换算规则:实现精确换算。

[0089] 具体算力单位换算流程包括:1)文本解析:使用正则表达式从原始字符串中分离数字、中文数量级词(万、亿、万亿等)、单位词(次 / 秒、FLOPS等)。2)系数匹配与计算:根据预定义的词典,将中文数量级词和单位词转换为对应的数学系数。3)公式执行:应用公式:标准值(PFLOPS)=原始数字×数量级系数×单位系数×10-15进行计算(1FLOPS=1次 / 秒,1PFLOPS=1015FLOPS)。4)合理性校验:将计算结果与预设的合理值域(如[0.01, 10000]PFLOPS)比对,对显著异常值进行标记或置空。

[0090] 最后利用规则库,对某些字段进行推断或补全,例如根据算力中心名称中的地域词补充缺失的location信息。

[0091] 例如:规则校验与增强模块加载预定义的YAML或数据库格式的规则库。对85条记录逐条处理;首先对于AI提取的算力描述,如"峰值算力达200P”,直接标准化为"computing_power”: 200。对于"5百万亿次浮点运算 / 秒”,则应用权利要求3所述公式:标准值=5×(1014)×1×10-15=0.5PFLOPS, 并将字段更新为"computing_power”:0.5。再检查computing_power值。规则库预设合理范围为[0.01, 10000]PFLOPS。若某个记录的值为0.0001(疑似单位识别错误)或500000(明显异常),则将该字段标记为"异常”,并在记录中添加校验标志。再将status字段中的"已投入使用”、"运行中”统一映射为"已运营”;将service_type中的"对外提供服务”映射为"公有云”。经过本步骤,85条记录被进一步清洗和增强,其中20条记录的computing_power字段为null或被标记为"异常”。

[0092] 多轮定向搜索补全模块分别连接至所述规则校验与增强模块的输出端和多源信息获取模块的输入端,用于评估经校验后信息的完整性,并在识别到关键信息字段缺失时,基于当前信息动态生成补充搜索词,并反馈至所述多源信息获取模块以启动新一轮的定向检索与信息提取流程,形成闭环反馈;具体包括:

[0093] (1)完整性评估:多轮定向搜索补全模块维护一个"必填字段”配置列表(例如,[‘name’,‘location.province’,‘computing_power’])。它对规则校验与增强模块输出的每条记录进行检查,若任一必填字段为空或为低置信度标识,则判定该记录不完整。

[0094] (2)补充搜索词动态生成:针对不完整记录,提取其已有高置信度字段(如name和location.city),与特定补全模板组合。例如,对缺失computing_power的记录,生成类似"{name} 技术白皮书 算力指标”或"{name} {location.city} 数据中心 规格”的补充查询词。此过程更具针对性。

[0095] (3)闭环反馈控制:将生成的补充搜索词集合,作为新一轮采集任务的输入,提交给多源信息获取模块。系统从而进入"检索→分析→校验→评估→再检索”的闭环。闭环的终止条件由两个"或”关系的条件控制:1)所有记录的必填字段总体填充率达到预设阈值(如95%);2)闭环执行轮次达到预设最大值(如3轮)。满足任一条件即退出闭环,进入后续流程。

[0096] 例如,多轮定向搜索补全模块根据预定义的"必填字段”配置列表评估记录的完整性。发现20条记录computing_power缺失或异常,判定需要补全。针对这些记录,结合其已有name和location字段,构造更具针对性的搜索词。例如,针对"临港智算中心”,生成搜索词:"临港智算中心 算力规模 PFLOPS”或"临港智算中心 官网 技术参数”。再将这20个补充搜索词列表,作为新一轮的任务指令,发送回多源信息获取模块。系统启动第二轮定向检索。第二轮检索可能只获取到15个相关的新网页,经S3-S5步骤处理后,成功补全了其中12条记录的算力信息。此时,系统判断必填字段填充率已达到设定阈值(例如,85条记录中仅剩8条缺失,填充率约91%,高于预设的90%阈值),遂终止闭环流程,将结果移交至下游数据融合模块。剩余8条未完全补全记录被标记为"部分可信”,其已填充字段进入主数据库,缺失字段留空并附置信度说明,供后续人工核查或专项补充采集使用。

[0097] 智能合并与去重模块连接至所述规则校验与增强模块和所述多轮定向搜索补全模块,用于对来自不同检索轮次、不同信息源的算力中心信息记录进行比对、合并与去重;所述比对基于融合文本相似度、地理信息一致性和算力信息匹配度的多维度特征策略,以判定多条记录是否指向同一算力中心实体,并对指向同一实体的记录进行信息融合与冗余消除;

[0098] 其中,文本相似度主要针对算力中心名称。采用最长公共子序列算法计算两个名称字符串的相似度得分,该算法对字符顺序调换具有一定容错性。

[0099] 地理信息一致性判定首先要求省级名称完全一致。对于市、区级名称,采用编辑距离等算法计算相似度,当相似度不低于第一预设阈值(例如0.7)时判定为一致。

[0100] 算力信息匹配度判定是在确保单位统一为PFLOPS的前提下,计算两条记录算力数值的相对误差。若误差小于第二预设阈值(例如15%),则判定为匹配。

[0101] 通过多维度特征融合策略为上述三个特征分配权重。通常,文本相似度的权重最高(如0.5-0.7),因为名称是最直接的标识;地理一致性权重次之(如0.3-0.4);算力匹配度权重相对较低(如0.1-0.2),因算力数据可能随时间变化或存在不同统计口径。加权计算得到总相似度得分。

[0102] 最后设定一个总分的判定阈值(如0.75)。当两条记录的总相似度得分超过该阈值时,判定它们指向同一实体。合并操作以一条记录为主,将另一条记录中的独有信息或更精确信息补充进来,实现信息融合,然后删除重复项。

[0103] 例如,此时系统共持有来自首轮和第二轮的、经过校验的算力中心记录约97条(85+12)。智能合并与去重模块采用上述多维度特征融合策略进行比对,具体包括:1)名称相似度:采用最长公共子序列算法计算。例如,"长三角AI计算中心”与"长三角人工智能计算中心”的LCS相似度得分经计算为0.82。2)地理一致性:检查location.province是否严格相等。city字段采用编辑距离计算相似度,若相似度>0.7(第一预设阈值),则判定为一致。3)算力匹配度:在统一为PFLOPS后,若两条记录的算力值相对误差小于15%(第二预设阈值),则判定为匹配。

[0104] 随后为三个特征分配权重:名称相似度(0.6)、地理一致性(0.3)、算力匹配度(0.1)。设定总分阈值0.75。例如,A、B两条记录名称得分0.82×0.6=0.49,地理完全一致得0.3,算力匹配得0.1,总分0.89>0.75,判定为同一实体。对于判定为同一实体的记录组,选取信息最完整(非空字段最多)的一条作为主记录,将其他记录中的独有信息(如另一个来源提供的精确投运日期)补充到主记录中,然后删除冗余副本。本步骤将97条记录合并去重为80条唯一、高质量的算力中心结构化数据。

[0105] 结果处理与输出模块连接至所述智能合并与去重模块,用于将最终合并去重后的、高质量的结构化算力中心信息数据进行汇总、格式化,并通过图形用户界面进行可视化展示和 / 或提供结构化数据文件导出功能。具体地,包括:

[0106] (1)数据汇总与格式化:接收来自智能合并与去重模块的最终数据集,将其转换为通用的结构化数据格式,如CSV、JSON或直接写入关系型数据库(如MySQL)或Elasticsearch中。

[0107] (2)可视化展示:提供图形用户界面(GUI),集成数据可视化组件。例如,利用地图库(如Leaflet)展示算力中心的地理分布;利用图表库(如ECharts)展示算力规模分布、运营状态比例等。

[0108] (3)交互与导出功能:在GUI中提供数据筛选、搜索、排序功能。同时,提供"一键导出”功能,允许用户将当前视图或查询结果以结构化数据文件(如Excel)的形式下载。

[0109] 例如,结果处理与输出模块将上述80条高质量算力中心数据转换为标准的JSONArray格式,并同时生成一份CSV文件,文件名为华东算力中心清单_20231030.csv。通过集成ECharts库的前端界面,自动生成一张标注了80个算力中心位置的中国地图,点击标记点可显示详情。同时,生成算力规模分布柱状图、运营状态饼图等。在图形用户界面提供该数据集的下载链接(CSV / JSON格式),并提供基于省份、算力区间的筛选查询功能。所有操作日志和最终数据报告被归档存储。结果输出完成后,任务调度模块收到最终状态上报,将本次任务标记为"成功完成”,并记录总耗时。系统恢复待命状态,等待下一次任务触发。

[0110] 本发明通过任务调度模块能够自动化、大规模地完成从信息发现到结构化提取的全流程,利用多源信息获取模块进行基于动态关键词的并行检索,并引入多轮定向搜索补全模块形成主动查漏补缺的闭环。从而有效克服了人工方式效率低下、传统爬虫覆盖面窄且被动、以及深度研究模型不适合广度收集的缺陷。使得系统能够自动运行,单次任务即可自动化扫描和处理数千个分散的信息源。多轮补全机制确保了对关键信息的持续追踪与获取,使得信息收集的广度(覆盖源)和深度(信息完整度)同步提升,能够及时响应算力中心新建、扩容或信息更新,为动态市场分析提供了坚实的数据基础。

[0111] 另外,本发明通过采用"AI智能分析模块和规则校验与增强模块”的协同处理流程。解决了单纯依赖大模型处理成本高、对领域知识(如复杂单位换算)掌握不精准,以及单纯依靠规则系统灵活性差、难以理解自然语言描述的问题。首先,通过规则初筛低成本排除噪音;其次,利用大语言模型的强大语义理解能力,灵活、准确地从复杂文本中抽取出初步结构化信息;最后,通过领域规则校验与增强模块,专门解决大模型不擅长的、确定性强的领域知识问题。从而确保了最终结果既具备语义理解的灵活性,又符合领域规则的严谨性。

[0112] 多轮定向搜索补全模块采用基于文本、地理、算力多维度特征融合的智能判定算法,并对来自不同轮次与来源的信息进行比对与决策,系统能够精准识别并合并描述同一实体的多条记录,从而有效解决了多源信息冗余与冲突的难题,最终生成统一、完整、高质量的结构化数据资产,极大提升了数据的可用性与价值。

[0113] 通过构建"评估-生成-反馈”的多轮定向搜索补全闭环,系统能够自动识别信息缺失项,并基于已有结果智能发起更具针对性的二次检索,从而实现了信息的自我补全与持续优化,显著增强了单次采集任务的完整性与智能性,使系统具备了动态适应与自我改进的能力。

[0114] 实施例二,如图5所示,本发明还公开了一种基于大模型的算力中心信息智能采集与提取分析方法,包括以下步骤:

[0115] S1:通过预设的触发机制,自动启动算力中心信息采集任务;具体实施中,所述触发机制可配置为:

[0116] (1)定时触发:基于操作系统级任务调度器(如Linux Cron)或应用级调度框架(如APScheduler),设置周期性执行计划(例如,每日凌晨2点),时间到达即自动创建并启动一个新的采集任务实例。

[0117] (2)事件触发:监听外部事件,如接收到包含特定指令的HTTP API调用、消息队列中的命令消息,或监控到源数据网站有更新公告时,自动触发任务执行。

[0118] 任务启动后,系统生成唯一的任务标识符,并初始化全局上下文,为后续步骤提供执行依据和状态跟踪基础。

[0119] S2:根据算力中心特征词库动态构建初始搜索词,执行多源并行检索,获取原始内容数据;其具体实施包括:

[0120] S201:动态构建初始搜索词:系统加载预先构建的"算力中心特征词库”,该词库包含行业通用术语(如"智算中心”、"超算中心”、"GPU集群”)和扩展关键词(如"人工智能算力”、"高性能计算”)。根据任务参数或通用策略,将这些特征词与目标地域词(如省份、城市名称)或通用查询后缀(如"建设进展”、"投运”)进行智能组合,生成一批初始搜索查询词列表。

[0121] S203:执行多源并行检索:系统同时向多个信息源发起查询请求。这通过以下方式实现:

[0122] 为每个主流搜索引擎(如百度、Bing)配置API客户端,并发提交搜索词。

[0123] 对特定的行业门户、政府公开信息平台,启动异步网络爬虫,模拟浏览器访问进行内容抓取。

[0124] 采用线程池或异步I / O框架管理所有并发请求,以最大化网络I / O效率,缩短整体采集时间。

[0125] 收集所有返回的原始内容,包括网页HTML源码、API返回的JSON / XML数据等,并将其统一定义为"原始内容数据”传递至下一步骤。

[0126] S3:对步骤S2获取的原始内容数据进行内容预处理,包括去噪、格式标准化与规则初筛;具体实施为:

[0127] S301:去噪:使用HTML解析库(如BeautifulSoup)去除网页内容中的无关标签(如<script>, <style>, <iframe>, 广告代码块等),提取纯净的正文文本。对于API返回的数据,解析其结构,提取文本字段。

[0128] S302:格式标准化:将所有文本内容统一转换为UTF-8编码。对过长的文本进行智能截断(例如,保留前10000个字符),以确保适配下游大语言模型的输入长度限制。统一换行符和空格格式。

[0129] S303:规则初筛:应用基于关键词的快速过滤规则。系统维护一个"核心关键词列表”,包含"算力”、"FLOPS”、"计算中心”、"服务器”等必现词。对每一份提取出的正文文本进行扫描,若文本中未包含任何列表中的关键词,则判定该内容与算力中心主题明显无关,予以直接丢弃,不进入后续深度分析流程。

[0130] S4:利用大语言模型,基于专用提示词,从步骤S3预处理后的内容中结构化提取算力中心的关键信息字段;具体实施为:

[0131] S401:构造专用提示词:

[0132] 系统指令:设定模型角色,例如:"你是一个专业的算力基础设施信息提取引擎,你的任务是从文本中精确抽取出指定的结构化信息。”

[0133] 用户指令:清晰定义待提取的字段(如"中心名称”、"所在地(省 / 市)”、"理论算力(PFLOPS)”、"运营状态”等)和要求输出的结构化格式(如JSON)。

[0134] 约束条件:嵌入领域规则,例如:"若算力表述为‘千P级’,则转换为1000;‘所在地’需区分省和市两级。”

[0135] S402:调用大语言模型:将构造好的提示词与步骤S3预处理后的每份有效文本内容相结合,形成完整的输入,调用大语言模型服务(如云端API或本地部署的大模型)。模型基于其对提示的理解和对文本的语义分析,输出符合指定格式的结构化信息。

[0136] S5:应用预定义的算力领域规则,对提取的关键信息字段进行校验与标准化,所述校验与标准化至少包括算力单位的统一换算;具体实施为:

[0137] S501:加载规则库:从配置文件或数据库中加载预定义的算力领域规则库。

[0138] S502:执行校验与标准化:

[0139] 逻辑校验:检查字段间的逻辑关系,例如,所在地是否包含有效的行政区划。

[0140] 格式规范化:统一表述,如将"已启用”、"运行中”统一规范为"已投产”。

[0141] 算力单位统一换算(核心):对提取的算力字符串,应用预定义的换算规则进行解析和计算。例如,识别"5百万亿次 / 秒”,解析数字"5”,数量级"百万亿”(系数1014),单位"次 / 秒”(系数1),根据公式标准值(PFLOPS)=5×1014×1×10-15=0.5进行计算,并将结果标准化为"0.5 PFLOPS”。同时对结果进行合理性校验。

[0142] S6:判断经步骤S5处理后的信息的完整性,若识别出关键信息字段缺失,则基于缺失字段动态生成补充搜索词,并返回步骤S2执行新一轮的定向信息获取与处理,形成"检索-分析-再检索”的闭环优化流程;且其流程有明确的终止条件。具体实施为:

[0143] S601:完整性判断:系统根据预定义的"必填字段集合”(如名称、所在地、算力),检查经过步骤S5处理后的每条记录。若记录中存在必填字段为空或无效,则标记为"不完整”。

[0144] S602:动态生成补充搜索词:针对所有不完整记录,系统提取其已确定的可靠信息(如准确的名称),与待补全的字段类型结合,生成更具针对性的补充搜索词。例如,对于已知名称"星河AI计算中心”但缺失算力的记录,生成搜索词"星河AI计算中心 算力规模PFLOPS”。

[0145] S603:发起新一轮处理:将生成的补充搜索词集合,作为新的输入,返回至步骤S2,开启新一轮的定向检索、预处理、AI提取和规则校验。此过程构成一个闭环。

[0146] S604:终止判断:该闭环流程并非无限循环。其终止依据权利要求9所述的二选一条件:

[0147] 条件一(填充率满足):当所有记录的必填字段总体填充率达到预设阈值(如95%)。

[0148] 条件二(轮次达限):当闭环执行的轮次达到预设的最大上限(如3轮)。只要满足任一条件,即跳出闭环,继续执行步骤S7。

[0149] S7:对通过多轮定向信息获取与处理所获得的多条信息,进行多维度特征融合的智能比对、合并与去重操作,所述多维度特征至少包括文本相似度、地理信息一致性和算力信息匹配度;具体实施为:

[0150] S701:特征计算与判定:

[0151] 文本相似度计算:采用最长公共子序列算法,计算两条记录中算力中心名称的相似度得分。

[0152] 地理信息一致性判定:首先,检查"省”级字段是否完全一致。其次,对"市 / 区”级字段,采用字符串相似度算法计算,若得分超过预设阈值(如0.7),则判定为一致。

[0153] 算力信息匹配度判定:在确保所有算力值已统一为PFLOPS的前提下,计算两个数值的相对误差,若误差小于可接受阈值(如10%),则判定为匹配。

[0154] S702:融合决策与操作:为上述三个维度分配权重(例如,文本:0.6,地理:0.3,算力:0.1),计算加权总分。设定合并阈值(如0.75)。当两条记录的总分超过阈值时,判定为描述同一实体,继而进行信息合并(以信息更完整或置信度更高的记录为主,补充另一条的缺失或更优信息)并删除冗余记录。

[0155] S8:将处理后的算力中心信息输出为结构化数据文件,并支持可视化展示与人工交互。具体实施为:

[0156] S801:结构化数据文件生成:将步骤S7处理后的最终数据集,序列化为标准的结构化数据文件,如CSV、JSON或Excel格式,便于后续导入数据库或直接分析。

[0157] S802:可视化展示:通过集成图表库(如ECharts、D3.js)的Web前端,将数据以地图(标注算力中心位置)、柱状图(算力分布)、饼图(状态比例)等形式进行直观展示。

[0158] S803:人工交互支持:在可视化界面提供数据筛选、查询、排序功能,并设有"导出”按钮,允许用户将当前视图下的数据或全部数据下载为文件。同时,可提供简单的编辑接口,供用户对自动处理的结果进行最终审核与微调。

[0159] 本实施例通过上述八个步骤的协同运作,构建了一套从任务触发到数据输出的完整智能化流程。该流程以自动化任务调度为起点,通过动态搜索词构建与多源并行检索实现信息的广泛获取,经预处理与规则初筛降低噪音干扰,再借助大语言模型的语义理解能力提取结构化信息,结合领域规则校验确保数据的准确性与标准化。多轮定向搜索补全机制形成闭环,有效提升信息完整性,而多维度特征融合的比对合并操作则解决了多源数据的冗余与冲突问题,最终以结构化文件与可视化交互方式呈现成果。整个系统实现了算力中心信息采集从人工依赖向全自动、智能化的转变,不仅大幅提升了信息处理效率与覆盖范围,更通过"采集-分析-优化”的闭环设计,确保了数据的高质量与动态适应性,为算力基础设施的规划、管理与市场分析提供了全面且可靠的数据支持。

[0160] 以上实施例仅用以说明本发明的技术方案,而非对其限制;尽管参照前述实施例对本发明进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本发明各实施例技术方案的精神和范围。< / script>

Claims

1. A system for intelligent acquisition, extraction, and analysis of computing center information based on a large model, characterized in that: include: The task scheduling module is used to automatically start the computing center information collection task according to the preset time strategy or event triggering conditions, and to monitor and log the execution status of the task in real time. The multi-source information acquisition module is connected to the task scheduling module and is used to receive task instructions, query multiple Internet information sources and databases in parallel based on dynamically generated search terms, and obtain initial web page content or data interface return content containing information related to the computing center. The content preprocessing module, connected to the multi-source information acquisition module, is used to clean the format of the acquired initial content and apply predefined rules based on keyword matching for fast filtering to remove content that is obviously unrelated to the theme of the computing center and output valid content. The AI ​​intelligent analysis module is connected to the content preprocessing module. It is used to receive the preprocessed valid content, construct special prompt words containing system instructions, user instructions and the text to be analyzed, call the large language model for semantic understanding and parsing, and extract predefined structured key information fields from the valid content. The rule verification and enhancement module is connected to the AI ​​intelligent analysis module and is used to load a professional knowledge rule base in the field of computing power, and to verify and standardize the key information fields extracted by the AI ​​intelligent analysis module, including at least the unified conversion of computing power units. The multi-round targeted search completion module is connected to the output end of the rule verification and enhancement module and the input end of the multi-source information acquisition module, respectively. It is used to evaluate the completeness of the verified information and, when the key information field is found to be missing, dynamically generate supplementary search terms based on the current information and feed them back to the multi-source information acquisition module to start a new round of targeted retrieval and information extraction process, forming a closed-loop feedback. The intelligent merging and deduplication module, connected to the rule verification and enhancement module and the multi-round targeted search completion module, is used to compare, merge and deduplicate computing center information records from different retrieval rounds and different information sources. The comparison is based on a multi-dimensional feature strategy that integrates text similarity, geographic information consistency and computing power information matching degree to determine whether multiple records point to the same computing center entity, and to perform information fusion and redundancy elimination on records that point to the same entity. The result processing and output module is connected to the intelligent merging and deduplication module. It is used to summarize and format the high-quality structured computing center information data after final merging and deduplication, and to visualize and / or provide structured data file export function through a graphical user interface.

2. The intelligent acquisition, extraction, and analysis system for computing center information based on a large model according to claim 1, characterized in that: The AI ​​intelligent analysis module includes at least the following preset special prompt words: The system instruction section defines the task role of the large language model as an information extraction expert for the computing center. The user instruction section is used to explicitly list the key information fields of the computing center to be extracted from the text and specify the output format; The processing rules section provides at least one domain constraint for computing center name normalization, geographic information hierarchy determination, computing data association verification, and data integrity checks.

3. The intelligent acquisition, extraction, and analysis system for computing center information based on a large model according to claim 1, characterized in that: The predefined unified conversion rule for computing power units in the rule verification and enhancement module specifically performs the following calculations: The computational power value of the original text description is converted into a value in PFLOPS using the following formula: Standard value (PFLOPS) = Original number × Order of magnitude factor × Unit factor × 10 -15 ; The magnitude coefficient is determined based on Chinese magnitude terms, and the unit coefficient is determined based on computing power unit terms. The results are then validated for reasonableness, and outliers exceeding the preset reasonable range are filtered out.

4. The intelligent acquisition, extraction, and analysis system for computing center information based on a large model according to claim 1, characterized in that: In the intelligent merging and deduplication module: The text similarity calculation specifically uses the longest common subsequence algorithm to calculate the similarity score between the names of computing centers; The geographic information consistency determination requires that the names of provincial-level administrative divisions in the information entries must be completely consistent, and the similarity of the names of municipal-level or more granular administrative divisions must not be less than a first preset threshold. The matching degree determination of computing power information requires that, after the computing power information has been uniformly converted into the unit PFLOPS, the numerical difference is within a second preset threshold range; The multi-dimensional feature strategy assigns different decision weights to text similarity, geographic information consistency, and computing power information matching degree, with text similarity having the highest weight.

5. The intelligent acquisition, extraction, and analysis system for computing center information based on a large model according to claim 1, characterized in that: In the multi-round targeted search completion module, the termination condition for closed-loop feedback is: the fill rate of all required fields reaches a preset completion threshold, or the number of rounds of execution reaches a preset maximum round threshold.

6. A method for intelligent acquisition, extraction, and analysis of computing center information based on a large model, characterized in that: Includes the following steps: S1: Automatically start the information collection task of the computing center through a preset triggering mechanism; S2: Dynamically construct initial search terms based on the feature word library of the computing power center, perform multi-source parallel retrieval, and obtain raw content data; S3: Perform content preprocessing on the raw content data obtained in step S2, including noise reduction, format standardization, and initial rule screening; S4: Using a large language model and based on dedicated prompt words, extract key information fields of the computing center from the preprocessed content in step S3 in a structured manner; S5: Apply predefined computing power domain rules to verify and standardize the extracted key information fields. The verification and standardization shall include at least the unified conversion of computing power units. S6: Determine the completeness of the information processed in step S5. If key information fields are found to be missing, dynamically generate supplementary search terms based on the missing fields and return to step S2 to perform a new round of targeted information acquisition and processing. S7: For multiple pieces of information obtained through multiple rounds of targeted information acquisition and processing, perform intelligent comparison, merging and deduplication operations based on multi-dimensional feature fusion. The multi-dimensional features include at least text similarity, geographic information consistency and computing power information matching degree. S8: Outputs the processed computing center information as a structured data file and supports visualization and human interaction.

7. The intelligent acquisition, extraction, and analysis method for computing center information according to claim 6, characterized in that: The construction of the dedicated prompt words in step S4 includes: Provide system instructions for defining large language models as information extraction experts for computing centers; Provide user commands that explicitly list the fields to be extracted and the output format; Provides constraints that include specific processing rules for computing power domains.

8. The intelligent acquisition, extraction, and analysis method for computing center information according to claim 6, characterized in that: The intelligent merging and deduplication operation of multi-dimensional feature fusion in step S7 specifically includes: Calculate the text similarity of computing center names among different information entries; Determine the consistency of geographical information among different information items, where provincial information must be completely consistent; After unifying the units, compare whether the computing power values ​​of different information items are within an acceptable error range; Based on the combined results of the above dimensions and a decision made according to the preset weight allocation, it is determined whether the information item describes the same computing center entity.

9. The intelligent acquisition, extraction, and analysis method for computing center information according to claim 6, characterized in that: The termination condition of the closed-loop optimization process described in step S6 is: the fill rate of the preset set of required fields meets the requirements, or the number of iterations reaches the upper limit.