A corpus generation method and device for a smart grid planning field
By acquiring formal knowledge from data in the field of smart power grid planning, identifying and clustering knowledge chains, generating corpus templates, and inputting them into a generative large model, the problems of professionalism, scalability, and scenario adaptability in smart power grid planning corpus generation are solved, achieving efficient corpus generation and agent training effects.
Patent Information
- Application Number
- CN202511543233.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing technologies struggle to generate high-quality, large-scale, scenario-adaptable, and logically consistent intelligent power grid planning corpora, resulting in poor agent training performance and low efficiency in the automated generation of business documents.
By acquiring formalized knowledge from data in the field of smart grid planning, identifying knowledge chains, clustering and merging them into the main chain, generating corpus templates, and inputting them into a generative large model, we can ensure the logical consistency and scenario adaptability of the corpus with the field of smart grid planning.
It has achieved high-quality, large-scale and diverse corpus generation, improved the accuracy of agent training and the efficiency of automated business document generation, and ensured the professionalism and scenario adaptability of the corpus.
Smart Images

Figure CN121009965B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart grid planning technology, and in particular to a corpus generation method and apparatus for smart grid planning. Background Technology
[0002] With the deepening of the construction of new power systems, the power grid structure is exhibiting complex characteristics such as a high proportion of new energy access, coexistence of multiple business users, and cross-regional interconnection and collaboration. As a core link in ensuring the safe operation of the power grid and optimizing resource allocation, the decision-making mode of intelligent power grid planning is transforming from traditional experience-driven to data-model-knowledge collaborative driven. In this process, high-quality corpora in the field of intelligent power grid planning have become a key support. They are not only the core data foundation for training AI models for grid optimization and fault repair plan generation systems, but also an important carrier for realizing the automated generation of planning business documents and the inheritance and reuse of industry expert knowledge. Their quality and scale directly determine the accuracy and efficiency of intelligent power grid planning.
[0003] Currently, the acquisition and generation of corpus in the field of smart power grid planning mainly relies on two technical approaches. The first is manual collection and organization, which involves domain experts sorting through historical planning project reports, business operation manuals, and expert experience notes to extract key business expressions to form corpus. This method requires a large investment of manpower and is limited by the scope of expert knowledge and organization efficiency. The average time to generate a single corpus can be several hours, and the annual generation is usually only a few thousand corpus entries, which is difficult to meet the needs of tens of thousands to hundreds of thousands of corpus entries required for training intelligent agents. The second is the transfer and application of general natural language generation technologies, which directly uses general generative models such as GPT and BERT, or generates corpus after a small amount of power grid text has been lightweighted and fine-tuned. However, such models do not deeply integrate the professional logic of power grid planning and can only output surface business expressions, failing to reflect domain-specific knowledge.
[0004] However, existing technological approaches suffer from significant bottlenecks, making it difficult to adapt to the specialized and scenario-based requirements of smart grid planning. Firstly, the capacity for large-scale generation of specialized corpora is insufficient, manual compilation is inefficient, and general models, lacking domain knowledge support, often produce corpora with errors in equipment parameters and confused business logic. Secondly, the corpora exhibit poor scenario adaptability. Grid planning encompasses multiple sub-scenarios such as fault repair, business expansion applications, and grid optimization, with significant differences in core processes and constraints across different scenarios. Existing technologies lack scenario-based constraint mechanisms, easily leading to mismatches between distribution network planning corpora and main grid scenarios. Thirdly, quality control is weak, lacking comprehensive oversight of the entire corpus generation chain (knowledge input, logic construction, internal...). The quality evaluation and anomaly correction mechanisms for generated data are insufficient to pinpoint the root cause of data problems (such as whether parameter errors stem from knowledge input bias or model generation bias). Low-quality data directly impacts the training effect of intelligent models. Fourth, the reuse rate of domain knowledge is low. Formal knowledge such as equipment ledgers and business rule bases accumulated in power grid planning has not formed a systematic knowledge-data conversion path. Structured knowledge is difficult to convert into natural language data with business logic, and unstructured expert experience cannot be efficiently integrated into the generation process, resulting in the waste of existing knowledge resources and data updates lagging behind business knowledge iterations. Summary of the Invention
[0005] In view of this, this application provides a corpus generation method and apparatus for the field of smart grid planning, in order to solve the problems of difficulty in large-scale generation of professional corpus in the field of smart grid planning, poor scene adaptability, weak quality controllability, and imbalance between diversity and constraints.
[0006] Specifically, this application is implemented through the following technical solution:
[0007] The first aspect of this application provides a corpus generation method for the field of smart grid planning, the method comprising:
[0008] Acquire formal knowledge of data in the field of smart grid planning;
[0009] Identify the formal knowledge to obtain multiple knowledge chains, each knowledge chain including multiple knowledge nodes, and each node being an entity in the formal knowledge;
[0010] The knowledge chains are clustered by identifying the link types of the knowledge chains, and the clustered knowledge chain classes are themed around the smart grid planning scenario.
[0011] Merge the knowledge chains within the same category to obtain the unique main chain corresponding to each knowledge chain class;
[0012] Generate corpus templates based on the main chain;
[0013] The corpus template is input into a generative large model to predict multiple different corpora in the field of smart grid planning corresponding to the corpus template.
[0014] A second aspect of this application provides a corpus generation device for the field of smart grid planning, the device comprising an acquisition module, an identification module, a clustering module, a merging module, a generation module, and a prediction module;
[0015] The acquisition module is used to acquire formal knowledge of data in the field of smart grid planning;
[0016] The identification module is used to identify the formal knowledge and obtain multiple knowledge chains, each knowledge chain including multiple knowledge nodes, and each node is an entity in the formal knowledge;
[0017] The clustering module is used to identify the link type of the knowledge chain and cluster the multiple knowledge chains. The clustered knowledge chain classes are themed around the smart grid planning scenario.
[0018] The merging module is used to merge various knowledge chains within the same category to obtain a unique main chain corresponding to each knowledge chain class;
[0019] The generation module is used to generate corpus templates based on the main chain;
[0020] The prediction module is used to input the corpus template into a generative large model and predict multiple different corpora in the field of smart grid planning corresponding to the corpus template.
[0021] The corpus generation method and apparatus for the field of smart power grid planning provided in this application use templates as carriers and generate unified templates that conform to knowledge constraints using formal knowledge, ensuring the consistency between the generated corpus and the logical knowledge in the field of smart power grid planning. Simultaneously, it generates knowledge chains for different scenarios, ensuring multi-scenario coverage. Finally, through differentiated prediction of template gaps, it guarantees sample diversity. Through a full-process design involving acquiring formal knowledge, constructing knowledge chains, clustering chains, merging main chains, generating templates, and predicting corpora, it accurately solves the problems of large-scale corpus generation difficulties, poor scenario adaptability, and weak professional logic in the field of smart power grid planning, achieving the goal of accurately transforming domain knowledge into high-quality professional corpus. Firstly, starting with acquiring formal knowledge of data in the field of smart power grid planning, it achieves a deep binding between domain knowledge and corpus generation, ensuring the professionalism of the corpus from the source. Formal knowledge in the field of smart grid planning is the core support for the professional attributes of the corpus. By using the formal knowledge of data in the field of smart grid planning as the input basis for the entire corpus generation process, the problem of traditional general generation technology being divorced from domain knowledge and outputting superficial text is avoided. By directly reusing the structured knowledge accumulated in the field of power grid, the subsequently generated corpus naturally has the professional expression of power grid planning, without the need for additional manual supplementation of professional information. This not only improves the efficiency of knowledge reuse, but also eliminates the generation of non-professional corpus from the source.
[0022] Secondly, through identification and clustering steps, precise adaptation of the corpus to the intelligent power grid planning scenario was achieved. By clarifying that the knowledge chain consists of entities in formal knowledge, and that the clustered chain classes are themed around the intelligent power grid planning scenario, the corpus generation has a clear scenario orientation. Subsequent corpus generated based on scenario-based chain classes can accurately match the business needs of different scenarios, avoiding scenario mismatch issues. Furthermore, by merging similar chains to obtain a unique main chain and generating corpus templates based on the main chain, a unified business logic framework is constructed for the corpus, ensuring logical consistency. The main chain is the extraction of the core business logic of similar chains, while the corpus template is the natural language presentation of the main chain. This ensures that all corpus generated based on the template strictly follows the core process and logical relationship of the main chain. The uniqueness of the main chain also avoids logical fragmentation of the corpus under similar scenarios, ensuring that corpus generated in different batches maintains a consistent business expression framework and improving the corpus's versatility in agent training.
[0023] Finally, the corpus templates were input into the generative large-scale model's prediction corpus, enabling the large-scale and diversified generation of professional corpora. The corpus templates provide clear business constraint boundaries for the large-scale model, allowing it to output corpora with different expressions in batches within these constraints, meeting the scalable training requirements of the intelligent agent. Furthermore, template constraints prevent the large-scale model from generating redundant content detached from business needs, ensuring that while the corpus is diverse, it always aligns with the professional logic and scenario requirements of power grid planning. Ultimately, this achieves efficient production of large-scale, scenario-based, and high-quality corpora, providing core data support for intelligent agent training in power grid intelligent planning and automated generation of business documents. Attached Figure Description
[0024] Figure 1 A flowchart of an embodiment of the corpus generation method for smart grid planning provided in this application;
[0025] Figure 2 This is a schematic diagram of the second embodiment of the corpus generation device for the field of smart grid planning provided in this application. Detailed Implementation
[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0027] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0028] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0029] The following specific embodiments are given to illustrate the technical solution of this application in detail.
[0030] Example 1
[0031] Figure 1This is a flowchart of an embodiment of the corpus generation method for smart grid planning provided in this application. Please refer to... Figure 1 The method provided in this embodiment may include:
[0032] S101. Obtain formal knowledge of data in the field of smart grid planning.
[0033] It should be noted that formalized data knowledge refers to a set of knowledge formed by organizing and expressing scattered, unstructured business information in the field of smart grid planning in a standardized and structured manner. Formalized data knowledge can include business rule knowledge (i.e., the technical standards, policy requirements, and constraints that power grid planning must follow), entity parameter knowledge (i.e., the core entities (equipment, regions, scenarios, etc.) involved in power grid planning and their attribute parameters), logical relationship knowledge (i.e., the relationships between different entities, rules, and scenarios), and historical case knowledge (i.e., reusable structured experience from completed power grid planning projects).
[0034] Specifically, formalizing knowledge through data can be achieved by analyzing authoritative materials within the power grid field, including but not limited to internal operational manuals of power grid companies, industry-issued technical standard documents (such as national standards and industry specifications related to the power industry), historical power grid planning project proposals, and expert-recorded summaries of business experience. During the acquisition process, it is crucial to ensure that the collected knowledge aligns with the actual business needs of intelligent power grid planning, possesses professionalism, accuracy, and completeness, and avoids logical deviations or business compliance issues in subsequent data generation due to missing or incorrect knowledge.
[0035] Furthermore, it should be noted that the formal knowledge obtained above has a structured characteristic, that is, the knowledge content must be presented in a form that can be recognized and parsed by a computer. For example, the equipment parameters and their corresponding relationships can be sorted out through tables, the order of business process nodes can be clarified through flowcharts, and business constraints can be clarified through rule entries. In this way, an operable basic data format can be provided for extracting knowledge nodes and establishing node relationships in subsequent steps.
[0036] S102. Identify the formal knowledge and obtain multiple knowledge chains, each knowledge chain including multiple knowledge nodes, and each node is an entity in the formal knowledge.
[0037] It should be noted that, firstly, independent knowledge nodes are extracted from the acquired formal knowledge, and then the connection relationship between nodes is established based on the power grid business logic (such as process sequence and parameter dependence). Finally, the orderly related nodes are combined into multiple knowledge chains to ensure that each chain can reflect the logical context of a specific business link or scenario of the power grid intelligent planning.
[0038] Among these steps, identifying formal knowledge refers to using structured analysis methods to extract the basic elements (i.e., knowledge nodes) that constitute the knowledge chain from the formal knowledge obtained in S101, and to sort out the logical relationships between these elements. This process is not simply about filtering knowledge, but rather about combining machine recognition (such as domain-adapted entity recognition algorithms) with manual verification to clarify the unit attributes of knowledge (such as whether it belongs to equipment, processes, or rules) and the types of relationships (such as chronological order or parameter matching), providing a basis for connecting subsequent nodes.
[0039] A knowledge chain refers to a sequence of knowledge nodes connected in a specific logical order, reflecting the business logic of power grid planning. Each knowledge chain corresponds to a specific business segment in intelligent power grid planning. For example, a user submits a business expansion request → performs capacity calculation → determines the transformer model → designs an access scheme, or fault location → isolates the faulty line → dispatches emergency repair equipment → performs maintenance operations. Its core function is to connect scattered knowledge nodes into a logical chain with business significance, providing a structured framework for subsequent scenario-based clustering and corpus generation.
[0040] A knowledge node refers to the basic unit that constitutes a knowledge chain, and each node corresponds to an entity in formal knowledge. These entities encompass core business elements in the power grid planning field, including equipment entities (such as transformers and transmission lines), process entities (such as capacity calculation and fault isolation), scenario entities (such as business expansion applications and grid optimization), and rule entities (such as load matching rules and power supply reliability standards). Each knowledge node not only contains basic information about the entity itself (e.g., a transformer node includes attributes such as rated capacity and voltage level), but also carries logical interfaces for association with other nodes (e.g., a transformer node needs to form a parameter dependency association with a capacity calculation node).
[0041] It should be noted that the execution process of this step can be divided into three sub-steps: splitting knowledge nodes, establishing node connection relationships, and integrating them to form a knowledge chain and standardizing it. In specific implementation, the formalized knowledge is identified to obtain multiple knowledge chains, including:
[0042] (1) Analyze the entities and attributes in the formal knowledge, identify the entities as knowledge nodes, and use the attributes as the descriptive information of the nodes.
[0043] It should be noted that breaking down formal knowledge into the smallest associative knowledge nodes avoids the confusion in subsequent chain construction logic caused by mixed knowledge dimensions. Furthermore, by supplementing with attributes, knowledge nodes not only include entity names but also carry business characteristics, ensuring that the generated corpus accurately reflects the professional attributes of power grid planning. In addition, the parameters, constraints, and operational requirements contained in the attribute descriptions are key bases for subsequently determining the parameter dependencies, rule constraints, and other relationships between nodes, safeguarding the business logic of the knowledge chain.
[0044] Specifically, entity recognition models trained on corpora in the power grid domain (such as BERT fine-tuning models and CRF conditional random field models) can automatically extract three types of core entities from formal knowledge. Among them, equipment entities are identified by the model through the names of equipment such as transformers, transmission lines, and substations, while filtering out non-core words (such as conjunctions such as "of" and "and"); process entities are identified through business process links such as capacity calculation, fault isolation, and scheme review, and the model achieves accurate extraction by annotating the combination structure of process verbs + business nouns (such as calculation-capacity, isolation-fault); rule entities are identified through constraint items such as N-1 security verification and load matching rules, and the model locates rule-type entities through rule keywords (such as need, should, standard, constraint).
[0045] It should also be noted that a relation extraction + structured mapping approach can be used to extract the attributes corresponding to entities and associate them with the entities. For equipment entities, using an entity-attribute relation extraction model (such as RE-BERT), the rated capacity and voltage level are extracted as attributes from statements such as "110kV transformer rated capacity 50MVA", and 50MVA and 110kV are extracted as attribute values. For process entities, using a process-attribute mapping table (pre-setting standard attributes for power grid business processes, such as cycle, response time, and execution standards), the cycle is extracted as an attribute from the capacity calculation cycle of 3 working days, and 3 working days is extracted as an attribute value. For rule entities, using a rule-constraint parsing tool, the constraint condition is extracted as an attribute from "transformer capacity must be ≥ 1.2 times the load", and "≥ 1.2 times the load" is extracted as an attribute value.
[0046] Furthermore, for complex knowledge that is difficult for models to accurately identify (such as non-standard equipment parameters and special scenario process attributes), manual verification can be conducted by power grid planning experts. Specifically, this includes verifying the accuracy of entity identification, supplementing missing attributes, and standardizing attribute descriptions.
[0047] (2) Establish connections between nodes based on the entity relationships defined in the formal knowledge to form a preliminary knowledge chain containing entity association sequences.
[0048] It should be noted that in this step, firstly, it is necessary to identify the various entity relationships defined in the formal knowledge. Secondly, based on the identified entity relationship types, a method combining rule matching and domain modeling is used to extract the specific associations between nodes from the formal knowledge. Then, directed connections between nodes are constructed according to preset connection rules, and finally integrated into a preliminary knowledge chain reflecting a certain business segment.
[0049] In the field of smart grid planning, defined entity relationships are the explicit association rules established through technical specifications, business processes, and parameter constraints. These entity relationships serve as the links connecting nodes. Specifically, entity relationships in smart grid planning mainly include temporal relationships, parameter dependencies, and rule constraint relationships.
[0050] Among them, temporal relationships are entity relationships defined based on the sequence of business processes, and often exist between process entities and process entities or between process entities and parameter entities. For example, in formal knowledge, "business expansion applications must first accept the request, then conduct capacity calculation, and finally design the access solution," which clarifies the temporal relationship of "request acceptance (process entity) → capacity calculation (process entity) → access solution design (process entity)." Another example is "before conducting capacity calculation, it is necessary to obtain the user's load data for the past 3 years," which clarifies the temporal relationship of "user's load data for the past 3 years (parameter entity) → capacity calculation (process entity)" (parameter acquisition comes first, process execution comes later).
[0051] Parameter dependencies are entity relationships defined based on the matching requirements of equipment parameters and business data. They often exist between parameter entities and equipment entities, or between parameter entities and rule entities. For example, in formal knowledge, "the rated capacity of the transformer must be ≥ the maximum load of the area × 1.2 times" clarifies the dependency relationship of "maximum load of the area (parameter entity, attribute: value) → 110kV transformer (equipment entity, attribute: rated capacity)" (the equipment parameter needs to be determined by the value calculation of the parameter entity); another example is "the current carrying capacity of the line must match the total rated current of the connected equipment", which clarifies the dependency relationship of "total rated current of the equipment (parameter entity) → transmission line (equipment entity, attribute: current carrying capacity)".
[0052] Rule-based constraints are entity relationships defined by industry standards and safety specifications, and often exist between rule entities and process entities, or between rule entities and equipment entities. For example, in formal knowledge, "the grid design must pass N-1 safety verification" clarifies the constraint relationship of "N-1 safety verification rule (rule entity) → grid design (process entity)" (process execution must comply with the rule requirements); another example is "the substation site selection must avoid areas prone to geological disasters," which clarifies the constraint relationship of "geological disaster avoidance rule (rule entity) → substation (equipment entity, attribute: site location)" (equipment attributes must meet the rule restrictions).
[0053] It should be noted that for entity relationship identification, a rule-matching + domain model-assisted approach can be used to extract entity relationships from formal knowledge. For temporal relationships, the order of nodes is located by identifying temporal keywords such as "before / after," "before / after," and "must be acquired / executed first." For parameter dependency relationships, the association between parameter entities and dependent entities is located by identifying dependency keywords such as "must be greater than or equal to," "must be less than or equal to," "determine," and "match." For rule constraint relationships, the association between rule entities and constrained entities is located by identifying constraint keywords such as "must conform," "must pass," and "prohibit." For complex relationships (such as multiple parameter dependencies), a relationship extraction model trained in the power grid domain (such as a Transformer-based RE model) is used to ensure that no relationship is missed in the identification process.
[0054] Furthermore, corresponding node connection rules are formulated for different types of entity relationships. For temporal relationships, entities are connected in the order of first occurrence → later occurrence, such as user load data for the past 3 years → capacity calculation → 110kV transformer selection. For parameter dependency relationships, entities are connected in the direction of parameter entity → dependent entity, such as regional maximum load of 40MW → 110kV transformer (rated capacity of 50MVA). For rule constraint relationships, entities are connected in the direction of rule entity → constrained entity, such as N-1 safety verification rule → grid design. At the same time, connection rules are explicitly prohibited, such as process entities cannot be connected in reverse (e.g., capacity calculation → demand acceptance is an invalid connection) and rule entities cannot depend on parameter entities (e.g., regional maximum load → N-1 rule is an invalid connection), to avoid logical errors.
[0055] Finally, the node sequences connected according to the rules are integrated into preliminary knowledge chains, with each chain corresponding to a complete business logic segment. These preliminary knowledge chains include single-relationship chains and multi-relationship hybrid chains. A single-relationship chain contains only one type of entity relationship, such as demand acceptance → capacity calculation → access scheme design (purely time-series relationship). A multi-relationship hybrid chain contains multiple entity relationships, such as user load data for the past 3 years (parameter entity) → capacity calculation (process entity, time-series first) → N-1 security verification rules (rule entity, constraint) → 110kV transformer (equipment entity, parameter dependency) → access scheme design (process entity, time-series second). This type of chain is closer to the complex business realities of power grid planning.
[0056] In this way, the scattered knowledge nodes are reorganized according to the inherent logic of the smart grid planning business, and the structural relationships between different knowledge units are restored, which can provide a reliable structured foundation for the subsequent generation of corpora that conform to the actual business logic.
[0057] (3) The preliminary knowledge chain is deduplicated and standardized to obtain multiple standardized knowledge chains.
[0058] It should be noted that the initial knowledge chain is directly generated based on entity relationships. Due to differences in the original formalized knowledge representation and errors in the identification process, two types of problems may exist: first, duplicate or highly similar chains lead to knowledge redundancy, increasing the computational cost of subsequent processing; second, the same entity or relationship is expressed inconsistently in different chains, such as 110kV transformers versus 110kV transformers, or scheme reviews versus approved schemes, affecting the uniformity and comparability of the knowledge chain. This step addresses these problems through deduplication and standardization, making the knowledge chain a standardized knowledge unit with a unified structure, clear logic, and no redundancy, more accurately reflecting the business logic of smart grid planning.
[0059] Specifically, for completely duplicated chains, identical chains can be directly deleted by comparing the node sequence and connection relationship of the knowledge chain; for chains with slightly different node descriptions but consistent core logic and business connotation, they are judged as substantial duplicates and merged; when a chain is a complete subsequence of another longer chain and there is no new business logic, it is judged as a redundant sub-chain and removed.
[0060] Furthermore, after deduplication, standardization is carried out. Based on the standard terminology system in the field of smart grid planning, the names of knowledge nodes are unified; the way the connection relationship between nodes is expressed is unified, and the standard symbols or descriptions of relationships such as timing, dependency, and constraint are clarified; the overall format of the knowledge chain is unified, including the order of node arrangement and the way attribute information is presented.
[0061] S103. Identify the link types of the knowledge chains and cluster the multiple knowledge chains. The clustered knowledge chain classes are themed around the smart grid planning scenario.
[0062] It should be noted that smart grid planning encompasses multiple sub-scenarios, such as fault repair, business expansion applications, and grid optimization. While the knowledge chains of different scenarios vary, they exhibit similar node connection logic (i.e., link types). For example, the chains in the fault repair scenario are primarily based on temporal relationships, while the chains in the grid optimization scenario often involve parameter dependencies and rule constraints. This step leverages this characteristic to first identify the link type of each knowledge chain (extracting its connection logic features), then group chains with similar link types into one category. Ultimately, each category of chains corresponds to a smart grid planning scenario, achieving precise matching between knowledge chains and business scenarios. This avoids scenario mismatches in subsequent corpus generation (e.g., using grid optimization corpus to adapt to the fault repair scenario).
[0063] It should be noted that the link type of a knowledge chain refers to the specific category of the connection relationship between nodes in the knowledge chain. This is a core feature distinguishing knowledge chains from different scenarios. Based on the business logic of smart grid planning, it is mainly divided into three categories: sequential association type, parameter dependency type, and rule constraint type. In actual chains, multiple types may be mixed, requiring the identification of all included link types as the chain's feature labels. Furthermore, a knowledge chain class themed around a smart grid planning scenario refers to a set formed after clustering, where all chains serve the same power grid planning scenario.
[0064] Specifically, clustering the multiple knowledge chains by identifying the link types of the knowledge chains includes:
[0065] (1) Extract the link types between nodes in each knowledge chain.
[0066] It should be noted that extracting the link types between nodes in a knowledge chain is fundamental to clustering operations, aiming to clarify the logical characteristics of each knowledge chain's connections. Link type refers to the specific category of the relationships between different nodes in a knowledge chain, reflecting the interaction patterns between entities in smart grid planning.
[0067] Specifically, each standardized knowledge chain needs to be analyzed node by node to identify the relationships between adjacent nodes. In the field of smart grid planning, common link types are mainly divided into three categories based on business logic: first, temporal correlation, where there is a clear sequential execution order between nodes, reflecting the sequential relationship of steps in the fault handling process; second, parameter dependency, where the parameters or attributes of the later node are determined by the parameters of the earlier node, reflecting the matching relationship between equipment parameters and load data; and third, rule constraint, where the relationships between nodes are restricted by industry norms or safety standards, reflecting the requirement that the planning scheme must comply with specific rules. During the extraction process, it is necessary to combine the entity relationships defined in formal knowledge, and determine the link type between all nodes in each knowledge chain through keyword recognition (such as prior to, depending on, and must comply with), and business logic judgment, providing a feature basis for subsequent chain similarity calculations.
[0068] (2) Configure weight coefficients for different link types, and calculate the similarity between knowledge chains based on the link types and weight coefficients.
[0069] Specifically, the configuration of weighting coefficients needs to be combined with the business characteristics of smart grid planning and set according to the representativeness of the link type to the business scenario. For example, in fault repair scenarios, time-series related links are crucial to the integrity of the process, so the weighting coefficient can be set higher (e.g., 0.5); while in grid optimization scenarios, rule-constrained links affect the compliance of the planning scheme, so the weighting coefficient can be increased accordingly (e.g., 0.4). Parameter-dependent links play a significant role in scenarios involving equipment selection, such as business expansion and installation, and the weighting coefficient also needs to be reasonably configured according to its importance (e.g., 0.3). The sum of the weighting coefficients is usually 1 to ensure the standardization of similarity calculation.
[0070] When calculating the similarity between knowledge chains, it is necessary to compare the link types contained in the two chains. The weight coefficients of each link type in each chain are used as feature vectors, and the similarity is calculated using vector similarity algorithms (such as cosine similarity or weighted summation). For example, if chain A contains temporal association (weight 0.5) and parametric dependency (weight 0.3), and chain B also contains the same link types with identical weight configurations, then the similarity between the two chains is 0.5 + 0.3 = 0.8, indicating that they are highly similar. In this way, the logical similarity between knowledge chains can be transformed into quantifiable values, providing a clear basis for clustering.
[0071] (3) Based on the number of clusters preset for the business scenarios of smart grid planning, knowledge chains with similarity higher than the preset threshold are divided into the same category using chain similarity as the metric, thus obtaining knowledge chain classes with business scenarios as the theme, and labeling each knowledge chain class with the corresponding scenario label.
[0072] It should be noted that, firstly, based on the core nodes and business logic of the knowledge chain, the scenario affiliation of each chain is determined, and the core business nodes in each standardized knowledge chain are extracted. Simultaneously, combining frequently occurring business terms within the chain, a mapping rule from core nodes to scenario tags is established. Then, based on this, all knowledge chains are divided into preset scenario categories. Furthermore, the number of clusters is preset according to the actual business needs of smart grid planning. Smart grid planning typically covers typical scenarios such as fault repair, business expansion applications, grid optimization, and new energy access; therefore, the corresponding number of clusters (e.g., 4 clusters) can be preset based on the number of these scenarios.
[0073] Subsequently, using chain similarity as a metric, clustering algorithms (such as K-means) are employed to group knowledge chains. Specifically, knowledge chains with similarity exceeding a preset threshold (e.g., 0.7) are grouped into the same category. For instance, all knowledge chains primarily consisting of time-series connections with high similarity are grouped together, as these chains often correspond to fault repair scenarios; while knowledge chains primarily consisting of parameter-dependent and rule-constrained connections may be categorized as business expansion applications or network structure optimization.
[0074] Finally, each clustered knowledge chain class is labeled with its corresponding scenario tag. By analyzing the common characteristics of the knowledge chains in the category (such as the main components of the link type and the core business nodes involved), the corresponding smart grid planning scenario is determined. The labeling needs to form a two-level classification system of primary scenario categories + secondary connection type subcategories; the primary scenario categories correspond to the core business scenarios of smart grid planning (such as fault repair, business expansion and installation, grid optimization, and new energy access); the secondary connection type subcategories are divided based on the differences in link types within the same scenario category (for example, under the fault repair category, it is divided into pure time-series subcategories based on the proportion of time-series related links ≥90%, and into time-series-parameter hybrid subcategories based on a mixture of time-series related links and parameter-dependent links). After the labeling is completed, each knowledge chain class becomes a knowledge set with a clear scenario and clear connection types, providing a hierarchical basis for the subsequent S104 merging of knowledge chains to generate the main chain. The specific merging method needs to be flexibly selected according to the number of secondary categories.
[0075] One possible scenario is that if a certain level of scenario category corresponds to only one second-level connection type subcategory (i.e., the connection types of all knowledge chains within this scenario are highly uniform, with no significant differences in sub-processes), then the first-level scenario category is used as the unit for merging. For example, in the first-level scenario category of business expansion and installation, if all knowledge chains use parameter-dependent and time-series-related connection types as their core connection types, corresponding to only one parameter-time-series hybrid subcategory, then all knowledge chains under this first-level scenario category are directly merged to generate a single main chain covering the core logic of the entire business expansion and installation scenario. This ensures that the main chain can fully represent the unified business process of this scenario, avoiding logical fragmentation due to excessive splitting.
[0076] The second possible scenario is that if a certain level of scenario category corresponds to multiple second-level connection type subcategories (i.e., there are significant differences in sub-processes within the scenario, and the connection types of different sub-processes are significantly different), then the subcategories are merged separately. For example, in the first-level scenario category of fault repair, there are two second-level subcategories: a pure time-series subcategory and a time-series-parameter hybrid subcategory. In this case, the two subcategories need to be merged separately. The chains of the pure time-series subcategory are merged into a fast fault repair main chain (focusing on process efficiency, without parameter verification), and the chains of the time-series-parameter hybrid subcategory are merged into a refined fault repair main chain (focusing on equipment parameter matching, including parameter verification). Ultimately, two main chains are generated in the fault repair scenario, each adapting to different sub-process requirements, ensuring that the subsequently generated corpus can cover the diverse business scenarios within the scenario.
[0077] In this way, by using the two flexible merging methods mentioned above, we can ensure the compatibility of the main chain with the smart grid planning scenario, while also taking into account the differences in sub-processes within the scenario.
[0078] S104. Merge the knowledge chains within the same category to obtain the unique main chain corresponding to each knowledge chain class.
[0079] It's important to note that the merging here is not a simple chain splicing, but rather a process of extracting commonalities and retaining core elements through deduplication, completion, and regularization. After clustering in step S103, all chains within the same knowledge chain class revolve around the same smart power grid planning scenario, and the link types and business logic between nodes are highly similar. At this point, structural analysis is performed on similar chains, and the frequency and process position of each node are statistically analyzed. Nodes with high frequency and located in key process positions are identified as core nodes, and the main trunk is constructed according to the power grid business logic sequence. Nodes with a frequency higher than the secondary threshold are identified as differentiating nodes, forming branch paths. Finally, each branch path is embedded into the corresponding position in the main trunk, forming a unique main chain containing the main trunk and branch paths. Through this merging process, while retaining the core business logic of the scenario, differentiated sub-processes under similar scenarios are integrated.
[0080] Specifically, the merging operation must strictly adhere to the business attributes of smart grid planning to avoid disrupting the core logic. It mainly follows three principles: the principle of prioritizing business logic, which requires that the core business processes / relationships of the scenario be fully preserved in the main chain, and key nodes should not be lost or the logical order reversed due to merging; the principle of minimum redundancy, which requires that homogeneous nodes or duplicate links that appear repeatedly in multiple chains within the same category be removed; and the principle of standardizing node descriptions, which requires that nodes with inconsistent descriptions but the same meaning in the same type of chain be standardized into standard terminology in the field of smart grid planning during merging.
[0081] In practice, knowledge chains within the same category are merged to obtain a unique main chain corresponding to each knowledge chain class, including:
[0082] (1) Perform structural analysis on multiple knowledge chains within the same category, extract the common nodes and fixed relationships of all chains, and form a master node cluster.
[0083] It should be noted that, firstly, for all knowledge chains of the same category (such as fault repair), each chain is broken down into its node sequence and connection order to form a structured chain graph (e.g., Chain 1: Fault alarm → Fault location → Fault isolation → Equipment repair; Chain 2: Fault alarm → Manual verification → Fault location → Fault isolation → Power restoration; Chain 3: Fault alarm → Fault location → Fault isolation → Power restoration). Next, through node-by-node comparison and sequence verification, nodes common to all chains (i.e., common nodes) are selected. By comparing the three chains, fault alarm, fault location, and fault isolation are nodes included in all chains; simultaneously, the connection order of these common nodes in each chain is verified, confirming that it is always Fault alarm → Fault location → Fault isolation (no reverse or skip connections). This fixed order is the common node association. Finally, the common nodes and fixed associations are integrated into a master node cluster, ensuring that the master node cluster reflects the core business logic of the same type of chain without changing the original connection order of common nodes in any single chain.
[0084] (2) Based on the fixed association relationship of the master node cluster, the main trunk of the main chain is constructed in combination with the power grid business logic.
[0085] It should be noted that the master node cluster is the foundation of the backbone, but it needs to supplement the core nodes required for the scenario in accordance with the business integrity requirements of smart grid planning (such nodes may not appear in all chains, but they are indispensable process links in the scenario). For example, in a fault repair scenario, restoring power supply is the ultimate goal of fault handling. Although chain 1 does not contain this node, according to the relevant specifications, restoring power supply needs to be added to the master node cluster as a necessary node; at the same time, the connection order of the nodes after supplementation should be verified to ensure that it conforms to the business logic. The common nodes formed after supplementation + necessary supplementary nodes + fixed association order constitute the backbone of the main chain, such as fault alarm → fault location → fault isolation → power restoration. This backbone retains the common logic of all chains, meets the integrity requirements of the scenario business, and the connection order of all nodes is consistent with the original order in a single chain (for example, the order of fault location → fault isolation is consistent in all chains, and the backbone follows this order).
[0086] (3) Extract the differentiated nodes and their corresponding relationships in each knowledge chain except for the main node cluster, and form the branch paths.
[0087] It should be noted that differentiated nodes refer to nodes that appear only in some chains and do not belong to the main node cluster. These nodes carry differences in sub-processes under the same scenario, and their association relationships in the original chain need to be extracted synchronously (ensuring that the connection order of a single chain is not changed). For example, comparing three chains, the maintenance equipment in chain 1 and the manual review in chain 2 are differentiated nodes. The association relationship of maintenance equipment in chain 1 is fault isolation → maintenance equipment → power restoration (after supplementing necessary nodes), and the association relationship of manual review in chain 2 is fault alarm → manual review → fault location. The combination of differentiated nodes + original chain association relationships is defined as a branch path, such as fault isolation → maintenance equipment → power restoration, or fault alarm → manual review → fault location. Each branch path completely retains the connection logic of differentiated nodes in the original chain, avoiding the confusion of sub-process order due to splitting.
[0088] (4) Embed the branch roads of each district into the corresponding positions of the trunk according to the original relationship to obtain a unique main chain containing the trunk and the branch paths.
[0089] It should be noted that the embedding position is determined by the association relationship of the branch paths. Find the overlapping nodes in the branch paths that coincide with the trunk nodes, and embed the branch path between the overlapping nodes. For example, in the branch path fault alarm → manual review → fault location, the fault alarm and fault location are overlapping nodes, so this branch path is embedded between the trunk fault alarm and fault location; in the branch path isolation fault → equipment repair → power restoration, the isolation fault and power restoration are overlapping nodes, so it is embedded between the trunk isolation fault and power restoration. After embedding, it is necessary to ensure that the connection order of the trunk nodes remains unchanged (still fault alarm → fault location → fault isolation → power restoration), and the original connection order of the branch paths remains unchanged (e.g., manual review is still after the fault alarm and before the fault location). The final unique main chain presents a trunk + branch structure. The trunk is fault alarm → fault location → fault isolation → power restoration, and the branch path is fault alarm → [manual review] → fault location, fault isolation → [equipment repair] → power restoration (the nodes in "[]" are differentiated nodes). This main chain retains the common core processes of all chains, fully covers the differences in sub-processes, and does not disrupt the inherent connection order of any single chain, providing a precise logical framework for generating scenario-based corpus templates that fit actual business needs.
[0090] S105. Generate a corpus template based on the main chain.
[0091] It should be noted that this step generates the corpus template based on the obtained main chain. Specifically, the core nodes and their logical relationships in the main chain are transformed into a fixed business expression framework, forming the backbone of the template. For distinguishing nodes in the main chain, optional expression branches are set at the corresponding positions in the template. Simultaneously, variable padding positions are reserved in the attribute information associated with the nodes. The final generated corpus template fully maps the structured logic of the main chain and conforms to the business expression habits of power grid planning.
[0092] It should be noted that the main chain, as the knowledge logic framework of a specific power grid intelligent planning scenario (such as fault repair and business expansion application), includes core nodes in the main body (such as fault alarm → fault location → fault isolation → power restoration) and distinguishing nodes in the branch paths (such as manual verification and equipment maintenance). The generation of the corpus template must fully map this structural feature. On the one hand, for the main body of the main chain, the core nodes and their associated logic are transformed into a fixed business expression framework. For example, fault alarm → fault location is transformed into "After a power grid fault occurs, the system automatically triggers a fault alarm. Maintenance personnel carry out fault location work based on the alarm information. The location result must include the fault type and specific location," ensuring that the template retains the core business process of the scenario. On the other hand, for the branch paths of the main chain, selectable expression branches are set at the corresponding positions. For example, between fault alarm and fault location, an optional expression "[If the alarm information is ambiguous, manual verification must be carried out first to confirm the authenticity of the alarm before location is carried out]" is added for the distinguishing node manual verification, which covers the differences in sub-processes without destroying the main logic.
[0093] Meanwhile, the corpus template needs to reserve variable padding spaces to support diverse subsequent generation. These padding spaces correspond to the attribute information of the main chain nodes (such as equipment parameters, time indicators, and numerical results). For example, after describing a fault location, set "location time ≤ __ minutes"; after describing power restoration, set "power supply reliability reaches __% after restoration". The settings of the padding spaces need to conform to the actual needs of power grid operations (such as time indicators conforming to operation and maintenance specifications, and numerical ranges matching equipment parameter standards). The final generated corpus template has structured logic (matching the main trunk and branches of the main chain), business professionalism (conforming to power grid planning expression habits), and generation flexibility (reserving variable padding space). It can be directly used as input for generative large models, ensuring that the subsequently generated corpus not only conforms to scenario requirements but also covers different combinations of business parameters and sub-processes.
[0094] S106. Input the corpus template into the generative large model to predict multiple different corpora in the field of smart grid planning corresponding to the corpus template.
[0095] It should be noted that after the corpus template is generated in step S105, the objects input into the generative large model have been transformed from knowledge chains into text-format corpus templates, rather than structured knowledge chains. During the generation process, it is important to pay attention to the correspondence between the template mask position and the original knowledge chain node. That is, the variable filling position in the template (i.e. the position to be predicted) should accurately correspond to the knowledge nodes and related content in the main chain generated in step S104, so as to ensure that the generated corpus still fits the business logic of the original chain.
[0096] Specifically, firstly, the compatibility between the corpus template and the generative large model is clarified. The corpus template is presented in the form of natural language text and has pre-set core business expressions corresponding to the scenario (such as "After a power grid failure occurs, [fault alarm method] is triggered, followed by [fault location operation]" in the fault repair scenario), optional sub-processes (such as "[If the alarm information is ambiguous, manual review is required]") and variable filling positions (i.e., mask positions). These mask positions are not randomly set, but correspond one-to-one with the knowledge nodes in the main chain (such as the fault alarm method mask position corresponds to the fault alarm node in the main chain, and the fault location operation mask position corresponds to the fault location node in the main chain). Each mask position is accompanied by business constraints associated with the original chain node (such as the fault location operation mask position is accompanied by the constraint that it must include the location tool + time ≤ 30 minutes, and this constraint originates from the attribute information of the fault location node in the main chain). Generative large models (usually Transformer architecture models fine-tuned from power grid domain corpora) first parse the expression framework and mask position rules of the text template, and then combine them with the business constraints of the original chain nodes to ensure that the generated content does not deviate from the professional logic of power grid planning.
[0097] In the specific generation process, the model's core capabilities revolve around text template mask filling, rather than chain parsing. Firstly, it ensures compliant filling: for the mask positions in the template corresponding to chain nodes, the model outputs content that conforms to constraints based on power grid domain knowledge (e.g., for the fault location operation mask position, it generates "Maintenance personnel used drone inspection + infrared thermometer for location, taking 25 minutes," which includes both the location tool and meets the time constraint). Secondly, it provides diverse expressions: without changing the core logic, it generates different wording or parameter values for the same mask position (e.g., for the fault location operation, it can generate "The fault point was determined using a portable fault locator, taking 28 minutes," differentiating itself from the former in terms of tools and time). Thirdly, it ensures textual coherence: based on the natural language context of the text template, the model ensures smooth connection between the mask position filling content and the fixed expressions in the template (e.g., "Subsequently carried out" and "After location was completed"), avoiding semantic breaks (e.g., generating "Subsequently carried out drone inspection to determine the fault location," rather than the isolated "Drone inspection").
[0098] In addition, it should be noted that the final output corpus for the field of smart grid planning must meet three requirements simultaneously: First, the text format must be compliant, with all corpus consisting of natural language text that conforms to the expression habits of power grid business; second, the node logic must be consistent, with the content generated by the template mask position being consistent with the business logic of the corresponding node in the original main chain; and third, the scenario adaptation must be accurate, with the corpus theme completely matching the power grid planning scenario (such as fault repair and business expansion application) corresponding to the template, without any scenario mismatch issues.
[0099] Specifically, the corpus template is input into a generative large model to predict multiple different corpora corresponding to the corpus template in the field of smart grid planning, including:
[0100] (1) Divide the corpus template into multiple consecutive positions to be predicted according to the order of the corresponding main chain. Each position corresponds to a knowledge node and related content in the main chain, and set a pre-condition logical constraint for each position.
[0101] Specifically, based on the node sequence of the main chain (e.g., fault alarm → fault location → fault isolation → power restoration), the corpus template is decomposed into consecutive positions to be predicted. For example, the template "After a power grid fault occurs, [fault alarm method] is triggered, followed by [fault location operation], and after location is completed, [fault isolation measures] are executed, ultimately achieving [power restoration effect]", can be divided into four positions to be predicted according to the main chain sequence: fault alarm method, fault location operation, fault isolation measures, and power restoration effect. Each position corresponds to a knowledge node and related content of the main chain (e.g., the fault alarm method corresponds to the specific implementation method of the fault alarm node).
[0102] Meanwhile, pre-conditions are set for each location to be predicted. The constraints must conform to the power grid business rules to ensure that the subsequent generation does not deviate from the actual scenario. For example, the pre-condition for locating faults is that it must include the location tool (such as an infrared thermometer or drone inspection) and the location time (≤30 minutes). The pre-condition for isolating faults is that it must match the fault type (such as disconnecting the switches on both sides of the line for line faults, and shutting down the faulty equipment for equipment faults). These constraints are not only the boundary conditions for the generated content, but also the key to ensuring the compliance of the corpus business.
[0103] (2) Input the corpus template and the first position to be predicted into the generative large model to generate the content of the first position to be predicted.
[0104] Specifically, the complete corpus template (including placeholders for all positions to be predicted) and the first position to be predicted (such as fault alarm method) are fed as joint input to the generative large model. The large model will first parse the overall scene of the template (such as fault repair) and the location of the first position, and then combine its learned power grid domain knowledge (such as common fault alarm methods) to generate content that meets the pre-constraints.
[0105] For example, if the prerequisite constraint for a fault alarm method is that it must include an automatic alarm trigger source (such as a SCADA system or online monitoring device), the model might generate content like "The SCADA system detects a sudden change in line current, automatically triggers a fault alarm, and the alarm information is simultaneously pushed to the mobile terminal of maintenance personnel." This satisfies the constraint of an automatic alarm trigger source and aligns with the actual alarm process of the power grid. The quality of the generated content in the first location directly affects the logical coherence of subsequent locations. Therefore, after generation, the model's logic must be implicitly validated (e.g., whether the content matches the scenario theme) to ensure there are no significant business deviations.
[0106] (3) For each subsequent position to be predicted, input the generated content, the pre-logical constraints of the current position to be predicted and the remaining part of the corpus template into the generative large model to generate the content of the current position.
[0107] Unlike the first location, the generation of subsequent locations to be predicted (such as fault location operations and fault isolation measures) depends on the context information of the already generated content. For example, if the first location content is "The SCADA system detects a sudden change in line current, automatically triggers a fault alarm, and the alarm information is simultaneously pushed to the mobile terminal of the maintenance personnel," then when generating the fault location operation, this already generated content, the pre-constraints of the fault location operation (including the location tool and time consumption), and the remaining part of the template (after the location is completed, [fault isolation measures] are executed, ultimately achieving [power restoration effect]) need to be jointly input into the large model.
[0108] The large model generates matching content based on the contextual logic of the generated content (such as the alarm originating from a sudden change in line current, and the location needs to be expanded to the line). For example, "The maintenance personnel rushed to the site with an infrared thermometer and carried out line location based on the drone inspection data. It took 22 minutes to determine that the fault point was the breakdown of the insulator on the #15 tower of the 110kV East Line". This content includes the location tools such as the infrared thermometer and the drone inspection, meets the constraint of time ≤30 minutes, and forms a logical closed loop with the generated cause of the line current sudden change alarm. This avoids the logical break problem of the alarm being a line fault but the location is for the substation equipment.
[0109] (4) After generating all the positions to be predicted in a single corpus, adjust the content according to the text differences between the generated corpora to ensure the distinguishability of the generated corpora.
[0110] Once all the locations to be predicted for a single corpus (such as a complete fault repair corpus) has been generated, it needs to be compared with the previously generated corpus of the same template in terms of textual differences. The comparison dimensions include differences in variable values (such as location time, equipment model) and differences in expression (such as operation description, result presentation).
[0111] For example, if the location time for corpus A is 22 minutes and the location tool is an infrared thermometer + drone, and the location time for newly generated corpus B is also 22 minutes and the location tool is described as drone inspection + infrared thermometer, and the difference between the two is less than a preset threshold (e.g., 30%), then the content of corpus B needs to be adjusted. The location time for corpus B can be modified to 28 minutes, and the location tool description can be optimized to a portable line fault locator combined with drone aerial photography, increasing the difference between corpus B and corpus A to over 50%. The adjustment process must ensure that it does not violate the pre-existing logical constraints (e.g., the time remains ≤30 minutes), achieving corpus diversity with different details in the same scenario while maintaining compliance.
[0112] (5) Repeat the position-by-position prediction and discrimination adjustment steps to obtain multiple corpora in the field of smart grid planning that meet the template logic constraints and have text discrimination.
[0113] Specifically, repeat steps (2)-(4), first generate new individual corpora in the order from the first position to the subsequent positions, and then adjust them by difference comparison to ensure their distinguishability from existing corpora until the number of generated corpora meets the requirements. In addition, the final output of multiple domain corpora must simultaneously meet the following requirements: all content of each corpus must meet the preconditions of the corresponding position to be predicted, and the context logic must be coherent (such as fault cause, location, isolation, and recovery measures forming a complete business chain), and any two corpora must have significant differences in variable values (such as time consumption, equipment model, effect indicators) or expression methods (difference degree ≥ preset threshold).
[0114] Furthermore, it should be noted that after predicting multiple different corpora corresponding to the corpus template in the field of smart grid planning, the following are included:
[0115] (1) The corpus generation method is segmented to obtain multiple corpus generation stages.
[0116] It's important to note that corpus generation is a complete process from knowledge acquisition to corpus output. The entire methodology needs to be divided into independent and evaluable corpus generation stages according to technical logic to ensure that quality issues at each stage are identifiable. Based on the preceding description, these stages typically include: knowledge acquisition, knowledge chain construction, chain clustering, main chain merging, corpus template generation, and large-scale model corpus prediction. Each stage corresponds to a key technical aspect of corpus generation. This segmentation allows for phased quality control, preventing quality issues from becoming difficult to trace due to stage confusion.
[0117] (2) Identify the subjects to be evaluated in each stage and calculate the quality score of each subject to be evaluated.
[0118] Each corpus generation stage has its core evaluation subject, which is the key output or operation that directly affects the corpus quality at that stage. Appropriate evaluation indicators and scoring methods need to be designed for different subjects. For example, in the knowledge acquisition stage, the evaluation subject is the completeness and accuracy of formal knowledge. Evaluation indicators include core entity coverage (such as the completeness of equipment and process entity extraction) and attribute parameter accuracy (such as the error rate of transformer rated capacity parameters), calculated by (1 - number of erroneous entities / parameters ÷ total number of entities / parameters) × 100. In the corpus template generation stage, the evaluation subject is the logical completeness and scenario adaptability of the template. Evaluation indicators include core node coverage (whether the template covers all core nodes in the main chain) and constraint clarity (whether the pre-defined logical constraints of the position to be predicted are clear), calculated by combining domain expert scoring and machine automatic verification (e.g., expert scoring accounts for 60%, and machine verification accuracy accounts for 40%).
[0119] By assigning quantitative scores to each entity to be evaluated, the quality level at each stage can be intuitively reflected, providing a data foundation for subsequent overall evaluation.
[0120] (3) Calculate the influence weight of each corpus generation stage on the predicted corpus of multiple smart grid planning domains.
[0121] The impact of different stages on the final corpus quality varies, necessitating the quantification of each stage's importance through influence weighting to ensure the overall evaluation results align with actual business practices. Weight calculations should be tailored to the business characteristics of smart grid planning, employing expert scoring combined with business relevance analysis. For instance, stages with the greatest impact on corpus logic and scenario adaptability (such as the main chain merging stage and corpus template generation stage) are assigned higher weights (e.g., 0.25 each) because they directly determine the core framework of the corpus; stages supporting the basic quality of the corpus (such as the knowledge acquisition stage and knowledge chain construction stage) are assigned medium weights (e.g., 0.2 each); and the large-model corpus prediction stage, which significantly impacts corpus diversity but whose core logic depends on preceding stages, is assigned a lower weight (e.g., 0.1). The sum of all stage weights is 1 to ensure the standardization of subsequent weighted calculations, and the weights can be fine-tuned according to the needs of actual business scenarios (such as fault repair and grid optimization) to reflect the flexibility of the evaluation.
[0122] (4) Evaluate the predicted corpus based on the weighted sum of the influence weights and the corresponding quality scores.
[0123] Specifically, the overall quality score of the corpus is obtained by weighted summation, and the formula is: Overall Quality Score = Σ (Quality Score of a Certain Stage × Influence Weight of that Stage). For example, if the score for the knowledge acquisition stage is 90 (weight 0.2), the score for the knowledge chain construction stage is 85 (weight 0.2), the score for the chain clustering stage is 88 (weight 0.1), the score for the main chain merging stage is 92 (weight 0.25), the score for the corpus template generation stage is 89 (weight 0.25), and the score for the large model corpus prediction stage is 91 (weight 0.1), then the overall quality score = 90 × 0.2 + 85 × 0.2 + 88 × 0.1 + 92 × 0.25 + 89 × 0.25 + 91 × 0.1 = 89.15.
[0124] The overall quality score can be directly used to determine whether the corpus meets the standards, and also provides a quantitative basis for selecting high-quality corpora to train agents, avoiding evaluation bias caused by subjective judgment. After obtaining the overall quality score of the corpus, it is necessary to further locate the root causes of quality problems and provide optimization directions to ensure the quality of subsequent corpus generation or agent training.
[0125] It should also be noted that, after evaluating the predicted corpus based on the weighted sum of the aforementioned influence weights and corresponding quality scores, the evaluation includes:
[0126] (1) Compare the scores of each stage with the preset stage qualification threshold, and select the stages with scores lower than the stage qualification threshold as abnormal stages.
[0127] Each corpus generation stage has an independent acceptable threshold (set according to power grid industry standards and business needs, such as 85 points for the knowledge acquisition stage and 90 points for the main chain merging stage). By comparing the actual scores of each stage with the thresholds, stages with scores below the acceptable thresholds are identified as abnormal stages. For example, if the chain clustering stage scores 82 points (threshold 85 points) and the large model corpus prediction stage scores 83 points (threshold 85 points), these two stages are judged as abnormal stages. The selection of abnormal stages is key to locating problems from the overall to the local, avoiding the neglect of implicit quality problems in local stages due to focusing only on the overall score (e.g., a low score in a certain stage may lower the overall score, but if it is not identified separately, it cannot be optimized in a targeted manner).
[0128] (2) For a determined abnormal stage, analyze the scores of multiple evaluation indicators included in the stage, and determine the evaluation indicators whose scores are lower than the corresponding qualified threshold as abnormal indicators.
[0129] It should be noted that each anomaly stage contains multiple evaluation indicators, and the scores of these indicators need to be further broken down to pinpoint specific problem areas. For example, if the chain clustering stage is an anomaly stage (score 82), its evaluation indicators include intra-class chain similarity (score 78, threshold 85), scene label matching degree (score 88, threshold 85), and clustering efficiency (score 85, threshold 80). In this case, the intra-class chain similarity score is lower than the threshold and is identified as an anomaly indicator. If the large model corpus prediction stage is an anomaly stage (score 83), among its evaluation indicators, corpus logical coherence (score 79, threshold 85) and expression diversity (score 86, threshold 85), corpus logical coherence is an anomaly indicator.
[0130] By breaking down the metrics, stage anomalies can be refined into metric anomalies, avoiding overly broad optimization directions (such as only knowing that there is a problem with chain clustering, but not knowing whether it is due to insufficient similarity or label mismatch).
[0131] (3) Based on the abnormal stage and the corresponding abnormal index, return the corresponding corpus generation stage that needs to be optimized.
[0132] By combining the correspondence between abnormal stages and abnormal indicators, specific optimization directions are derived, and the corpus generation stages that need adjustment are identified. For example, for the abnormal intra-cluster chain similarity in the chain clustering stage, the optimization direction is to adjust the weight coefficient of the link type (such as increasing the weight of time-series related links) and optimize the similarity calculation threshold. The stage that needs optimization is S103 (chain clustering stage). For the abnormal corpus logical coherence in the large model corpus prediction stage, the optimization direction is to strengthen the pre-logical constraints of the position to be predicted (such as adding constraint descriptions of parameter dependencies) and perform secondary fine-tuning of the generative large model in the power grid domain. The stage that needs optimization is S106 (large model corpus prediction stage).
[0133] It should be noted that the returned optimization guidelines need to be specific to the stage and operation to ensure that technical personnel can directly implement and adjust them.
[0134] After completing the corpus quality evaluation and anomaly optimization, it is necessary to select qualified corpus data to train the intelligent power grid planning agent and realize the actual business application of the corpus data.
[0135] Furthermore, after predicting and obtaining multiple different corpora corresponding to the corpus template in the field of smart grid planning, the process also includes:
[0136] (1) Select corpora with a comprehensive quality score not lower than the preset training threshold to construct a training dataset.
[0137] It's important to note that training the AI agent requires high-quality corpus support. Therefore, a training threshold is set (usually higher than the acceptable corpus threshold, such as 90 points), and corpora with an overall quality score ≥ the training threshold are selected to form the training dataset. For example, from 1000 generated corpora, 800 corpora with an overall quality score ≥ 90 points are selected, and 200 corpora with scores below the threshold are removed. When constructing the training dataset, attention must be paid to data balance. That is, ensure that the proportion of corpora from different sub-scenarios matches the actual business (e.g., 30% for fault repair scenarios, 25% for business expansion and installation scenarios, and 45% for network optimization scenarios), avoiding an overabundance of one type of corpus that leads to an unbalanced AI agent (e.g., only proficient in fault repair, not network optimization).
[0138] (2) The training dataset is classified according to the business scenarios of smart grid planning to form a scenario-based training subset corresponding to each scenario.
[0139] Smart grid planning encompasses multiple sub-scenarios, and the logic and business requirements of the corpus differ significantly across scenarios. Therefore, the training dataset needs to be categorized according to scenario, forming scenario-based training subsets. For example, the 800 training corpora can be divided into three subsets based on scenario: fault repair (240 corpora, including fault location, isolation, and recovery), business expansion application (200 corpora, including demand acceptance, capacity calculation, and equipment selection), and grid optimization (360 corpora, including load forecasting, N-1 verification, and economic evaluation). Each subset needs to be clearly labeled with a scenario tag, allowing for targeted input into the model during subsequent training. This ensures the intelligent agent can learn the specific business logic for different scenarios, improving scenario adaptability.
[0140] (3) Input the scenario-based training subset into the training model containing the corpus feature extraction and scenario knowledge fusion mechanism for iterative training, and monitor the completion rate of the agent's planning task in the corresponding scenario through the validation set.
[0141] It should be noted that the training model needs to be domain-adaptive, therefore a dual mechanism of corpus feature extraction and scenario knowledge fusion is designed. The corpus feature extraction mechanism extracts core business features from the corpus (such as transformer rated capacity of 50MVA and positioning time of 25 minutes) using a Transformer-based model and transforms them into vectors that the model can recognize. The scenario knowledge fusion mechanism integrates scenario-specific knowledge (such as isolation before repair in fault repair and safety before economy in grid optimization) into the model in the form of a knowledge graph to ensure that the model understands the scenario logic.
[0142] During training, each scenario-based training subset is input into the model one by one, and iterative training is carried out (e.g., 100 rounds of iteration). After each round of training, a validation set (e.g., 10% of the corpus extracted from the scenario-based subset, which does not participate in training) is used to monitor the completion rate of the agent's planning tasks in the corresponding scenario (e.g., whether it can generate compliant fault repair plans and network optimization plans based on the corpus).
[0143] (4) When the completion rate of the planning task of the intelligent agent reaches the preset standard, the training is stopped and an intelligent agent adapted to the intelligent planning business of the power grid is obtained.
[0144] It should be noted that a task completion standard is set (e.g., overall completion ≥ 95%, completion of each sub-scenario ≥ 90%), and the training effect of the agent is continuously monitored. If, after a certain iteration, the overall task completion rate of the agent reaches 95%, and the completion rates of fault repair, business expansion application, and grid optimization scenarios are 93%, 92%, and 96% respectively, all ≥ 90%, then training is stopped. If the standard is not met, iteration continues (e.g., adding 20 more iterations), or model parameters are adjusted (e.g., optimizing the learning rate) until the standard is met. The resulting agent can adapt to multiple business scenarios of smart grid planning and can generate compliant and accurate planning solutions based on the input scenario requirements.
[0145] In addition, it should be noted that the validation set is used to monitor the agent's completion rate of planning tasks in the corresponding scenario, including:
[0146] (1) Verify the planning results of the agent based on the preset task completion evaluation index.
[0147] It should be noted that the multi-dimensional task completion evaluation indicators cover the core capabilities of the intelligent agent. These indicators may include: logical compliance rate (the proportion of planning results that conform to the power grid business logic); parameter accuracy rate (the proportion of accurate equipment parameters and numerical indicators in the planning results); and scenario adaptability (the matching ratio between the planning results and the requirements of the input scenario). The planning results output by the intelligent agent are scored using a validation set, and the scores for each indicator are calculated.
[0148] (2) When any task completion evaluation index is lower than the corresponding threshold, locate the abnormal stage in the training process of the agent that is associated with the task completion evaluation index, associate the scenario-based training subset corresponding to the abnormal stage with the abnormal index, and generate repair prompt information.
[0149] If a certain indicator score is lower than a preset threshold (e.g., a logical compliance rate threshold of 95% but an actual score of 92%), then the abnormal stage associated with that indicator is identified, i.e., the link in the training process that might have caused the indicator to be low. Specifically, if the logical compliance rate is low, it may be associated with the corpus template generation stage (incomplete template logical constraints) or the scenario-based training subset (logical deviation in a certain scenario corpus). Further, the scenario-based training subset corresponding to the abnormal stage (e.g., the fault repair subset) is associated with the abnormal indicator (low logical compliance rate), generating repair prompts. For example, if the fault repair subset corpus contains logical deviations that were not repaired after isolation, then such compliant corpus needs to be supplemented, and the pre-isolation logical constraints from isolation to repair need to be strengthened in the corpus template generation stage.
[0150] (3) Optimize the training dataset and training model according to the repair prompt information, and re-execute iterative training until the agent reaches the preset standard in all task completion evaluation indicators.
[0151] Based on the preceding explanations, optimize accordingly according to the repair suggestions. For optimizing the training dataset, supplement a certain number (e.g., 50) of compliant data points related to post-fault repair isolation and maintenance, and replace the data points with logical deviations (e.g., 10) in the fault repair subset. For optimizing the training model, adjust the constraint settings in the data point template generation stage, strengthen the logical connection between isolation and maintenance, regenerate the template, and supplement the training data. After optimization, re-execute iterative training, and monitor the task completion evaluation indicators again. Repeat this process until all indicators reach the preset standards (e.g., logical compliance rate 96%, parameter accuracy 93%, scenario adaptability 95%), ensuring that the agent possesses stable and comprehensive smart grid planning capabilities.
[0152] The method provided in this embodiment first uses formalized knowledge such as business rules, entity parameters, and logical associations in the field of smart grid planning as input. By parsing the entities and attributes in the formalized knowledge and establishing node associations that conform to the business logic of the power grid, a standardized knowledge chain is constructed. Then, the chain is clustered in a scenario-based manner by identifying the link type to ensure that each type of chain fits the core business logic of a specific planning scenario. Subsequently, similar chains are merged to generate a unique main chain, and a corpus template is constructed based on the main chain. This further solidifies the professional knowledge of the power grid field into the internal framework of the corpus, so that the generated corpus naturally has the professional attributes of power grid planning. This effectively avoids problems such as equipment parameter errors, reversed process order, and scenario misalignment that are prone to occur in general generation techniques, ensuring that the corpus can accurately match the actual business needs of smart grid planning.
[0153] Secondly, by using automated knowledge parsing algorithms, chain clustering algorithms, and main chain merging logic, the inefficient traditional method of manually organizing corpora is replaced, significantly improving the conversion efficiency from knowledge to corpus templates. At the same time, relying on the batch prediction capabilities of generative large models, under the pre-constraint logic of corpus templates and the discrimination adjustment mechanism, multiple corpora with differentiated expressions and compliant parameters can be quickly output. This not only meets the large-scale training needs of agents for tens of thousands to hundreds of thousands of corpora, but also ensures the diversity of corpora on the basis of compliance, avoiding the problem of insufficient generalization ability of agents caused by homogenization of corpora.
[0154] Finally, in the knowledge chain construction stage, deduplication and standardization processes are used to ensure the consistency of the knowledge chain. In the corpus prediction stage, positional logical constraints and adjustments for differences in generated content are used to ensure the logical coherence and quality stability of the corpus. At the same time, this embodiment establishes a systematic transformation path from formal knowledge to knowledge chain to main chain to corpus template to domain corpus. This allows for the efficient reuse of knowledge resources accumulated in the power grid planning field, such as equipment ledgers, business rule bases, and historical cases. This not only reduces the knowledge input cost of corpus generation but also allows for dynamic tracking of knowledge iteration, further enhancing the practical value and timeliness of the corpus. This provides reliable data support for the training of intelligent agents in power grid intelligent planning and the automated generation of business documents.
[0155] Example 2
[0156] Corresponding to the aforementioned embodiment of a corpus generation method for the field of smart grid planning, this application also provides an embodiment of a corpus generation device for the field of smart grid planning.
[0157] Figure 2 This is a schematic diagram of the second embodiment of the corpus generation device for smart grid planning provided in this application. Please refer to... Figure 2 The apparatus provided in this embodiment includes an acquisition module 210, an identification module 220, a clustering module 230, a merging module 240, a generation module 250, and a prediction module 260.
[0158] The acquisition module 210 is used to acquire formal knowledge of data in the field of smart grid planning.
[0159] The identification module 220 is used to identify the formal knowledge and obtain multiple knowledge chains, each knowledge chain including multiple knowledge nodes, and each node is an entity in the formal knowledge;
[0160] The clustering module 230 is used to identify the link type of the knowledge chain and cluster the multiple knowledge chains. The clustered knowledge chain classes are themed around the smart grid planning scenario.
[0161] The merging module 240 is used to merge various knowledge chains within the same category to obtain a unique main chain corresponding to each knowledge chain class;
[0162] The generation module 250 is used to generate a corpus template based on the main chain;
[0163] The prediction module 260 is used to input the corpus template into a generative large model and predict multiple different corpora in the field of smart grid planning corresponding to the corpus template.
[0164] The apparatus of this embodiment can be used to perform... Figure 1 The steps of the method embodiment shown are similar in principle and process, and will not be repeated here.
[0165] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0166] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0167] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A corpus generation method for the field of smart power grid planning, characterized in that, The method includes: Acquire formal knowledge of data in the field of smart grid planning; Identify the formal knowledge to obtain multiple knowledge chains, each knowledge chain including multiple knowledge nodes, and each node being an entity in the formal knowledge; The knowledge chains are clustered by identifying the link types of the knowledge chains, and the clustered knowledge chain classes are themed around the smart grid planning scenario. Merge the knowledge chains within the same category to obtain the unique main chain corresponding to each knowledge chain class; Generate corpus templates based on the main chain; The corpus template is input into a generative large model to predict multiple different corpora in the field of smart grid planning corresponding to the corpus template. The step of inputting the corpus template into a generative large model to predict multiple different corpora corresponding to the corpus template in the field of smart grid planning includes: The corpus template is divided into multiple consecutive positions to be predicted according to the order of the corresponding main chain. Each position corresponds to a knowledge node and related content in the main chain, and a pre-condition logical constraint is set for each position. Input the corpus template and the first position to be predicted into the generative large model to generate the content of the first position to be predicted. For each subsequent position to be predicted, the generated content, the pre-defined logical constraints of the current position to be predicted, and the remaining part of the corpus template are input into the generative large model to generate the content at the current position. After generating all the positions to be predicted for a single corpus, the content is adjusted based on the textual differences between the generated corpora to ensure the discriminative power of the generated corpora. By repeating the position-by-position prediction and discrimination adjustment steps, multiple corpora in the field of smart grid planning that conform to template logic constraints and have text discrimination are obtained.
2. The method according to claim 1, characterized in that, After the prediction obtains multiple different power grid smart planning domain corpora corresponding to the corpus template, it includes: The corpus generation method is segmented to obtain multiple corpus generation stages; Identify the entities to be evaluated in each stage and calculate the quality score for each of the entities to be evaluated. Calculate the influence weights of each corpus generation stage on the predicted corpora of multiple smart grid planning domains; The predicted corpus is evaluated based on the weighted sum of the influence weights and corresponding quality scores.
3. The method according to claim 2, characterized in that, The evaluation of the predicted corpus based on the weighted sum of the influence weights and corresponding quality scores includes: The scores of each stage are compared with the preset stage qualification threshold, and stages with scores lower than the stage qualification threshold are selected as abnormal stages. For a identified abnormal phase, analyze the scores of multiple evaluation indicators included in the phase, and identify the evaluation indicators whose scores are lower than the corresponding qualified threshold as abnormal indicators. Based on the abnormal stage and the corresponding abnormal indicators, return the corresponding corpus generation stage that needs to be optimized.
4. The method according to claim 1, characterized in that, After the prediction obtains multiple different power grid smart planning domain corpora corresponding to the corpus template, it includes: A training dataset is constructed by selecting corpora with a comprehensive quality score not lower than a preset training threshold; The training dataset is classified according to the business scenarios of smart grid planning, forming scenario-based training subsets corresponding to each scenario. The scenario-based training subset is input into the training model containing the corpus feature extraction and scenario knowledge fusion mechanism for iterative training, and the completion rate of the agent's planning task in the corresponding scenario is monitored through the validation set. When the completion rate of the planning task of the intelligent agent reaches the preset standard, the training stops, and an intelligent agent adapted to the intelligent planning business of the power grid is obtained.
5. The method according to claim 4, characterized in that, The step of monitoring the completion rate of the agent's planning task in the corresponding scenario through the validation set includes: The planning results of the intelligent agent are verified based on the preset task completion evaluation index; When any task completion evaluation index is lower than the corresponding threshold, locate the abnormal stage in the training process of the intelligent agent that is associated with the task completion evaluation index, associate the scenario-based training subset corresponding to the abnormal stage with the abnormal index, and generate repair prompt information. Based on the repair prompts, optimize the training dataset and training model, and re-execute iterative training until the agent reaches the preset standard on all task completion evaluation metrics.
6. The method according to claim 1, characterized in that, The identification of the formal knowledge yields multiple knowledge chains, including: The entities and attributes in the formal knowledge are analyzed, the entities are identified as knowledge nodes, and the attributes are used as the descriptive information of the nodes; Based on the entity relationships defined in the formal knowledge, connections between nodes are established to form a preliminary knowledge chain containing entity association sequences; The initial knowledge chain is deduplicated and standardized to obtain multiple standardized knowledge chains.
7. The method according to claim 1, characterized in that, The process of identifying the link types of the knowledge chains and clustering the multiple knowledge chains includes: Extract the link types between nodes in each knowledge chain; Configure weight coefficients for different link types, and calculate the similarity between knowledge chains based on the link types and weight coefficients; Based on the preset number of clusters for the business scenarios of smart grid planning, and using chain similarity as the metric, knowledge chains with similarity higher than a preset threshold are classified into the same category, resulting in knowledge chain classes themed around business scenarios. Each knowledge chain class is then labeled with a corresponding scenario tag.
8. The method according to claim 1, characterized in that, The process of merging knowledge chains within the same category to obtain a unique main chain corresponding to each knowledge chain class includes: Structural analysis is performed on multiple knowledge chains within the same category to extract common nodes and fixed relationships among all chains, forming a master node cluster; Based on the fixed association relationship of the master node cluster, the backbone of the main chain is constructed in combination with the power grid business logic; Extract the differentiated nodes and their corresponding relationships from each knowledge chain, excluding the main node cluster, to form branch paths; Embed the branch paths of each district into the corresponding positions of the main trunk according to their original relationships to obtain a unique main chain that includes the main trunk and the branch paths.
9. A corpus generation device for the field of smart grid planning, characterized in that, The device includes an acquisition module, an identification module, a clustering module, a merging module, a generation module, and a prediction module; The acquisition module is used to acquire formal knowledge of data in the field of smart grid planning; The identification module is used to identify the formal knowledge and obtain multiple knowledge chains, each knowledge chain including multiple knowledge nodes, and each node is an entity in the formal knowledge; The clustering module is used to identify the link type of the knowledge chain and cluster the multiple knowledge chains. The clustered knowledge chain classes are themed around the smart grid planning scenario. The merging module is used to merge various knowledge chains within the same category to obtain a unique main chain corresponding to each knowledge chain class; The generation module is used to generate corpus templates based on the main chain; The prediction module is used to input the corpus template into the generative large model and predict multiple different corpora in the field of smart grid planning corresponding to the corpus template. The step of inputting the corpus template into a generative large model to predict multiple different corpora corresponding to the corpus template in the field of smart grid planning includes: The corpus template is divided into multiple consecutive positions to be predicted according to the order of the corresponding main chain. Each position corresponds to a knowledge node and related content in the main chain, and a pre-condition logical constraint is set for each position. Input the corpus template and the first position to be predicted into the generative large model to generate the content of the first position to be predicted. For each subsequent position to be predicted, the generated content, the pre-defined logical constraints of the current position to be predicted, and the remaining part of the corpus template are input into the generative large model to generate the content at the current position. After generating all the positions to be predicted for a single corpus, the content is adjusted based on the textual differences between the generated corpora to ensure the discriminative power of the generated corpora. By repeating the position-by-position prediction and discrimination adjustment steps, multiple corpora in the field of smart grid planning that conform to template logic constraints and have text discrimination are obtained.
Citation Information
Patent Citations
Large model intelligent agent system based on thinking clustering planning and domain knowledge retrieval
CN118861305A
Block chain-based enterprise data processing method and system
CN120768522A