Risk identification method and device based on hierarchical function intention tree, equipment and medium

CN122736325APending Publication Date: 2026-09-11GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610924074.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]本发明提供一种基于层级功能意图树的风险识别方法、装置、设备及介质,以解决现有项目重复建设风险识别准确性较低的技术问题

Benefits of technology

[0009]The aforementioned risk identification method, apparatus, equipment, and medium based on a hierarchical functional intent tree can obtain project documents for newly submitted projects and historical projects in a historical project database. These project documents are then subjected to hierarchical parsing and structured segmentation to obtain a document structure tree. Functional segments are identified for each text unit in the document structure tree, and functional elements are extracted from the identified functional description segments to obtain structured functional units. These structured functional units are then aggregated layer by layer to form multiple functional clusters, and a corresponding parent node function name is generated for each functional cluster. A hierarchical functional intent tree for the newly submitted projects and historical projects is constructed based on the functional clusters and the parent node function names. Based on the hierarchical functional intent tree, semantic alignment is performed between each functional node in the newly submitted projects and each functional node in the historical projects. Consistency constraints are then applied, combining the path information and business object information of each functional node, to obtain the correspondence between functional nodes in the newly submitted projects and functional nodes in the historical projects. The functional coverage of the newly submitted projects relative to the historical projects is calculated based on the correspondence, and the redundant construction risk score is determined based on the functional coverage. In this invention, by performing hierarchical parsing and structured segmentation of project documents, identifying functional segments and extracting elements, a hierarchical functional intent tree is constructed. Then, by combining semantics, path and business object consistency, cross-project functional node alignment is achieved, functional coverage is calculated and the risk score of redundant construction is determined. This can accurately capture the essence of functions and effectively solve the pain points of long project documents, multiple levels, complicated expressions and difficulty in identifying implicit duplication, thereby improving the accuracy of identifying the risk of redundant construction in projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736325A_ABST
    Figure CN122736325A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of risk identification, and discloses a risk identification method and device based on a hierarchical function intent tree, equipment and medium, the method constructs a hierarchical function intent tree by performing hierarchical analysis and structured segmentation on project documents, function segment identification and element extraction, and then realizes cross-project function node alignment by combining semantic, path and business object consistency, calculates the function coverage and determines the repeated construction risk score result, can accurately capture the function essence, effectively solve the pain points of long project document length, multiple levels, miscellaneous expression and difficult identification of implicit repetition, thereby improving the accuracy of project repeated construction risk identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of risk identification technology, and in particular to a risk identification method, apparatus, device, and medium based on a hierarchical functional intent tree. Background Technology

[0002] As the digital transformation of power grid enterprises continues to advance, the number of IT projects related to business scenarios such as equipment operation and maintenance, production management, marketing services, dispatch support, safety supervision and quality control, and comprehensive management is constantly increasing. During the project initiation phase, it is typically necessary to submit application materials such as requirement documents, feasibility studies, and construction plans. These application materials are generally characterized by their length, multiple chapter levels, dense business terminology, scattered functional descriptions, and inconsistent terminology. The same type of construction objectives often appear in different business languages, technical terms, or different chapter organization methods, leading to overlaps between projects in terms of functional substance, business objects, data objects, or system boundaries, which are difficult to identify quickly and accurately manually. Especially in scenarios involving cross-unit, cross-professional, and cross-batch applications, reviewers face the realities of large-scale historical projects, broad comparison scopes, and numerous implicit duplications. Identifying the risk of duplicate project construction has become a crucial aspect of power grid IT project initiation review.

[0003] Currently, existing technologies for project duplication detection and requirements analysis mainly employ keyword matching, text similarity calculation, and vector retrieval. These methods primarily focus on capturing superficial similarities at the word and sentence level. When dealing with long, structured project documents with scattered descriptions across chapters, they struggle to accurately grasp high-level functional semantics and their hierarchical relationships, easily becoming limited to shallow comparisons at the word or sentence level. This results in low accuracy in identifying project duplication risks. Summary of the Invention

[0004] This invention provides a risk identification method, apparatus, device, and medium based on a hierarchical functional intent tree to solve the technical problem of low accuracy in identifying risks associated with repeated construction in existing projects.

[0005] Firstly, a risk identification method based on a hierarchical functional intent tree is provided, including: Obtain project documents for newly submitted projects and historical projects in the historical project database, and perform hierarchical parsing and structured segmentation on the project documents to obtain a document structure tree; Functional segment identification is performed on each text unit in the document structure tree, and functional elements are extracted from the identified functional description segments to obtain structured functional units. The structured functional units are aggregated layer by layer to form multiple functional clusters, and a corresponding parent node functional name is generated for each functional cluster. Based on the functional clusters and the parent node functional names, a hierarchical functional intent tree for the newly submitted project and each historical project is constructed. Based on the hierarchical functional intent tree, semantic alignment is performed between each functional node in the newly submitted project and each functional node in the historical project. Consistency constraints are then applied by combining the path information and business object information of each functional node to obtain the correspondence between the functional nodes in the newly submitted project and the functional nodes in the historical project. The functional coverage of the newly declared project relative to the historical project is calculated based on the correspondence, and the result of the duplication construction risk score is determined based on the functional coverage.

[0006] Secondly, a risk identification device based on a hierarchical functional intent tree is provided, including: The parsing and segmentation unit is used to obtain project documents of newly submitted projects and historical projects in the historical project database, and to perform hierarchical parsing and structured segmentation on the project documents to obtain a document structure tree. The identification unit is used to identify the functional segments of each text unit in the document structure tree, and extract functional elements from the identified functional description segments to obtain structured functional units. An aggregation building unit is used to aggregate the structured functional units layer by layer to form multiple functional clusters, and generate a corresponding parent node functional name for each functional cluster. Based on the functional clusters and the parent node functional names, a hierarchical functional intent tree for the newly submitted project and each historical project is constructed. The alignment unit is used to perform semantic alignment between each functional node in the newly submitted project and each functional node in the historical project based on the hierarchical functional intent tree, and to perform consistency constraints by combining the path information and business object information of each functional node, so as to obtain the correspondence between the functional nodes in the newly submitted project and the functional nodes in the historical project. The determining unit is used to calculate the functional coverage of the newly declared project relative to the historical project based on the correspondence, and to determine the duplication construction risk score result based on the functional coverage.

[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the risk identification method based on a hierarchical functional intent tree described above.

[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the risk identification method based on a hierarchical functional intent tree described above.

[0009] The aforementioned risk identification method, apparatus, equipment, and medium based on a hierarchical functional intent tree can obtain project documents for newly submitted projects and historical projects in a historical project database. These project documents are then subjected to hierarchical parsing and structured segmentation to obtain a document structure tree. Functional segments are identified for each text unit in the document structure tree, and functional elements are extracted from the identified functional description segments to obtain structured functional units. These structured functional units are then aggregated layer by layer to form multiple functional clusters, and a corresponding parent node function name is generated for each functional cluster. A hierarchical functional intent tree for the newly submitted projects and historical projects is constructed based on the functional clusters and the parent node function names. Based on the hierarchical functional intent tree, semantic alignment is performed between each functional node in the newly submitted projects and each functional node in the historical projects. Consistency constraints are then applied, combining the path information and business object information of each functional node, to obtain the correspondence between functional nodes in the newly submitted projects and functional nodes in the historical projects. The functional coverage of the newly submitted projects relative to the historical projects is calculated based on the correspondence, and the redundant construction risk score is determined based on the functional coverage. In this invention, by performing hierarchical parsing and structured segmentation of project documents, identifying functional segments and extracting elements, a hierarchical functional intent tree is constructed. Then, by combining semantics, path and business object consistency, cross-project functional node alignment is achieved, functional coverage is calculated and the risk score of redundant construction is determined. This can accurately capture the essence of functions and effectively solve the pain points of long project documents, multiple levels, complicated expressions and difficulty in identifying implicit duplication, thereby improving the accuracy of identifying the risk of redundant construction in projects. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a risk identification method based on a hierarchical functional intent tree according to an embodiment of the present invention; Figure 2 yes Figure 1 A schematic diagram of a specific implementation of step S110; Figure 3 yes Figure 1 A schematic diagram of a specific implementation of step S120; Figure 4 yes Figure 1 A schematic diagram of a specific implementation of step S130; Figure 5 yes Figure 1 A schematic diagram of a specific implementation of step S140; Figure 6 yes Figure 1 A schematic diagram of a specific implementation of step S150; Figure 7 This is a schematic block diagram of a risk identification device based on a hierarchical functional intent tree according to an embodiment of the present invention; Figure 8 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] The risk identification method based on a hierarchical functional intent tree provided in this invention can be applied to either a client or a server. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. Currently, the accuracy of identifying risks associated with redundant construction in existing projects is low. To address this problem, this invention proposes a risk identification method based on a hierarchical functional intent tree. This method constructs a hierarchical functional intent tree by performing hierarchical parsing and structured segmentation of project documents, identifying functional segments, and extracting elements. It then combines semantics, path, and business object consistency to achieve cross-project functional node alignment, calculates functional coverage, and determines the risk score for redundant construction. This method can accurately capture the essence of functions and effectively solves the pain points of long project documents, multiple levels, complex descriptions, and difficulty in identifying implicit duplication, thereby improving the accuracy of identifying risks associated with redundant construction in projects. The invention will be described in detail below through specific embodiments.

[0014] Please see Figure 1 As shown, Figure 1 A flowchart of a risk identification method based on a hierarchical functional intent tree provided in an embodiment of the present invention includes the following steps: S110-S140.

[0015] S110. Obtain the project documents of each historical project in the newly submitted project and historical project database, and perform hierarchical parsing and structured segmentation on the project documents to obtain the document structure tree.

[0016] Specifically, obtain project documents for newly submitted projects, including requirements documents, feasibility study documents, and construction plans. At the same time, retrieve archived project documents for each historical project from the historical project database. It should be noted that the project documents can be in Word format, PDF format, or online forms.

[0017] Among them, such as Figure 2 As shown, step S110 includes steps S111-S116: S111. Identify the document format type of the project document; S112. If the document format type is a Word document, then read the title style tags in the project document to obtain format feature information; S113. If the document format type is PDF, then extract the format feature information based on the title number, font size, indentation position and line break position in the project document. S114. If the document format type is an online form, then read the text content within the online form; S115. Based on the format feature information or the text content, identify the hierarchical relationship between chapters in the project document, establish a tree-like hierarchical structure between projects, chapters, subsections and paragraphs, and form a document hierarchical framework. S116. Concatenate each paragraph node in the document hierarchy framework with its parent title path to generate the text unit with hierarchical context information, and summarize all the text units to obtain the document structure tree.

[0018] Specifically, the format of the project document is determined by the file extension, file header identifier, or document metadata to identify whether it is a Word document, PDF document, or an online form. If the document format is Word, the document parsing interface is called to read the heading style tags in the project document, including style levels such as Heading 1, Heading 2, and Heading 3, as well as formatting features such as bolding and font size variations, to obtain formatting feature information. It should be noted that the heading style tags in Word documents clearly define the chapter hierarchy and can be directly extracted and used. If the document format is PDF, formatting feature information is extracted based on the heading numbers, font sizes, indentation positions, and line break positions in the project document. Specifically, heading numbers such as "1", "1.1", and "1.1.1" reflect the chapter hierarchy; font sizes such as size 1 and size 2 distinguish between headings and body text; indentation positions determine the paragraph hierarchy; and line break positions identify chapter boundaries. If the document format is an online form, the text content within the online form is read. Understandably, online forms typically have a fixed field structure and hierarchical division, allowing direct access to chapter information. It should also be noted that parent-child node relationships are determined through numbering rules, style hierarchy, or indentation depth.

[0019] The document hierarchical framework concatenates each paragraph node with its parent title path to generate text units containing hierarchical contextual information. For example, if a paragraph is located under the path "3 Construction Content / 3.2 Inspection Management / 3.2.1 Defect Handling", then this text unit contains the complete title path and paragraph text. All text units are then aggregated to obtain the document structure tree. The document structure tree is a hierarchical organization of project documents represented by a tree-like data structure. Its root node corresponds to the project name, intermediate nodes correspond to chapter titles at each level, and leaf nodes correspond to paragraph text units. Each node contains four types of information: text content, title path, position in the original text, and project number. The purpose of the document structure tree is to convert the original unstructured project document into structured data, enabling subsequent functional segment identification to utilize both paragraph text information and the semantics of the chapter, avoiding misjudgments based solely on local words.

[0020] S120. Perform functional segment identification on each text unit in the document structure tree, and extract functional elements from the identified functional description segments to obtain structured functional units.

[0021] Specifically, such as Figure 3As shown, step S120 includes steps S121-S124: S121, adding a classification layer at the end of the large language model for paragraph type classification, and setting an instruction template to constrain the structured output of functional elements; S122, constructing a training sample set for functional segment recognition and functional element extraction, and performing supervised fine-tuning of the large language model based on the training sample set; S123, inputting each text unit in the document structure tree into the fine-tuned large language model, and outputting the paragraph type classification result of the text unit through the classification layer, and filtering out the text units whose paragraph type classification result is the functional description segment to obtain the target text unit; S124, inputting the target text unit into the fine-tuned large language model according to the format requirements of the instruction template to extract three types of functional elements including functional actions, business objects, and processing results, forming the structured functional unit. It should be noted that the basic model of the large language model can be Qwen2.5-7B-Instruct, which has strong Chinese understanding ability and long context processing ability. The classification layer classifies the hidden state vector of the input sample to obtain the probability of each class. The calculation formula is as follows: P(y|x) = softmax(W c h + b c ) Where x represents the input sample, h represents the semantic representation vector output by the fine-tuned large language model, and W c and b c Here, y represents the paragraph type classification label, which is a parameter for the classification layer. Paragraph type classification labels include functional description paragraphs, background introduction paragraphs, construction goal paragraphs, implementation guarantee paragraphs, investment calculation paragraphs, and other paragraphs. The instruction template is used to constrain the model's output structured results, ensuring that functional elements are presented in a unified format. It should also be noted that the training sample set comes from labeled power grid information project documents and is fine-tuned using LoRA, enabling domain adaptation without altering most of the original parameters, facilitating incremental training on the power grid project corpus. The fine-tuned large language model has the ability to identify functional description paragraphs from project documents. Understandably, when the probability corresponding to a functional description paragraph is greater than a preset threshold, the paragraph is identified as a functional description paragraph. For the selected functional description paragraphs, the instruction template guides the fine-tuned large language model to output structured data, extracting key functional information from the natural language description. For example, for the paragraph "Support on-site photo upload and generate defect work orders," the model output is: functional action equals photo upload and generation; business object equals on-site defect information and work orders; processing result equals forming a closed-loop handling entry point. The construction content, originally described in a scattered manner using natural language, has been transformed into comparable structured functional units.

[0022] S130. The structured functional units are aggregated layer by layer to form multiple functional clusters, and a corresponding parent node functional name is generated for each functional cluster. The hierarchical functional intent tree of the newly applied project and each historical project is constructed based on the functional clusters and the parent node functional names.

[0023] Specifically, such as Figure 4 As shown, step S130 includes steps S131-S134: S131, using the structured functional unit as the leaf node of the hierarchical functional intent tree, semantically encoding the functional actions, business objects, and processing results of the leaf node to obtain a semantic encoding result; S132, performing path encoding on the title path of the leaf node to obtain a path encoding result, and fusing the semantic encoding result and the path encoding result to obtain a comprehensive representation vector of the leaf node; S133, calculating the similarity between the leaf nodes based on the comprehensive representation vector, and merging leaf nodes with similarity higher than a preset threshold layer by layer starting from the leaf node at the bottom of the hierarchical result to form a multi-level functional cluster set; S134, inputting the leaf nodes in the functional cluster set into the large language model, generating the corresponding parent node function name under the constraints of a preset prompt template, creating a parent node with the parent node function name, and organizing the parent node as the upper-level node and the leaf nodes in the functional cluster set according to the subordinate relationship to obtain the hierarchical functional intent tree of the newly submitted project and each historical project. It should be noted that, in the specific implementation of step S131, each structured functional unit is reorganized into a unified input text, which is then input into a supervised fine-tuned large language model to extract semantic representation vectors. This supervised fine-tuned large language model is based on the Qwen2.5-7B-Instruct architecture and performs domain adaptation on the power grid project corpus using LoRA, accurately capturing deep semantic information of functional actions, business objects, and processing results. Understandably, the title path reflects the chapter position of the structured functional unit in the project document, such as "Construction Content / Inspection Management / Defect Handling". Path encoding is processed using an independent encoder to obtain path semantic vectors. In the process of fusing the semantic encoding results and the path encoding results, the content semantic vector and the path semantic vector are normalized separately and then linearly combined by the learned weight parameters, so that the comprehensive representation vector simultaneously contains functional content semantics and hierarchical contextual information.

[0024] It should also be noted that for any two leaf nodes, the similarity calculation comprehensively considers three dimensions: overall semantic proximity, business object similarity, and title path similarity, with each dimension having a weight coefficient of 1. A hierarchical agglomerative clustering algorithm is used during the leaf node merging process, iteratively merging from bottom to top, prioritizing the merging of the node pairs with the highest similarity each time. To avoid overly scattered or coarse clustering results, an intra-cluster consistency score is further calculated for each candidate functional cluster, using the following formula: in, Represents the k-th candidate functional cluster C k The intra-cluster consistency score is used to measure the semantic tightness of nodes within a candidate functional cluster. This represents the k-th candidate functional cluster; Indicates the number of leaf nodes included; The normalization coefficient is used to average the sum of pairwise similarities, compressing the total score to between 0 and 1, facilitating a unified threshold for judgment; f i f j Represents candidate functional cluster C k The i-th and j-th leaf nodes within; s ij Represents the leaf node f i with f j The overall node similarity between them.

[0025] when When the threshold is exceeded, it is assumed that several leaf nodes in the cluster can indeed belong to the same upper-level function; otherwise, the original fine-grained nodes are retained and not merged. The prompt template is fixed as: "The following functions belong to the same upper-level function. Please summarize a parent function name of no more than ten characters and output only the name." For example, for the group of leaf nodes "Defect Reporting, Work Order Dispatch, Rectification Tracking, Acceptance and Closure", the model outputs "Defect Closed-Loop Handling" as the parent node name. A parent node is created using the parent node function name. The parent node is organized as the upper-level node and the leaf nodes in the function cluster set according to the subordinate relationship. The above aggregation and naming process is repeated until the project root node is generated, resulting in the hierarchical function intent tree of the newly submitted project and each historical project.

[0026] S140. Based on the hierarchical functional intent tree, semantic alignment is performed between each functional node in the newly submitted project and each functional node in the historical project. Consistency constraints are then applied by combining the path information and business object information of each functional node to obtain the correspondence between the functional nodes in the newly submitted project and the functional nodes in the historical project.

[0027] Specifically, such as Figure 5As shown, step S140 includes steps S141-S144: S141, based on the hierarchical functional intent tree, construct node description text for each functional node in the newly submitted project and each functional node in the historical project, wherein the node description text includes the path where the node is located, the name of the parent node, the name of the current node, and functional element information; S142, use a dual-tower coding model to encode each functional node in the newly submitted project and each functional node in the historical project, respectively, to obtain a new project node vector and a historical project node vector, based on the new project node vector and the historical project node vector... The recall score is calculated using point vectors, and multiple candidate node pairs are selected based on the recall score. S143: For each candidate node pair, the input cross-coding model is jointly encoded to obtain a fine-grained ranking score. S144: The path consistency score and business object consistency score between the candidate node pairs are calculated. The recall score, fine-grained ranking score, path consistency score, and business object consistency score are weighted and fused to obtain a comprehensive alignment score. Based on the comprehensive alignment score, the correspondence between functional nodes in the newly submitted project and functional nodes in the historical project is determined. It should be noted that the construction of the node description text is not a simple concatenation of node names, but rather an integration of multi-level semantic information. The path of a node is recorded layer by layer from the project root node to the current node, forming a complete hierarchical link, such as "distribution network inspection system / construction content / inspection management / defect handling". The parent node name originates from the functional cluster naming result generated during the hierarchical functional intent tree construction phase. The current node name is a summary label of the structured functional unit at the leaf node level and the parent node functional name of the functional cluster at the intermediate node level. Functional element information is extracted directly from structured functional units, including the specific content of three slots: functional action, business object, and processing result, so that the node description has both structural and functional semantics.

[0028] Furthermore, the dual-tower coding model preferably adopts the BGE-M3 architecture, fine-tuned through domain contrastive learning, which performs excellently on Chinese semantic understanding tasks. During fine-tuning, contrastive learning training is performed using power grid project corpora. Positive sample pairs are different expressions of the same function, while negative sample pairs are literally similar but different functional descriptions. The encoding output is a fixed-dimensional semantic vector. When calculating the recall score, for each newly submitted project functional node, all nodes in the historical project database are traversed, cosine similarity is calculated, and the nodes are sorted. The top K nodes are selected as candidate nodes, with the K value dynamically adjusted according to the size of the historical project database to form candidate node pairs. The cross-coding model preferably adopts the bge-reranker-v2-m3 architecture. The input format is the concatenation of new project node text and historical project node text using a delimiter. The model performs bidirectional attention encoding on the concatenated sequence, extracts hidden state vectors at specific positions, and outputs a fine-ranking score after linear transformation and activation function. This model can capture fine-grained correspondences between two node texts and effectively identify synonymous scenarios, such as "task dispatch" and "job task issuance".

[0029] It should be further explained that the path consistency score extracts the title paths of the two nodes, comparing them layer by layer starting from the root node, and calculating the depth of the common prefixes. The business object consistency score extracts the business object terms of the two nodes, calculating term overlap or semantic vector similarity. The recall score, ranking score, path consistency score, and business object consistency score are weighted and fused to obtain a comprehensive alignment score, with the sum of all weight coefficients equal to 1.

[0030] In one embodiment, such as this embodiment, the step of determining the correspondence between functional nodes in the new application project and functional nodes in the historical project based on the comprehensive alignment score includes: comparing the comprehensive alignment score of each candidate node pair with a preset alignment threshold, and selecting candidate node pairs whose comprehensive alignment score is greater than or equal to the preset alignment threshold as node pairs to be matched; performing constraint matching between the set of functional nodes in the new application project and the set of functional nodes in the historical project based on the comprehensive alignment score of the node pairs to be matched, so as to select the matching combination with the highest comprehensive alignment score as the optimal matching result; and determining the correspondence between the new project node and the historical project node in the optimal matching result. It should be noted that by traversing all candidate node pairs, selecting candidate node pairs whose comprehensive alignment score is greater than or equal to the preset alignment threshold as node pairs to be matched, and directly eliminating candidate node pairs below the threshold, low-quality matching candidates can be filtered out, reducing the amount of matching computation and improving matching accuracy. Further, the constraint matching uses a constrained bipartite graph matching method to solve for the optimal matching, which can be expressed as: in, M represents the optimal matching result; u represents the set of all possible matching combinations; v represents the set of functional nodes in the newly submitted project; v represents the set of functional nodes in the historical project. This represents the overall alignment score. Matching constraints include: each newly submitted project functional node can match at most one historical project functional node, and each historical project functional node can be matched by at most one newly submitted project functional node; matching is only allowed between nodes at the same or adjacent levels. These constraints are implemented through preprocessing of bipartite graph edges; node pairs that do not meet the level constraints are not matched. The algorithm iteratively solves the problem until convergence, selecting the matching combination with the highest overall alignment score as the optimal matching result. It should be noted that the correspondence is stored in key-value pairs, where the key is the identifier of the newly submitted project functional node, and the value is the identifier of the matched historical project functional node and the overall alignment score.

[0031] S150. Calculate the functional coverage of the newly declared project relative to the historical project based on the correspondence, and determine the redundant construction risk score result based on the functional coverage.

[0032] Specifically, such as Figure 6 As shown, step S150 includes steps S151-S153: S151, assigning node weights to each functional node in the functional node set of the newly submitted project, and calculating the functional coverage of the newly submitted project relative to the historical project based on the node weights and the correspondence; S152, inputting the functional coverage and the risk characteristics derived from the functional coverage into the risk calibration model, and outputting the probability of duplicate construction risk; S153. Determine the risk level based on the probability of redundant construction risk and the preset threshold to obtain the redundant construction risk score result. It should be noted that the node weight is determined comprehensively based on the node's level and its position in the project document's chapter. Generally, nodes located in core chapters such as "Construction Content," "Functional Design," and "Business Process" have higher weights, while nodes in supplementary descriptions have lower weights. Intermediate-level nodes representing a complete business module typically have higher weights than leaf nodes representing only a single step. Core functional modules such as "Defect Closed-Loop Handling" and "Inspection Task Management" have higher weights than general functional nodes such as "Field Display" and "List Query." This method ensures that core functional modules occupy a higher proportion in the overall score, preventing them from being biased by detailed functions. Node weights can be expressed as: in, Indicates the hierarchical weight of the node. This represents the chapter confidence level of a node. and These are the weighting coefficients.

[0033] To further clarify, functional coverage includes single-project weighted coverage and library-level coverage. Single-project weighted coverage examines whether each functional node of a newly submitted project has a high-quality corresponding node in a historical project. The calculation formula is as follows: in, This represents the weighted coverage of a single project, which is the weighted coverage of the newly submitted project A relative to the historical project database B. This represents the set of functional nodes for a newly submitted project; w u This represents the node weight of the u-th functional node; u represents a specific functional node in the newly submitted project. Indicates the overall aligned score; This represents the set of functional nodes in a historical project; v represents a specific functional node within that historical project. If a high-quality corresponding node exists, it indicates that the corresponding function has been covered by the historical project; otherwise, it indicates that the function is still part of the newly added content. The coverage status of all nodes is weighted, summed, and normalized to obtain a coverage score between 0 and 1.

[0034] Library-level coverage examines whether each functional node of a newly submitted project can be found in the entire historical project database. Even if duplicate content is scattered across multiple historical projects, as long as the core functional nodes can mostly be found with high-scoring correspondences in the historical project database, the library-level coverage will still be high, thus accurately reflecting the risk of decentralized duplication of construction. Library-level coverage aims to address the most common technical challenge in project approval, namely, that duplicate content is often constructed and accumulated in a scattered manner, but it already exists overall. The formula for library-level coverage is as follows: in, This indicates library-level coverage, where K represents the historical project library. The meanings of the other parameters are the same as those in the single-project weighted coverage calculation formula, and will not be repeated here for the sake of simplicity.

[0035] It should be added that risk characteristics include core node hit rate, continuous function chain hit rate, and new node ratio. The continuous function chain hit rate refers to whether the parent node and multiple child nodes in a business process can simultaneously find corresponding relationships in historical projects. If the entire chain is matched, it indicates that the business function is not a localized coincidence but has been systematically built. The core node hit rate refers to the proportion of core functional nodes in the hierarchical functional intent tree of a newly submitted project that have corresponding nodes in the historical project database. Core functional nodes are usually located in the middle layer of the hierarchical functional intent tree, representing complete business function modules, such as "defect closed-loop handling," "inspection task management," and "inspection task dispatch," rather than individual operational steps at the bottom level. The new node ratio refers to the proportion of functional nodes in a newly submitted project whose corresponding relationships cannot be found in the historical project database.

[0036] It should also be noted that the risk calibration model uses a logistic regression model, outputting a probability of duplication risk between 0 and 1. The high-risk threshold is typically set to 0.7, and the low-risk threshold to 0.3. A duplication risk probability higher than the high-risk threshold is considered high, lower than the low-risk threshold is considered low, and in the middle range is considered medium risk. Furthermore, the distribution of single-project coverage and library-level coverage is used to further subdivide duplication types. If most core functions correspond to a single historical project, it is considered overall duplication; if only a single functional module highly overlaps with a historical project, it is considered partial duplication; if different functions are covered by multiple historical projects but there are still a few newly added functions, it is considered reusable duplication. The final result is the duplication risk score.

[0037] This invention addresses the challenges of lengthy, hierarchical, and fragmented project initiation documents for power grid information technology projects. Traditional text similarity methods can only capture superficial similarities at the word and sentence level, failing to effectively identify implicit duplication of construction work where expressions differ but functions are essentially the same. This invention constructs a hierarchical functional intent tree by performing hierarchical parsing and structured segmentation of project documents, identifying functional segments, and extracting elements. It then aligns cross-project functional nodes by combining semantics, path, and business object consistency, calculates functional coverage, and determines the duplication risk score. This transforms traditional text-based duplication detection into functional-based duplication identification, accurately uncovering implicit duplication risks and improving the accuracy of risk identification.

[0038] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0039] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.

[0040] In one embodiment, a risk identification device 200 based on a hierarchical functional intent tree is provided, which corresponds one-to-one with the risk identification method based on a hierarchical functional intent tree in the above embodiments. As shown in Figure 7, the risk identification device 200 based on a hierarchical functional intent tree includes a parsing and segmentation unit 201, an identification unit 202, an aggregation and construction unit 203, an alignment unit 204, and a determination unit 205. Detailed descriptions of each functional module are as follows: The parsing and segmentation unit 201 is used to obtain project documents of newly submitted projects and historical projects in the historical project database, and to perform hierarchical parsing and structured segmentation on the project documents to obtain a document structure tree. The identification unit 202 is used to identify the functional segments of each text unit in the document structure tree, and extract functional elements from the identified functional description segments to obtain structured functional units. The aggregation building unit 203 is used to aggregate the structured functional units layer by layer to form multiple functional clusters, and generate a corresponding parent node functional name for each functional cluster. Based on the functional clusters and the parent node functional names, a hierarchical functional intent tree for the newly submitted project and each historical project is constructed. Alignment unit 204 is used to perform semantic alignment between each functional node in the newly submitted project and each functional node in the historical project based on the hierarchical functional intent tree, and to perform consistency constraints by combining the path information and business object information of each functional node, so as to obtain the correspondence between the functional nodes in the newly submitted project and the functional nodes in the historical project. The determining unit 205 is used to calculate the functional coverage of the newly declared project relative to the historical project based on the correspondence, and to determine the duplication construction risk score result based on the functional coverage.

[0041] The aforementioned risk identification device based on a hierarchical functional intent tree can be implemented as a computer program, which can, for example... Figure 8 It runs on the computer device shown.

[0042] Please see Figure 8 , Figure 8 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. The computer device 300 is a device capable of identifying the risk of repeated project construction.

[0043] See Figure 8 The computer device 300 includes a processor 302, a memory, and a network interface 305 connected via a system bus 301. The memory may include a non-volatile storage medium 303 and internal memory 304.

[0044] The non-volatile storage medium 303 may store an operating system 3031 and a computer program 3032. When the computer program 3032 is executed, it causes the processor 302 to execute a risk identification method based on a hierarchical functional intent tree.

[0045] The processor 302 provides computing and control capabilities to support the operation of the entire computer device 300.

[0046] The internal memory 304 provides an environment for the execution of the computer program 3032 in the non-volatile storage medium 303. When the computer program 3032 is executed by the processor 302, the processor 302 can execute a risk identification method based on a hierarchical functional intent tree.

[0047] This network interface 305 is used for network communication with other devices. Those skilled in the art will understand that... Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device 300 to which the present invention is applied. The specific computer device 300 may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0048] The processor 302 is used to run a computer program 3032 stored in a memory to implement any embodiment of the risk identification method based on a hierarchical functional intent tree described above.

[0049] It should be understood that, in this embodiment of the invention, the processor 302 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0050] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by a processor in the computer system to implement the process steps of the embodiments of the above methods.

[0051] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program. When executed by a processor, the computer program causes the processor to perform any embodiment of the risk identification method based on a hierarchical functional intent tree described above.

[0052] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0053] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0054] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0055] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0056] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0057] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0058] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Since these modifications and variations fall within the scope of the claims and their equivalents, this invention also intends to include these modifications and variations.

[0059] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A risk identification method based on a hierarchical functional intent tree, characterized in that, include: Obtain project documents for newly submitted projects and historical projects in the historical project database, and perform hierarchical parsing and structured segmentation on the project documents to obtain a document structure tree; Functional segment identification is performed on each text unit in the document structure tree, and functional elements are extracted from the identified functional description segments to obtain structured functional units. The structured functional units are aggregated layer by layer to form multiple functional clusters, and a corresponding parent node functional name is generated for each functional cluster. Based on the functional clusters and the parent node functional names, a hierarchical functional intent tree for the newly submitted project and each historical project is constructed. Based on the hierarchical functional intent tree, semantic alignment is performed between each functional node in the newly submitted project and each functional node in the historical project. Consistency constraints are then applied by combining the path information and business object information of each functional node to obtain the correspondence between the functional nodes in the newly submitted project and the functional nodes in the historical project. The functional coverage of the newly declared project relative to the historical project is calculated based on the correspondence, and the result of the duplication construction risk score is determined based on the functional coverage.

2. The risk identification method based on hierarchical functional intent tree as described in claim 1, characterized in that, The step of performing hierarchical parsing and structured segmentation of the project document to obtain a document structure tree includes: Identify the document format type of the project document; If the document format type is a Word document, then the title style tags in the project document are read to obtain format feature information; If the document format is PDF, then the format feature information is extracted based on the title number, font size, indentation position, and line break position in the project document. If the document format type is an online form, then read the text content within the online form; Based on the format feature information or the text content, identify the hierarchical relationship between chapters in the project document, establish a tree-like hierarchical structure between projects, chapters, sections and paragraphs, and form a document hierarchical framework. Each paragraph node in the document hierarchy framework is concatenated with its parent title path to generate a text unit with hierarchical context information. All text units are then summarized to obtain the document structure tree.

3. The risk identification method based on a hierarchical functional intent tree as described in claim 1, characterized in that, The step of identifying functional segments in each text unit of the document structure tree and extracting functional elements from the identified functional description segments to obtain structured functional units includes: A classification layer is added at the end of the large language model for paragraph type classification, and instruction templates are set to constrain the structured output of functional elements; Construct a training sample set for functional segment recognition and functional element extraction, and perform supervised fine-tuning of the large language model based on the training sample set; Each text unit in the document structure tree is input into the fine-tuned large language model, and the paragraph type classification result of the text unit is output through the classification layer. The text units whose paragraph type classification result is the functional description segment are filtered out to obtain the target text unit. The target text unit is input into the fine-tuned large language model according to the format requirements of the instruction template to extract three types of functional elements, including functional actions, business objects, and processing results, to form the structured functional unit.

4. The risk identification method based on a hierarchical functional intent tree as described in claim 3, characterized in that, The steps of aggregating the structured functional units layer by layer to form multiple functional clusters, generating a corresponding parent node functional name for each functional cluster, and constructing a hierarchical functional intent tree for the newly submitted project and each historical project based on the functional clusters and the parent node functional names include: The structured functional unit is used as the leaf node of the hierarchical functional intent tree. The functional actions, business objects and processing results of the leaf node are semantically encoded to obtain the semantic encoding result. The title path of the leaf node is path encoded to obtain the path encoding result, and the semantic encoding result and the path encoding result are fused to obtain the comprehensive representation vector of the leaf node; The similarity between the leaf nodes is calculated based on the comprehensive representation vector. Starting from the leaf node at the bottom of the hierarchical result, leaf nodes with similarity higher than a preset threshold are merged layer by layer to form a multi-level functional cluster set. The leaf nodes in the functional cluster set are input into the large language model. Under the constraints of the preset prompt template, the corresponding parent node functional name is generated. A parent node is created with the parent node functional name. The parent node is used as the upper-level node and organized with the leaf nodes in the functional cluster set according to the subordinate relationship to obtain the hierarchical functional intent tree of the newly applied project and each historical project.

5. The risk identification method based on a hierarchical functional intent tree as described in claim 1, characterized in that, The step of semantically aligning each functional node in the newly submitted project with each functional node in the historical project based on the hierarchical functional intent tree, and combining the path information and business object information of each functional node to apply consistency constraints, to obtain the correspondence between functional nodes in the newly submitted project and functional nodes in the historical project, includes: Based on the hierarchical functional intent tree, node description text is constructed for each functional node in the newly submitted project and each functional node in the historical project. The node description text includes the path where the node is located, the name of the parent node, the name of the current node, and functional element information. A dual-tower coding model is used to encode each functional node in the newly submitted project and each functional node in the historical project, respectively, to obtain a new project node vector and a historical project node vector. A recall score is calculated based on the new project node vector and the historical project node vector, and multiple candidate node pairs are selected based on the recall score. For each candidate node pair, the candidate node pair is jointly encoded into the cross-coding model to obtain the fine-ranking score; Calculate the path consistency score and business object consistency score between the candidate node pairs, and then weight and fuse the recall score, the ranking score, the path consistency score, and the business object consistency score to obtain a comprehensive alignment score. Based on the comprehensive alignment score, determine the correspondence between the functional nodes in the newly submitted project and the functional nodes in the historical project.

6. The risk identification method based on a hierarchical functional intent tree as described in claim 5, characterized in that, The step of determining the correspondence between functional nodes in the newly submitted project and functional nodes in the historical project based on the comprehensive alignment score includes: The comprehensive alignment score of each candidate node pair is compared with a preset alignment threshold, and the candidate node pairs whose comprehensive alignment score is greater than or equal to the preset alignment threshold are taken as node pairs to be matched. Based on the comprehensive alignment score of the node pairs to be matched, constraint matching is performed between the set of functional nodes of the newly declared project and the set of functional nodes of the historical project, so as to select the matching combination with the highest comprehensive alignment score as the optimal matching result. The new project nodes and historical project nodes in the optimal matching results are determined as the corresponding relationships.

7. The risk identification method based on a hierarchical functional intent tree as described in claim 1, characterized in that, The steps of calculating the functional coverage of the newly declared project relative to the historical project based on the correspondence, and determining the redundant construction risk score result based on the functional coverage, include: Assign node weights to each functional node in the set of functional nodes of the newly submitted project, and calculate the functional coverage of the newly submitted project relative to the historical projects based on the node weights and the correspondence. The functional coverage and the risk characteristics derived from the functional coverage are input into the risk calibration model, and the probability of redundant construction risk is output. The risk level is determined based on the probability of redundant construction risk and a preset threshold to obtain the redundant construction risk score result.

8. A risk identification device based on a hierarchical functional intent tree, characterized in that, include: The parsing and segmentation unit is used to obtain project documents of newly submitted projects and historical projects in the historical project database, and to perform hierarchical parsing and structured segmentation on the project documents to obtain a document structure tree. The identification unit is used to identify the functional segments of each text unit in the document structure tree, and extract functional elements from the identified functional description segments to obtain structured functional units. An aggregation building unit is used to aggregate the structured functional units layer by layer to form multiple functional clusters, and generate a corresponding parent node functional name for each functional cluster. Based on the functional clusters and the parent node functional names, a hierarchical functional intent tree for the newly submitted project and each historical project is constructed. The alignment unit is used to perform semantic alignment between each functional node in the newly submitted project and each functional node in the historical project based on the hierarchical functional intent tree, and to perform consistency constraints by combining the path information and business object information of each functional node, so as to obtain the correspondence between the functional nodes in the newly submitted project and the functional nodes in the historical project. The determining unit is used to calculate the functional coverage of the newly declared project relative to the historical project based on the correspondence, and to determine the duplication construction risk score result based on the functional coverage.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the risk identification method based on a hierarchical functional intent tree as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the risk identification method based on a hierarchical functional intent tree as described in any one of claims 1 to 7.