Characteristic attribute marking method and system for power transmission and transformation project design file
By constructing a characteristic attribute system of power transmission and transformation engineering design files, using natural language processing and image recognition technology, combined with pre-trained models and machine learning models, the reliability and accuracy of power transmission and transformation engineering design files are solved, and an efficient automated marking method is realized.
Patent Information
- Application Number
- CN202510393459.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the marking method of power transmission and transformation engineering design documents has problems such as low reliability, poor accuracy and narrow application scope, especially because manual marking depends on personal quality and poor automation marking flexibility.
Natural language processing and image recognition technology are adopted, combined with pre-trained models and machine learning models, and the feature attribute system of power transmission and transformation engineering design files is constructed. Through preprocessing, feature attribute extraction, detailed branch level division, scene tree construction and correlation analysis, automated feature attribute marking is realized.
It realizes high reliability, accuracy and wide applicability of power transmission and transformation engineering design documents, improves marking efficiency and accuracy, and is suitable for design file management and retrieval of complex power systems.
Smart Images

Figure CN120296421A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electrical automation, and particularly relates to a method and system for marking characteristic attributes of design documents for power transmission and transformation projects. Background Art
[0002] With the development of economic technology and the improvement of people's living standards, electric energy has become an essential secondary energy source in people's production and life, bringing endless convenience to people's production and life. Therefore, ensuring the stable and reliable supply of electric energy has become one of the most important tasks of the power system.
[0003] In the field of power transmission and transformation project design, the management and retrieval of design documents are key links to ensure the smooth progress of power transmission and transformation projects, and they are also of great significance to the power system. At present, with the continuous expansion of the scale and the increase in complexity of the power system, the quantity and variety of design documents have increased exponentially.
[0004] Currently, for the marking of design documents for power transmission and transformation projects, generally two schemes of manual marking and automated marking are adopted. For the manual marking scheme, generally, designers manually classify and mark the design documents of power transmission and transformation projects according to experience; however, the manual marking scheme highly depends on the personal qualities of the marking personnel, and its reliability and accuracy are relatively low, and the efficiency is also low. For the automated marking scheme, generally, a marking scheme based on predefined rules is adopted, but the flexibility of such a scheme is poor and the applicable range is narrow. Summary of the Invention
[0005] One object of the present invention is to provide a method for marking characteristic attributes of design documents for power transmission and transformation projects with high reliability, good accuracy and wide applicable range.
[0006] Another object of the present invention is to provide a system for implementing the method for marking characteristic attributes of design documents for power transmission and transformation projects.
[0007] The method for marking characteristic attributes of design documents for power transmission and transformation projects provided by the present invention includes the following steps:
[0008] S1. Construct a characteristic attribute system for design documents of power transmission and transformation projects;
[0009] S2. Preprocess the target design documents of power transmission and transformation projects;
[0010] S3. Extract characteristic attributes from the preprocessed target design documents of power transmission and transformation projects based on natural language processing and image recognition technologies;
[0011] S4. Mark and verify the characteristic attributes obtained in step S3 to complete the marking of characteristic attributes of the target design documents of power transmission and transformation projects.
[0012] The construction of the characteristic attribute system for the power transmission and transformation project design documents described in step S1 specifically includes the following steps:
[0013] The constructed characteristic attribute system for the power transmission and transformation project design documents includes text structure tags, text language semantic tags, and text business tags;
[0014] Among them, the text structure tags include classification numbers, standard names, chapters, articles, paragraphs, figures, tables, and mathematical formulas; the text language semantic tags include scope, type, and content; the text business tags include project information, professional information, document type, and document content, where the project information includes project name, voltage level, and design stage, the professional information includes primary electrical, secondary electrical, and civil engineering, the document type includes drawings, specifications, calculation books, and equipment and material lists, and the document content includes equipment model, technical parameters, and installation method.
[0015] The preprocessing of the target power transmission and transformation project design documents described in step S2 specifically includes the following steps:
[0016] Preprocess the target power transmission and transformation project design documents to ensure the consistency of the document content; the preprocessing includes format unification and standardization processing.
[0017] Based on natural language processing and image recognition technologies, the feature attribute extraction of the preprocessed target power transmission and transformation project design documents includes the following steps:
[0018] Through a pre-trained model, predict and perform similarity scoring on the feature attributes corresponding to the text content of the target power transmission and transformation project design documents;
[0019] According to the similarity scoring, perform the division of the detailed branch levels, and classify the text content into the corresponding detailed branch levels to form the first classification data;
[0020] For each detailed branch, use a hierarchical method for recursive classification to form the second classification data;
[0021] According to the first classification data and the second classification data, construct a scenario tree;
[0022] According to the constructed scenario tree, perform the relevance analysis of the text content;
[0023] According to the relevance analysis results, perform classification annotation on the text content.
[0024] The step of predicting and performing similarity scoring on the feature attributes corresponding to the text content of the target power transmission and transformation project design documents through a pre-trained model specifically includes the following steps:
[0025] Using a pre-trained model, predict the characteristic attributes corresponding to the text content of the power transmission and transformation engineering design documents and perform similarity scoring;
[0026] According to the predicted similarity score, retain the text content with a similarity score higher than the set threshold;
[0027] For the retained text content, classify and save it according to the characteristic attribute system of the power transmission and transformation engineering design documents constructed.
[0028] The above-mentioned method of dividing the detailed branch level according to the similarity score and classifying the text content into the corresponding detailed branch level to form the first classification data specifically includes the following steps:
[0029] In the characteristic attribute system of each category, divide the detailed branch level according to the predicted similarity score; among them, the content with a similarity score higher than or equal to the first threshold with the characteristic attribute is classified as a high branch; the content with a similarity score lower than the first threshold and higher than or equal to the second threshold is classified as a medium branch; the content with a similarity score lower than the second threshold is classified as a low branch;
[0030] Classify the text content into the corresponding detailed branch level to form the first classification data.
[0031] The above-mentioned method of using a hierarchical method to perform recursive classification for each detailed branch to form the second classification data specifically includes the following steps:
[0032] For each detailed branch, perform secondary extraction on the text content therein;
[0033] According to the extracted text content, calculate the similarity between the text content and the characteristic attribute;
[0034] If the similarity of some text content with the characteristic attribute is lower than the set threshold, then classify this part of the text content as non-characteristic text content;
[0035] Classify all non-characteristic text content using a hierarchical method to form the second classification data.
[0036] The above-mentioned method of constructing a scenario tree according to the first classification data and the second classification data specifically includes the following steps:
[0037] According to the first classification data and the second classification data, adopt a top-down scheme to construct a scenario tree:
[0038] Create a root node: Set the processing scenario as S, then the root node R is represented as R = f root (S), where f root () is a function defining the creation of the root node;
[0039] Perform child node expansion: Create first-level child nodes based on the first-layer classification data; assume there are n categories in the first-layer classification data, denoted as C1, C2,..., C n ; For each category C i , create the corresponding first-level child node N i as N i = f sub (R, C i ), where f sub () is a function defining child node expansion; the first-level child nodes represent different given characteristic attribute categories;
[0040] Perform addition of branch nodes: Assume the first-level child node N i has m non-characteristic attribute categories, denoted as D i1 , D i2 ,..., D im ; Add detailed branch nodes B i to each first-level child node N ij as B ij = f branch (N i , D ij ), where f branch () is a function for adding branch nodes; the branch nodes are used to reflect non-characteristic attribute content at different depth levels;
[0041] Complete the construction of the scenario tree through the creation of the root node, child node expansion, and addition of branch nodes.
[0042] Perform relevance analysis of text content based on the constructed scenario tree, specifically including the following steps:
[0043] Represent node A in the scenario tree as vector a and node B as vector b; calculate the content relevance evaluation value between node A and node B using the following formula:
[0044]
[0045] In the formula, MSS(A, B) is the content relevance evaluation value between node A and node B; a * b represents the dot product of vector a and vector b; γ is a set characteristic correction factor; ||a|| represents the modulus of vector a; ||b|| represents the modulus of vector b; δ is a set balance coefficient; len(a) is the text length corresponding to vector a; len(b) is the text length corresponding to vector b;
[0046] According to the obtained content relevance evaluation values of Node A and Node B, perform the elimination operation: If the content relevance evaluation value is greater than the set threshold, it is determined that the content of Node A and Node B is repetitive, and one of the nodes is eliminated to improve the accuracy and refinement of the content of the scenario tree;
[0047] Calculate the content occupancy value of Node A using the following formula:
[0048]
[0049] In the formula, IV(A) is the content occupancy value of Node A; ω is the set weight value; n is the number of all nodes pointing to Node A; IV(T i ) is the content occupancy value of the i-th node T i pointing to Node A; γ is the node - to - node association stability coefficient; C(T i ) is the out - link number of the i-th node pointing to Node A; freq(A,T i ) is the occurrence frequency of Node A and Node T i in the power grid business association scenario; freq(T j ,T i ) is the occurrence frequency of Node T j and Node T i in the power grid business association scenario;
[0050] Construct a content association network through the above steps to analyze the connection strength and direction of each text content in the association network. For example, if Node A represents the text content of a certain power grid professional term, and it has a close association with the text contents of various different description forms corresponding to this power grid professional term, and occupies a core position in the association network, this indicates that the text content of this power grid professional term has a high association occupancy, and plays a key supporting role in the processing decisions of subsequent text contents of various different description forms.
[0051] Calculate the weight value of the node using the following formula:
[0052]
[0053] In the formula, αregion is the set correction coefficient; L is the level of the detailed branch; βdevice is the set weight adjustment factor; D is the depth - layer information; uniqueness(L,D) is the uniqueness index representing the relationship between the level of the detailed branch and the depth - layer information; Total uniqueness is the sum of the uniqueness indices of all node text contents;
[0054] Calculate the uniqueness index using the following steps:
[0055] Tokenize the text content of each node to obtain a keyword set; for the node text content S, the corresponding keyword set is Ks;
[0056] Calculate the keyword scarcity coefficient sc(k) using the following formula:
[0057]
[0058] where N is the total number of keywords in the scenario tree text library; n(k) is the number of occurrences of keyword k in the text library;
[0059] After vectorizing the text, calculate the similarity sim(S, T) between the node text content S and other text content T using the following formula:
[0060]
[0061] where Kt is the keyword set corresponding to other text content T;
[0062] Calculate the uniqueness index uniqueness(S, T) between the node text content S and other text content T using the following formula:
[0063]
[0064] where m is the total number of texts in the scenario tree;
[0065] The detailed branch level is divided according to the degree of detail of the text content. For example, from general descriptions to specific parameter settings, different levels are assigned; the depth layer information reflects the depth of the position of the node in the scenario tree; nodes located in the deeper layer of the scenario tree and with a high detailed branch level, such as the text content regarding the precise parameter adjustment of a certain device in a specific substation, will have a relatively higher weight value; in this way, the importance of key nodes in the entire scenario tree can be highlighted.
[0066] According to the results of the relevance analysis, classify and label the content of this article, which specifically includes the following steps:
[0067] Use a pre-trained machine learning model to automatically label each feature attribute;
[0068] The machine learning model includes several layers of neurons. The forward propagation process of the l-th layer of neurons is expressed as h l = δ(W l h l-1 + b l ), where h l is the output vector of the hidden layer of the l-th layer of neurons, δ() is the set activation function, W l is the hidden layer weight matrix of the l-th layer of neurons, bl It is the hidden layer bias vector of the l-th layer neurons.
[0069] The present invention also provides a system for implementing the method for marking characteristic attributes of the power transmission and transformation engineering design documents, including a system construction module, a file preprocessing module, an attribute extraction module, and an attribute marking module; the system construction module, the file preprocessing module, the attribute extraction module, and the attribute marking module are connected in series in sequence; the system construction module is used to construct a characteristic attribute system of the power transmission and transformation engineering design documents and upload data information to the file preprocessing module; the file preprocessing module is used to preprocess the target power transmission and transformation engineering design documents according to the received data information and upload the data information to the attribute extraction module; the attribute extraction module is used to extract characteristic attributes of the preprocessed target power transmission and transformation engineering design documents based on natural language processing and image recognition technologies and upload the data information to the attribute marking module; the attribute marking module is used to mark and verify the obtained characteristic attributes according to the received data information to complete the marking of the characteristic attributes of the target power transmission and transformation engineering design documents.
[0070] The method and system for marking characteristic attributes of the power transmission and transformation engineering design documents provided by the present invention, through extracting and analyzing the characteristic attributes of the target power transmission and transformation engineering design documents, not only realize the marking of the characteristic attributes of the power transmission and transformation engineering design documents, but also have high reliability, good accuracy, and good applicability. Brief Description of the Drawings
[0071] Figure 1 It is a schematic flow chart of the method of the present invention.
[0072] Figure 2 It is a schematic diagram of the functional modules of the system of the present invention. Detailed Embodiments
[0073] As Figure 1 shown is a schematic flow chart of the method of the present invention: The method for marking characteristic attributes of the power transmission and transformation engineering design documents disclosed by the present invention includes the following steps:
[0074] S1. Construct a characteristic attribute system of the power transmission and transformation engineering design documents; specifically including the following steps:
[0075] The constructed characteristic attribute system of the power transmission and transformation engineering design documents includes text structured tags, text language semantic tags, and text service tags;
[0076] Among them, the text structure tags include classification number, standard name, chapter, article, paragraph, figure, table and mathematical formula; the text language semantic tags include scope, type and content; the text business tags include project information, professional information, document type and document content, where the project information includes project name, voltage level and design stage, the professional information includes primary electrical, secondary electrical and civil engineering, the document type includes drawings, specifications, calculation sheets and equipment and material lists, and the document content includes equipment model, technical parameters and installation method;
[0077] S2. Preprocess the target power transmission and transformation project design documents; specifically, it includes the following steps:
[0078] Preprocess the target power transmission and transformation project design documents to ensure the consistency of the document content; the preprocessing includes format unification and standardization processing, such as document format conversion, text recognition, image processing, etc., and perform a preliminary analysis of the documents to extract key information such as text, images, and tables in the documents;
[0079] S3. Based on natural language processing and image recognition technologies, extract the characteristic attributes of the preprocessed target power transmission and transformation project design documents; it includes the following steps:
[0080] Through a pre-trained model, predict and perform similarity scoring on the characteristic attributes corresponding to the text content of the target power transmission and transformation project design documents; specifically, it includes the following steps:
[0081] Adopt a pre-trained model to predict and perform similarity scoring on the characteristic attributes corresponding to the text content of the power transmission and transformation project design documents;
[0082] According to the predicted similarity score, retain the text content with a similarity score higher than the set threshold;
[0083] Classify and save the content of the retained text according to the constructed characteristic attribute system of the power transmission and transformation project design documents;
[0084] According to the similarity score, perform the division of the detailed branch levels, and classify the text content into the corresponding detailed branch levels to form the first classification data; specifically, it includes the following steps:
[0085] In the characteristic attribute system of each category, perform the division of the detailed branch levels according to the predicted similarity score; among them, the content with a similarity score higher than or equal to the first threshold (preferably 80 points) with the characteristic attribute is classified as the high branch; the content with a similarity score lower than the first threshold and higher than or equal to the second threshold (preferably 60 points) is classified as the middle branch; the content with a similarity score lower than the second threshold is classified as the low branch;
[0086] Classify the text content into corresponding detail branch levels to form first classification data;
[0087] For each detail branch, a hierarchical method is used for recursive classification to form the second layer of classification data; specifically, the following steps are included:
[0088] For each detail branch, perform secondary extraction on the text content;
[0089] According to the extracted text content, the similarity between the text content and the feature attribute is calculated;
[0090] If the similarity between some text content and the characteristic attribute is lower than the set threshold, the text content is classified as non-characteristic text content;
[0091] All non-feature text content is recursively classified using a hierarchical method to form the second layer of classification data;
[0092] According to the first classification data and the second classification data, a scene tree is constructed; specifically, the steps include:
[0093] According to the first classification data and the second classification data, a top-down approach is adopted to construct the scene tree:
[0094] Create a root node: Set the processing scenario to S, then the root node R is represented by R = f root (S), where f root () is a function to define the root node; this root node represents the entire processing scene, and all subsequent nodes will be derived from it; it is the starting point and core of the entire scene tree, carrying the role of commanding the overall situation and laying the basic framework for the construction of subsequent nodes;
[0095] Expand the child nodes: Create the first-level child nodes based on the first-level classification data; set the first-level classification data to have n categories, represented by C1, C2, ..., C n ; For each category C i , create the corresponding first-level child node N i N i =f sub (R,C i ), where f sub() is a function that defines the extension of child nodes; the first-level child nodes represent different given characteristic attribute categories; for example, in scenarios related to power grid review, if the first-level classification data involves categories such as standardized basic and general standards, technical guidance general rules or guidelines for each specialty, then the corresponding first-level child nodes will expand around these categories; each child node focuses on a specific characteristic attribute field. In this way, the entire processing scenario is initially divided according to different characteristic dimensions, making the scenario tree start to show a branching structure and providing a direction for further refinement;
[0096] Add branch nodes: Set the first-level child node N i There are m non-characteristic attribute categories, denoted as D i1 , D i2 ,..., D im ; For each first-level child node N i Add detailed branch nodes B ij For B ij = f branch (N i , D ij ), where f branch () is a function for adding branch nodes; branch nodes are used to reflect non-characteristic attribute content at different depth levels; for example: in scenarios related to power grid review, assuming the child node is about technical guidance general rules or guidelines for each specialty, then the second-level classification data may involve non-characteristic attributes such as local regulations and internal management requirements; for these data, add corresponding branch nodes to the technical guidance general rules or guidelines child nodes for each specialty, and describe in detail different file information under local regulations, as well as different files under the internal management regulations branch, etc. Through such operations, continuously enrich the details of the scenario tree so that it can cover various types of information in the text content to be processed more comprehensively and meticulously;
[0097] Through the creation of the root node, the extension of child nodes, and the addition of branch nodes, the construction of the scenario tree is completed; this scenario tree can clearly show the hierarchical relationship and internal connection between various types of data, providing a solid foundation and convenient tool for subsequent data analysis, decision-making, etc. based on this scenario tree;
[0098] According to the constructed scenario tree, perform relevance analysis on the text content; specifically, it includes the following steps:
[0099] Represent node A in the scenario tree with vector a and node B with vector b; calculate the content relevance evaluation value of node A and node B using the following formula:
[0100]
[0101] Where MSS(A,B) is the content relevance evaluation value between node A and node B; a*b represents the dot product of vectors a and b; γ is a set feature correction factor; ||a|| represents the norm of vector a; ||b|| represents the norm of vector b; δ is a set balance coefficient; len(a) is the text length corresponding to vector a; len(b) is the text length corresponding to vector b;
[0102] According to the obtained content relevance evaluation value between node A and node B, perform the elimination operation: if the content relevance evaluation value is greater than the set threshold, it is determined that the contents of node A and node B are repeated, and one of the nodes is eliminated to improve the accuracy and refinement of the content of the scenario tree;
[0103] The content occupancy value of node A is calculated using the following formula:
[0104]
[0105] Where IV(A) is the content occupancy value of node A; ω is a set weight value; n is the number of all nodes pointing to node A; IV(T i ) is the content occupancy value of the i-th node T i pointing to node A; γ is the node-to-node association stability coefficient; C(T i ) is the number of outgoing links of the i-th node pointing to node A; freq(A,T i ) is the occurrence frequency between node A and node T i in the power grid service association scenario; freq(T j ,T i ) is the occurrence frequency between node T j and node T i in the power grid service association scenario;
[0106] If the text content of a power grid fault warning is closely related to the text content of multiple fault handling measures and occupies a core position in the association network, this indicates that the warning text content has a high association occupancy and plays a key supporting role in subsequent fault handling decisions;
[0107] Construct a content association network through the above links to analyze the connection strength and direction of each text content in the entire network; for example, node A represents the text content of a certain power grid professional term, which is closely related to the text content of various different description forms corresponding to this power grid professional term and occupies a core position in the association network, indicating that the text content of this power grid professional term has a high association occupancy and plays a key supporting role in subsequent processing decisions for various different description forms of text content.
[0108] The weight value of the node is calculated using the following formula:
[0109]
[0110] Where αregion is the set correction coefficient; L is the level of the detail branch; βdevice is the set weight adjustment factor; D is the depth layer information; uniqueness(L,D) is the uniqueness index representing the relationship between the level of the detail branch and the depth layer information; Total uniqueness is the sum of the uniqueness indices of all node text contents.
[0111] The following steps are used to calculate the uniqueness index:
[0112] The text content of each node is tokenized to obtain a keyword set; for the node text content S, the corresponding keyword set is Ks;
[0113] The keyword scarcity coefficient sc(k) is calculated using the following formula:
[0114]
[0115] Where N is the total number of keywords in the scenario tree text library; n(k) is the number of times the keyword k appears in the text library;
[0116] After vectorizing the text, the similarity sim(S,T) between the node text content S and other text content T is calculated using the following formula:
[0117]
[0118] Where Kt is the keyword set corresponding to the other text content T;
[0119] The uniqueness index uniqueness(S,T) between the node text content S and other text content T is calculated using the following formula:
[0120]
[0121] Where m is the total number of texts in the scenario tree;
[0122] The detail branch level is divided according to the degree of detail of the text content. For example, from general descriptions to specific parameter settings, different levels are assigned; the depth layer information reflects the depth of the node's position in the scenario tree. Nodes located deeper in the scenario tree and with a high level of detail branch, such as text content about precise parameter adjustment of a certain device in a specific substation, will have a relatively higher weight value. In this way, the importance of key nodes in the entire scenario tree can be highlighted.
[0123] After completing the above series of analyses, the content of this article will be classified and labeled, specifically including the following steps:
[0124] According to the results of the relevance analysis, classify and label the content of this article; specifically including the following steps:
[0125] Use a pre-trained machine learning model to automatically label each feature attribute;
[0126] The described machine learning model includes several layers of neurons, and the forward propagation process of the l-th layer of neurons is expressed as h l = δ(W l h l-1 + b l ), where h l is the output vector of the hidden layer of the l-th layer of neurons, δ() is the set activation function, W l is the weight matrix of the hidden layer of the l-th layer of neurons, and b l is the bias vector of the hidden layer of the l-th layer of neurons;
[0127] Through such layer-by-layer operations, the algorithm can deeply analyze various details of the feature attributes, such as its structural characteristics, data format, and associations with other attributes, and then generate metadata covering extremely detailed marking information; these metadata not only accurately label the categories of the feature attributes, but also include relevant confidence values, marking bases, and any possible supplementary explanations;
[0128] S4. Mark and verify the feature attributes obtained in step S3 to complete the marking of the feature attributes of the target power transmission and transformation project design document;
[0129] Specifically in implementation, the verification process can adopt the method of manual review or cross-validation.
[0130] Finally, the feature attribute information can be associated with the design document and stored in the database in the *.xml format for convenient subsequent retrieval, analysis, and application.
[0131] Such as Figure 2The following is a schematic diagram of the functional modules of the system of the present invention: The system for implementing the method for marking the characteristic attributes of the power transmission and transformation project design documents disclosed in the present invention includes a system construction module, a file preprocessing module, an attribute extraction module, and an attribute marking module; the system construction module, the file preprocessing module, the attribute extraction module, and the attribute marking module are connected in series in sequence; the system construction module is used to construct the characteristic attribute system of the power transmission and transformation project design documents and upload the data information to the file preprocessing module; the file preprocessing module is used to preprocess the target power transmission and transformation project design documents according to the received data information and upload the data information to the attribute extraction module; the attribute extraction module is used to extract the characteristic attributes of the preprocessed target power transmission and transformation project design documents based on natural language processing and image recognition technologies according to the received data information and upload the data information to the attribute marking module; the attribute marking module is used to mark and verify the obtained characteristic attributes according to the received data information to complete the marking of the characteristic attributes of the target power transmission and transformation project design documents.
Claims
1. A method for marking characteristic attributes of power transmission and transformation engineering design documents, comprising the following steps: S1. Construct a characteristic attribute system for power transmission and transformation engineering design documents; S2. Preprocess the target power transmission and transformation engineering design documents; S3. Based on natural language processing and image recognition technologies, extract characteristic attributes from the preprocessed target power transmission and transformation engineering design documents; S4. Mark and verify the characteristic attributes obtained in step S3 to complete the marking of the characteristic attributes of the target power transmission and transformation engineering design documents.
2. The method for marking characteristic attributes of the power transmission and transformation project design document according to claim 1, wherein The step S1 of constructing a characteristic attribute system for power transmission and transformation engineering design documents specifically includes the following steps: The constructed characteristic attribute system for power transmission and transformation engineering design documents includes text structured tags, text language semantic tags, and text service tags; Among them, the text structured tags include classification numbers, standard names, chapters, articles, paragraphs, figures, tables, and mathematical formulas; the text language semantic tags include scope, type, and content; the text service tags include project information, professional information, document type, and document content, where the project information includes project name, voltage level, and design stage, the professional information includes primary electrical, secondary electrical, and civil engineering, the document type includes drawings, specifications, calculation books, and equipment material lists, and the document content includes equipment models, technical parameters, and installation methods.
3. The method for marking characteristic attributes of the power transmission and transformation project design document according to claim 2, wherein The step S3 of extracting characteristic attributes from the preprocessed target power transmission and transformation engineering design documents based on natural language processing and image recognition technologies includes the following steps: Through a pre-trained model, predict and perform similarity scoring on the characteristic attributes corresponding to the text content of the target power transmission and transformation engineering design documents; According to the similarity scoring, divide the detail branch levels and classify the text content into the corresponding detail branch levels to form the first classification data; For each detail branch, use a hierarchical method for recursive classification to form the second classification data; According to the first classification data and the second classification data, construct a scenario tree; According to the constructed scenario tree, perform relevance analysis on the text content; According to the relevance analysis results, classify and label the text content.
4. The method for marking characteristic attributes of power transmission and transformation project design documents according to claim 3, wherein The step of predicting and performing similarity scoring on the characteristic attributes corresponding to the text content of the target power transmission and transformation engineering design documents through a pre-trained model specifically includes the following steps: Use a pre-trained model to predict and perform similarity scoring on the characteristic attributes corresponding to the text content of the power transmission and transformation engineering design documents; According to the predicted similarity scoring, retain the text content with a similarity scoring higher than the set threshold; Classify and save the content of the retained text according to the constructed characteristic attribute system of the power transmission and transformation engineering design documents.
5. The method for marking characteristic attributes of power transmission and transformation project design documents according to claim 4, characterized in that The step of dividing the detail branch levels according to the similarity scoring and classifying the text content into the corresponding detail branch levels to form the first classification data specifically includes the following steps: In the feature attribute system of each category, according to the predicted similarity score, the level of the detailed branch is divided; among them, the content with a similarity score higher than or equal to the first threshold to the feature attribute is divided into a high branch; the content with a similarity score lower than the first threshold and higher than or equal to the second threshold is divided into a medium branch; the content with a similarity score lower than the second threshold is divided into a low branch; The text content is classified into the corresponding detailed branch level to form the first classification data.
6. The method for marking characteristic attributes of power transmission and transformation project design documents according to claim 5, wherein For each of the detailed branches, a hierarchical method is used for recursive classification to form the second-level classification data, which specifically includes the following steps: For each of the detailed branches, the text content therein is extracted again; According to the extracted text content, the similarity between the text content and the feature attribute is calculated; If the similarity of some text content to the feature attribute is lower than the set threshold, then this part of the text content is classified as non-feature text content; All the non-feature text content is recursively classified using a hierarchical method to form the second-level classification data.
7. The method for marking characteristic attributes of power transmission and transformation project design documents according to claim 6, wherein The construction of the scenario tree is carried out according to the first classification data and the second classification data, which specifically includes the following steps: According to the first classification data and the second classification data, a top-down scheme is adopted to construct the scenario tree: Create the root node: Set the processing scenario as S, then the root node R is represented as R = f root (S), where f root () is a function defining the creation of the root node; Perform child node expansion: Create first-level child nodes based on the first-layer classification data; Assume there are n categories in the first-layer classification data, denoted as C1, C2,..., C n ; For each category C i , create the corresponding first-level child node N i as N i = f sub (R, C i ), where f sub () is a function defining child node expansion; The first-level child nodes represent different given characteristic attribute categories; Add branch nodes: Set the first-level child node N i There are m non-characteristic attribute categories, denoted as D i1 , D i2 ,..., D im ; For each first-level child node N i Add detailed branch nodes B ij For B ij = f branch (N i , D ij ), where f branch () is the function for adding branch nodes; Branch nodes are used to reflect non-characteristic attribute content at different depth levels; The construction of the scenario tree is completed through the creation of the root node, the expansion of the child nodes, and the addition of the branch nodes.
8. The method for marking characteristic attributes of power transmission and transformation project design documents according to claim 7, wherein The relevance analysis of the text content is carried out according to the constructed scenario tree, which specifically includes the following steps: The node A in the scenario tree is represented by the vector a, and the node B is represented by the vector b; the following formula is used to calculate the content relevance evaluation value of the node A and the node B: In the formula, MSS(A,B) is the content relevance evaluation value of the node A and the node B; a*b represents the dot product of the vector a and the vector b; γ is the set feature correction factor; ||a|| represents the modulus of the vector a; ||b|| represents the modulus of the vector b; δ is the set balance coefficient; len(a) is the text length corresponding to the vector a; len(b) is the text length corresponding to the vector b; According to the obtained content relevance evaluation value of the node A and the node B, an elimination operation is carried out: if the content relevance evaluation value is greater than the set threshold, it is determined that the content of the node A and the node B is repeated, and one of the nodes is eliminated to improve the accuracy and refinement of the content of the scenario tree; The following formula is used to calculate the content occupancy value of the node A: where IV(A) is the content occupancy value of node A; ω is the set weight value; n is the number of all nodes pointing to node A; IV(T i ) is the content occupancy value of the i-th node T i pointing to node A; γ is the correlation stability coefficient between nodes; C(T i ) is the out-link number of the i-th node pointing to node A; freq(A,T i ) is the occurrence frequency between node A and node T i in the power grid service association scenario; freq(T j ,T i ) is the occurrence frequency between node T j and node T i in the power grid service association scenario; The following formula is used to calculate the weight value of the node: In the formula, αregion is the set correction coefficient; L is the level of the detailed branch; βdevice is the set weight adjustment factor; D is the depth layer information; uniqueness(L,D) is the uniqueness index between the level of the detailed branch and the depth layer information; Total uniqueness is the sum of the uniqueness indexes of all node text contents; The following steps are used to calculate the uniqueness index: The text content of each node is tokenized to obtain a keyword set; for the node text content S, the corresponding keyword set is Ks; The keyword scarcity coefficient sc(k) is calculated using the following formula: where N is the total number of keywords in the scenario tree text library; n(k) is the number of occurrences of keyword k in the text library; After vectorizing the text, the similarity sim(S, T) between the node text content S and other text content T is calculated using the following formula: where Kt is the set of keywords corresponding to other text content T; The uniqueness index uniqueness(S, T) between the node text content S and other text content T is calculated using the following formula: where m is the total number of texts in the scenario tree.
9. The method for marking characteristic attributes of power transmission and transformation project design documents according to claim 8, characterized in that According to the relevance analysis results, the text content is classified and labeled, specifically including the following steps: Using a pre-trained machine learning model to automatically label each feature attribute; The described machine learning model includes several layers of neurons, where the forward propagation process of the neurons in the l-th layer is expressed as h l = δ(W l h l-1 + b l ), where h l is the output vector of the hidden layer of the neurons in the l-th layer, δ() is the set activation function, W l is the weight matrix of the hidden layer of the neurons in the l-th layer, and b l is the bias vector of the hidden layer of the neurons in the l-th layer.
10. A system for implementing the method for marking characteristic attributes of the power transmission and transformation project design document according to any one of claims 1 to 9, characterized in that Including a system construction module, a file preprocessing module, an attribute extraction module, and an attribute labeling module; the system construction module, the file preprocessing module, the attribute extraction module, and the attribute labeling module are connected in series in sequence; the system construction module is used to construct a feature attribute system for the power transmission and transformation project design document and upload the data information to the file preprocessing module; The file preprocessing module is used to preprocess the target power transmission and transformation project design document according to the received data information and upload the data information to the attribute extraction module; The attribute extraction module is used to extract feature attributes from the preprocessed target power transmission and transformation project design document based on natural language processing and image recognition technologies according to the received data information and upload the data information to the attribute labeling module; The attribute labeling module is used to label and verify the obtained feature attributes according to the received data information to complete the feature attribute labeling of the target power transmission and transformation project design document.