Intelligent repeated detection method and system for power grid project report
By using multimodal information fusion and deep attention coding technology, the problem of low efficiency in duplicate detection in power grid engineering reports has been solved, achieving high-precision and high-robust duplicate detection and saving review time.
Patent Information
- Application Number
- CN202510941460.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional manual plagiarism checking and review methods are inefficient in power grid engineering reports, lack cross-modal collaboration, and result in long review cycles.
By employing multimodal information fusion technology, a multimodal unit set of power grid engineering reports is constructed through deep attention encoding and cross-modal reasoning. This allows for refined processing of text, tables, and images, and the use of conditional generative adversarial networks for duplicate detection.
It achieves high-precision and robust repetitive detection of implicit synonym substitution, numerical variants and graphical fine-tuning scenarios in power grid engineering reports, saving review time and solving the problem of insufficient cross-modal collaboration.
Smart Images

Figure CN120950984A_ABST
Abstract
Description
Technical Field
[0001] As the State Grid Corporation of China accelerates its intelligent and high-quality development, the number of feasibility study reports and project proposals for power grid engineering projects has exploded. According to statistics from State Grid Hunan Electric Power Co., Ltd., the number of feasibility study reports submitted in the province in 2025 increased by more than 40% compared with the previous four years. Each report is more than 100 pages long, containing thousands of text paragraphs, dozens of tables, and multiple flowcharts / wiring diagrams. This poses a serious challenge to traditional manual plagiarism checking and review, such as long review time and insufficient cross-modal collaboration. Background Technology
[0002] In view of the above problems, the purpose of this invention is to provide an intelligent duplicate detection method and system for power grid engineering reports, which can quickly detect and identify duplicates in power grid engineering reports, thereby saving review time and solving the problem of insufficient cross-modal collaboration.
[0003] The first aspect of this invention provides an intelligent duplicate detection method for power grid engineering reports, comprising:
[0004] Obtain multidimensional data from power grid engineering reports;
[0005] The multidimensional data in the power grid engineering report is preprocessed to obtain a set of text regions in the power grid engineering report;
[0006] Based on the set of text regions in the power grid engineering report, construct a multimodal unit set for the power grid engineering report;
[0007] The multimodal unit set of the power grid engineering report is optimized to obtain a refined unit set of the power grid engineering report;
[0008] The text unit features in the refined unit set of the power grid engineering report are subjected to bidirectional / bi-order attention encoding, and the image and table unit features in the refined unit set of the power grid engineering report are embedded to obtain the multimodal vector of the power grid engineering report.
[0009] The multimodal vectors in the power grid engineering report are then subjected to final cross-modal embedding to obtain a unified and fused semantic vector.
[0010] The unified and fused semantic vector is input into the preset plagiarism detection model to obtain a similarity score.
[0011] Different duplicate alerts are triggered based on whether the duplicate similarity score falls into different preset similarity score ranges.
[0012] In this solution, the step of constructing a multimodal unit set of the power grid engineering report based on the text region set of the power grid engineering report specifically includes:
[0013] Extract text region set R text The initial text unit, initial table unit, and initial image unit in the text;
[0014] After merging and semantically reconstructing the initial text units, cleaning and placeholder processing are performed to obtain the processed text units.
[0015] The power grid table in the initial table cell is adjusted and the values are standardized to obtain the processed table cell.
[0016] The primitives and edges in the initial image unit are detected to obtain the processed image unit graph;
[0017] By combining text units, table units, and graph units, we obtain the multimodal unit set U of the power grid engineering report, and its formula is:
[0018] in The unique identifier is represented by type∈{text,table,graph}, label identifies the chapter or business semantics, ptr identifies the original text location pointer, and n represents the nth text region.
[0019] In this scheme, the formula for merging and semantically reconstructing the initial text units is as follows:
[0020] Where T k This represents the text unit resulting from the merging of two initial text units, where T i T j Let represent the initial text units of the i-th text region and the j-th text region. This represents a string concatenation operation. Where R i R represents the i-th text region. j τ represents the j-th text region; τ represents the merging threshold.
[0021] In this solution, the formula for cleaning and placeholder processing is: T' k =Clean(T k =RegexReplace(T) k ,S,__VAR__), where S represents the noise and variable field pattern set, __VAR__ represents the uniform placeholder symbol, and T' k Represents text unit T k Text cells after cleaning and placeholder processing.
[0022] In this scheme, the step of performing bidirectional / bi-order attention encoding on the text unit features in the refined unit set of the power grid engineering report specifically includes:
[0023] Extracting unit u' from the refined unit set i The cleaned character sequence T' i ;
[0024] Character sequence T' i Divided into a list of tokens, Where L i =|T′ i |;
[0025] The token list is embedded and mapped using a preset training term, and the formula is as follows: e I,k =Embed(t I,k ), t I,k Let L be the k-th token of the i-th unit. i This represents the number of tokens in that unit. Embedding layer d = 768;
[0026] Self-attention is applied to all token vectors within the same paragraph to obtain local feature vectors. Its formula is Where Q = W Q E,K=W K E,V = W V E represents the linear mapping of query, key, and value, respectively; Represented as learnable weights; d k This represents the attention head dimension; softmax(·) represents the normalization function.
[0027] For multiple paragraphs within the same chapter, cross-paragraph attention is introduced to obtain global feature vectors. The formula is: in in Indicates cross-paragraph mapping, d s Represents the global attention dimension;
[0028] The local feature vector and the global feature vector are concatenated to obtain the text vector H. text (u i ′), its formula is: Where [·;·] denotes vector concatenation, Pool(.) represents average pooling operation, and the output is...
[0029] In this solution, the step of embedding image and table unit features from the refined unit set of the power grid engineering report specifically includes:
[0030] Extracting unit u from the refined unit set j 'Image data I j ;
[0031] Image data I j Cross-modal alignment is performed using a pre-defined visual encoder to obtain the image vector H. vis (u j ′), its formula is: H vis (u j ′)=QwenVL vis (I j ),in This represents a visual feature mapping, d' = 768;
[0032] Extracting unit u' from the refined unit set k Each cell
[0033] Based on the preset dual-branch embedding and fusion, the table vector H is obtained according to cell m. tab (u' k The formula is: in
[0034] Embed θ This represents the text embedding fine-tuned on the corpus of power grid cost and parameter tables, where NormNum(v,u) represents the numerical normalization function. Mapped to a real number vector.
[0035] In this scheme, the step of performing final cross-modal embedding on the multimodal vectors of the power grid engineering report to obtain a unified fused semantic vector specifically includes:
[0036] The vector is first linearly projected onto the model input dimension d' = 1024, and type embedding is added to obtain the projected vector Z. i Its formula is: Z i =W proj H i +E type (t i ),in Original feature vector, d = 768, This represents the learnable projection matrix. Indicates type embedding, distinguishing between text, tables, and images.
[0037] Traverse all vectors to obtain the projection sequence Z = [Z1, ..., Zn]. N ];
[0038] The projected sequences are fused and finally embedded across modalities to obtain a unified fused semantic vector H. uni (u).
[0039] A second aspect of the present invention provides an intelligent duplicate detection system for power grid engineering reports, comprising a memory and a processor. The memory stores a program for an intelligent duplicate detection method for power grid engineering reports. When the processor executes the program for the intelligent duplicate detection method for power grid engineering reports, it performs the following steps:
[0040] Obtain multidimensional data from power grid engineering reports;
[0041] The multidimensional data in the power grid engineering report is preprocessed to obtain a set of text regions in the power grid engineering report;
[0042] Based on the set of text regions in the power grid engineering report, construct a multimodal unit set for the power grid engineering report;
[0043] The multimodal unit set of the power grid engineering report is optimized to obtain a refined unit set of the power grid engineering report;
[0044] The text unit features in the refined unit set of the power grid engineering report are subjected to bidirectional / bi-order attention encoding, and the image and table unit features in the refined unit set of the power grid engineering report are embedded to obtain the multimodal vector of the power grid engineering report.
[0045] The multimodal vectors in the power grid engineering report are then subjected to final cross-modal embedding to obtain a unified and fused semantic vector.
[0046] The unified and fused semantic vector is input into the preset plagiarism detection model to obtain a similarity score.
[0047] Different duplicate alerts are triggered based on whether the duplicate similarity score falls into different preset similarity score ranges.
[0048] In this solution, the step of constructing a multimodal unit set of the power grid engineering report based on the text region set of the power grid engineering report specifically includes:
[0049] Extract text region set R text The initial text unit, initial table unit, and initial image unit in the text;
[0050] After merging and semantically reconstructing the initial text units, cleaning and placeholder processing are performed to obtain the processed text units.
[0051] The power grid table in the initial table cell is adjusted and the values are standardized to obtain the processed table cell.
[0052] The primitives and edges in the initial image unit are detected to obtain the processed image unit graph;
[0053] By combining text units, table units, and graph units, we obtain the multimodal unit set U of the power grid engineering report, and its formula is:
[0054] in The unique identifier is represented by type∈{text,table,graph}, label identifies the chapter or business semantics, ptr identifies the original text location pointer, and n represents the nth text region.
[0055] In this scheme, the formula for merging and semantically reconstructing the initial text units is as follows:
[0056] Where T k This represents the text unit resulting from the merging of two initial text units, where T i T j Let represent the initial text units of the i-th text region and the j-th text region. This represents a string concatenation operation. Where R i R represents the i-th text region. j τ represents the j-th text region; τ represents the merging threshold.
[0057] This invention discloses an intelligent duplicate detection method and system for power grid engineering reports. By organically integrating multimodal information such as text, tables, and images in power grid engineering reports, and based on deep attention encoding, large-model cross-modal reasoning, and conditional generative adversarial networks, it achieves high-precision and robust duplicate detection of implicit synonym substitution, numerical variants, and graphic fine-tuning scenarios in power grid engineering documents, saving review time and solving the problem of insufficient cross-modal collaboration. Attached Figure Description
[0058] Figure 1 A flowchart of an intelligent duplicate detection method for power grid engineering reports according to the present invention is shown;
[0059] Figure 2 A block diagram of an intelligent duplicate detection system for power grid engineering reports according to the present invention is shown. Detailed Implementation
[0060] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0061] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0062] Figure 1 A flowchart of an intelligent duplicate detection method for power grid engineering reports according to the present invention is shown.
[0063] like Figure 1 As shown, this invention discloses an intelligent duplicate detection method for power grid engineering reports, comprising:
[0064] S101, Obtain multi-dimensional data from the power grid engineering report;
[0065] Specifically, the multidimensional data in the power grid engineering report includes text, tables, and image data;
[0066] S102, preprocess the multidimensional data in the power grid engineering report to obtain the text area set of the power grid engineering report;
[0067] Specifically, the open-source model LayoutLMv3 is used to recognize the layout of the document page and output a set of text regions R. text R text ={R i |R i =(x i ,y i ,w i ,h i )}, where R i Represents the i-th text block region, (x i ,y i () represents the coordinates (in pixels) of the top-left corner of the text block, (w i ,h i ) represents the width and height (in pixels) of the text block.
[0068] Furthermore, a fine-tuning corpus was constructed for power grid industry terminology, such as transformer capacity, line load, and reactive power compensation, and LayoutLMv3 was subjected to domain-adaptive fine-tuning using the following loss function: in This represents the cross-entropy loss, used for basic category (body text / title / caption) classification; Indicates focusing loss, used to enhance the identification of low-frequency power grid terms; λfoc >1 indicates a focus loss weight, used to adjust the importance of grid terms; ), where y c For the true category one-hot, p c Predict probabilities for the model; Where p t The probability of correctly predicting the class is given by α and γ, which are hyperparameters.
[0069] S103, Based on the set of text regions in the power grid engineering report, construct a set of multimodal units for the power grid engineering report;
[0070] According to an embodiment of the present invention, the step of constructing a multimodal unit set of a power grid engineering report based on the text region set of the power grid engineering report specifically includes:
[0071] Extract text region set R text The initial text unit, initial table unit, and initial image unit in the text;
[0072] After merging and semantically reconstructing the initial text units, cleaning and placeholder processing are performed to obtain the processed text units.
[0073] The formula for cleaning and placeholder processing is: T' k =Clean(T k =RegexReplace(T) k ,S,__VAR__), where S represents the noise and variable field pattern set, __VAR__ represents the uniform placeholder symbol, and T' k Represents text unit T k Text cells after cleaning and placeholder processing.
[0074] The formula for merging and semantically reconstructing the initial text units is as follows:
[0075] Where T k This represents the text unit resulting from the merging of two initial text units, where T i T j Let represent the initial text units of the i-th text region and the j-th text region. This represents a string concatenation operation. Where R i R represents the i-th text region. j τ represents the j-th text region; τ represents the merging threshold.
[0076] Specifically, to meet business requirements, merging the two adjacent descriptions "220kV main transformer expanded to 315MVA" and "total capacity 630MVA" ensures the continuity of paragraphs with the same meaning. This involves merging text blocks that are spatially adjacent and have high layout relevance based on similarity.
[0077] Furthermore, to enable the model to better distinguish key sections of the power grid, the Transformer-based merging module was fine-tuned, and the loss function was defined as: Where e i ,e j Represents text block T i ,T j The embedding vector ‖·‖ represents the Euclidean norm; e i =Embed(T i ), generated by a pre-trained Transformer; IoU(R i ,R j As a weight, a higher IoU indicates that it should be merged.
[0078] The power grid table in the initial table cell is adjusted and the values are standardized to obtain the processed table cell.
[0079] Specifically, TabStruct is used for table inspection, outputting a cell matrix M, where M = [m ij ],m ij =(r i ,c j ,content ij ), where r i c j Indicates row and column index, content ij This represents text or numerical values in the cell; then, for the capacity / model fields in the "Main Equipment List" and "Cost Estimation Table," a dedicated fine-tuning set is constructed, using weighted loss, with the following formula:
[0080] in Let S represent the cross-entropy loss of the table classification, and α represent the set of industry-sensitive units. spec >1 indicates a sensitive weight, Pred(m ij Truth(m) ij This represents the embedding vector prediction and the actual annotation, where S contains key fields such as "transformer capacity" and "line length"; then, the units and precision of the numerical table cells are standardized, and the formula is content′. ij =NormalizeNum(content ij ,U), where U represents the standardization target unit.
[0081] The primitives and edges in the initial image unit are detected to obtain the processed image unit graph;
[0082] Specifically, using a pre-defined visual relationship detection model, the node set V and edge set E are identified to construct an image unit graph representation, graph = (V, E), E = {(v...} i ,v j )∣conf(v i ,v j )>δ}, where δ=0.7 represents the confidence threshold, and conf(·) represents the model output connection confidence.
[0083] By combining text units, table units, and graph units, we obtain the multimodal unit set U of the power grid engineering report, and its formula is:
[0084] in The unique identifier is represented by type∈{text,table,graph}, label identifies the chapter or business semantics, ptr identifies the original text location pointer, and n represents the nth text region.
[0085] S104, optimize the multimodal unit set of the power grid engineering report to obtain a refined unit set of the power grid engineering report;
[0086] Specifically, this step proposes a hybrid segmentation method based on document spatial layout and semantic thematic coherence for the refined division of logical units in power grid engineering project reports. This supports subsequent accurate comparison and traceability analysis across document units, as chapter titles in power grid feasibility study reports often include "investment estimate".
[0087] Accurately distinguishing between headings and body text using key terms such as "project progress" is a prerequisite for subsequent aggregation by business module. For each unit... Extract the following spatial layout features to construct a feature vector. Where x i ,y i Representation unit u i Top left corner coordinates (pixels), w i ,h i Indicates the width and height (in pixels) of a unit; t i ∈{text,table,graph} represents the cell type; Indicates a title instruction (1 = title); Indicates paragraph indentation level; The figure caption is represented by 1 (Figure Caption); the feature vector is optimized by introducing a spatial classification network architecture and fine-tuning, and its loss function is: in y represents the model's predicted probability distribution for the three classes; i ∈{e1,e2,e3} represents the true category one-hot label; Represents cross-entropy loss; This represents the focusing loss, used to enhance title detection; α and γ are hyperparameters; λ hd =3 indicates the title loss amplification factor.
[0088] Furthermore, by increasing the clustering weights of business themes such as "investment estimation" and "construction organization," it is ensured that related units under different modalities are accurately classified into the same theme, and for each unit u i Different encoders are invoked according to their type, and fine-tuned on the power grid scenario corpus, with the text unit set to E. text (u i )=TransEnc φ (T′ i ), TransEnc φ The parameter φ represents the Transformer encoder used for fine-tuning in power grid-related sections such as "substation expansion" and "transmission line renovation"; T i 'Represents unit u i After cleaning the text string, fine-tune the target. Where P represents a set of paragraphs on the same business topic; Set the table cell to E table (u i The formula is: in Representation unit u i The set of all cells in the table; Embed θ This refers to the embedding network fine-tuned on labeled tables such as "Cost Estimation Table" and "Equipment Parameter Table"; the image unit is set to E. graph (u i The formula is: E graph (u i ) = GraphEnc ψ (V i E i ), where V i E i Represents the set of nodes and edges in a flowchart; GraphEnc ψ This represents a graph neural network encoder that fine-tunes diagrams such as "construction to acceptance" and "first-time wiring," outputting a set of refined units based on the aforementioned operations. id i 'A unique identifier after refinement; type' i 'Refined type; label' i 'Business themes after refinement; ptr' i = (page number, (x i ,y i ,w i ,h i The refined positioning pointer.
[0089] S105, perform bidirectional / bi-order attention encoding on the text unit features in the refined unit set of the power grid engineering report, and embed the image and table unit features in the refined unit set of the power grid engineering report to obtain the multimodal vector of the power grid engineering report;
[0090] The step of performing bidirectional / bi-order attention encoding on the text unit features in the refined unit set of the power grid engineering report specifically includes:
[0091] Extracting unit u' from the refined unit set i The cleaned character sequence T' i ;
[0092] Character sequence T' i Divided into a list of tokens, Where L i =|T′ i |;
[0093] The token list is embedded and mapped using a preset training term, and the formula is as follows: e I,k =Embed(t I,k ), t I,k Let L be the k-th token of the i-th unit. i This represents the number of tokens in that unit. Embedding layer d = 768;
[0094] Self-attention is applied to all token vectors within the same paragraph to obtain local feature vectors. Its formula is Where Q = W Q E,K=W K E,V = W V E represents the linear mapping of query, key, and value, respectively; Represented as learnable weights; d k This represents the attention head dimension; softmax(·) represents the normalization function.
[0095] For multiple paragraphs within the same chapter, cross-paragraph attention is introduced to obtain global feature vectors. The formula is: in in Indicates cross-paragraph mapping, d s Represents the global attention dimension;
[0096] The local feature vector and the global feature vector are concatenated to obtain the text vector H. text (u i ′), its formula is: Where [·; ·] denotes vector concatenation, Pool(·) denotes average pooling operation, and the output...
[0097] The step of embedding image and table unit features from the refined unit set of the power grid engineering report specifically includes: extracting unit u from the refined unit set. j 'Image data I j ;
[0098] Image data I j Cross-modal alignment is performed using a pre-defined visual encoder to obtain the image vector H. vis (u j ′), its formula is: H vis (u j ′)=QwenVL vis (I j ),in This represents a visual feature mapping, d' = 768;
[0099] Extracting unit u' from the refined unit set k Each cell
[0100] Based on the preset dual-branch embedding and fusion, the table vector H is obtained according to cell m. tab (u' k The formula is: in
[0101] Embed θ This represents the text embedding fine-tuned on the corpus of power grid cost and parameter tables, where NormNum(v,u) represents the numerical normalization function. Mapped to a real number vector.
[0102] Specifically, text vectors, table vectors, and image vectors are combined to obtain a multimodal vector set.
[0103] S106, Perform final cross-modal embedding on the multimodal vectors of the power grid engineering report to obtain a unified fused semantic vector;
[0104] The step of performing final cross-modal embedding of the multimodal vectors from the power grid engineering report to obtain a unified fused semantic vector specifically includes:
[0105] The vector is first linearly projected onto the model input dimension d' = 1024, and type embedding is added to obtain the projected vector Z. i Its formula is: Z i =W proj H i +E type (t i ),in Original feature vector, d = 768, This represents the learnable projection matrix. Indicates type embedding, distinguishing between text, tables, and images.
[0106] Traverse all vectors to obtain the projection sequence Z = [Z1, ..., Zn]. N ];
[0107] The projected sequences are fused and finally embedded across modalities to obtain a unified fused semantic vector H. uni (u).
[0108] Specifically, units of different types within the same business section are concatenated in the original text order to form a cross-modal fusion sequence, X. c =[H text (u′1),…,H text (u′ m ),H tab (u′ m+1 ),…,H tab (u′ m+t ),H vis (u′ m+t+1 ),...], where X c This represents the fusion sequence of the c-th business chapter, where m represents the number of text units, t represents the number of table units, and subsequent indexing extends to the image units.
[0109] Furthermore, after fusing the projection sequences, the fused output F is obtained. fused Its formula is F fused =[head1; ...;head 16 W O ,in Where k = 1…16, This indicates the output fusion matrix. This represents the parameters for each head.
[0110] Furthermore, the fused output F is obtained. fused After that, F fused Based on this, 12 layers of TransformerBlock are stacked, each containing self-attention and feedforward networks: F (0) =F fused , This represents the features of the i-th vector after L=12 layers of inference. Then, a unified high-dimensional vector is output for any unit u, and after fusion, they share the same semantic space, resulting in a unified fused semantic vector. The formula is:
[0111] S107, Input the unified and fused semantic vector into the preset plagiarism detection model to obtain the duplication similarity score;
[0112] Specifically, the similarity score is set as S. ij If we designate the two reports as A and B, then the set of logical units corresponding to report A is: The logical unit set of report B is as follows b}, then the unified fusion vector corresponding to each unit u is Construct the discriminant condition vector The plagiarism detection model outputs a similarity score. Where S ij Let represent the repetition similarity score of unit pair (i,j), and D(·) represent the trained discriminator function. This represents vector concatenation, where a and b represent the total number of logical units in the corresponding document.
[0113] S108: Based on the repetition similarity score falling into different preset similarity score ranges, different repetition warnings are triggered.
[0114] Specifically, for example, the preset similarity score range is divided into greater than or equal to 0.98, less than 0.98 but greater than or equal to 0.95, and less than 0.95 but greater than or equal to 0.90, and when S ij When the value is ≥0.98, a high-risk repetitive warning message is triggered; when the value is ≤0.95, a warning message is triggered. ij When S < 0.98, a moderate repetitive warning message is triggered; when 0.90 ≤ S ijWhen the value is less than 0.95, a low-level repetitive warning message is triggered. Furthermore, different colors are set to distinguish different levels of warning messages, such as red for high-risk repetitive warning messages, orange for medium-level repetitive warning messages, and yellow for low-level repetitive warning messages.
[0115] According to an embodiment of the present invention, the method further includes: after triggering a moderate or high-risk repeated warning message, calculating a structured similarity index for the table unit and the image unit respectively to quickly locate the differences, and setting the similarity of the table unit to S. tab Its formula is Where e a ,e b This represents cell text or numerical embedding; cos(·) represents cosine similarity; when the similarity of table cells is greater than or equal to the set table cell similarity threshold, the similar parts in the corresponding table cells will be marked; set the image similarity to S. graph Its formula is
[0116] According to an embodiment of the present invention, it further includes: based on the Levenshtein distance d between the initial text units in Report A and Report B. lev Its formula is in and Let represent the initial text units of the i-th text region in report A and the j-th text region in report B, respectively; based on a preset distance threshold, d lev The system performs a judgment to automatically highlight added, deleted, or modified characters in text paragraphs.
[0117] According to an embodiment of the present invention, the method further includes: before performing a duplication similarity score on reports A and B, extracting the names of reports A and B; if the extracted names of reports A and B are the same, then reports A and B are determined to be duplicate reports and the duplication similarity score is terminated; if the extracted names of reports A and B are different, then the power grid equipment in reports A and B is further extracted and set as the implementation object; if either of the two reports has an implementation object and the implementation content is repeated, then the two reports are determined to be duplicate reports and the duplication similarity score is terminated; if neither of the two reports has an implementation object and the implementation content is repeated, then a duplication similarity score is performed on the two reports.
[0118] It should be noted that the implementation objects refer to the names of power grid equipment such as substations, transformers, lines, and branch lines. The determination of duplicate implementation objects is based on the subject, number, and segment number of the implementation object. Specifically: when two implementation objects have the same subject and no number or segment number, the corresponding two implementation objects are directly determined to be duplicates; when two implementation objects have the same subject and the implementation object contains a number and segment number, and if the number and segment number have an inclusion relationship, the corresponding two implementation objects are determined to be duplicates; for example, "10kV Sheye Line Badou 2 Substation Branch Line P11-P12", The "10kV Sheye Line Badou 2 Substation Branch Line" is the main body of the implementation. If the main body is consistent and has numbering or segment information, the segment information "P11-P12" is compared. For example, P1-P12, P5-P6, and P5-P6 are included in P1-P12. The determination of duplicate implementation content is based on the same operation action for the corresponding content of the same type. For example, if Project A changes conductor type A and Project B changes conductor type B, conductor type A and conductor type B are of the same type, and the operation action is replacement, then the corresponding implementation content is duplicated. If Project A changes the conductor and Project B changes the transformer, then the corresponding implementation content is considered non-duplicate.
[0119] According to an embodiment of the present invention, the method further includes: when a placeholder for consecutive X appears in the technical_solution or construction_contents of any implementation object in the comparison report, the number of occurrences is recorded; if the number of occurrences is greater than a preset first number threshold, the corresponding report is set to be duplicated with an existing report, for example, the preset first number threshold is 2.
[0120] Figure 2 A block diagram of an intelligent duplicate detection system for power grid engineering reports according to the present invention is shown.
[0121] like Figure 2 As shown, a second aspect of the present invention provides an intelligent duplicate detection system 2 for power grid engineering reports, including a memory 21 and a processor 22. The memory stores a program for an intelligent duplicate detection method for power grid engineering reports. When the processor executes the program for the intelligent duplicate detection method for power grid engineering reports, it performs the following steps:
[0122] Obtain multidimensional data from power grid engineering reports;
[0123] The multidimensional data in the power grid engineering report is preprocessed to obtain a set of text regions in the power grid engineering report;
[0124] Based on the set of text regions in the power grid engineering report, construct a multimodal unit set for the power grid engineering report;
[0125] The multimodal unit set of the power grid engineering report is optimized to obtain a refined unit set of the power grid engineering report;
[0126] The text unit features in the refined unit set of the power grid engineering report are subjected to bidirectional / bi-order attention encoding, and the image and table unit features in the refined unit set of the power grid engineering report are embedded to obtain the multimodal vector of the power grid engineering report.
[0127] The multimodal vectors in the power grid engineering report are then subjected to final cross-modal embedding to obtain a unified and fused semantic vector.
[0128] The unified and fused semantic vector is input into the preset plagiarism detection model to obtain a similarity score.
[0129] Different duplicate alerts are triggered based on whether the duplicate similarity score falls into different preset similarity score ranges.
[0130] In this solution, the step of constructing a multimodal unit set of the power grid engineering report based on the text region set of the power grid engineering report specifically includes:
[0131] Extract text region set R text The initial text unit, initial table unit, and initial image unit in the text;
[0132] After merging and semantically reconstructing the initial text units, cleaning and placeholder processing are performed to obtain the processed text units.
[0133] The power grid table in the initial table cell is adjusted and the values are standardized to obtain the processed table cell.
[0134] The primitives and edges in the initial image unit are detected to obtain the processed image unit graph;
[0135] By combining text units, table units, and graph units, we obtain the multimodal unit set U of the power grid engineering report, and its formula is:
[0136] in The unique identifier is represented by type∈{text,table,graph}, label identifies the chapter or business semantics, ptr identifies the original text location pointer, and n represents the nth text region.
[0137] In this scheme, the formula for merging and semantically reconstructing the initial text units is as follows:
[0138] Where T kThis represents the text unit resulting from the merging of two initial text units, where T i T j Let represent the initial text units of the i-th text region and the j-th text region. This represents a string concatenation operation. Where R i R represents the i-th text region. j τ represents the j-th text region; τ represents the merging threshold.
[0139] This invention discloses an intelligent duplicate detection method and system for power grid engineering reports. By organically integrating multimodal information such as text, tables, and images in power grid engineering reports, and based on deep attention encoding, large-model cross-modal reasoning, and conditional generative adversarial networks, it achieves high-precision and robust duplicate detection of implicit synonym substitution, numerical variants, and graphic fine-tuning scenarios in power grid engineering documents, saving review time and solving the problem of insufficient cross-modal collaboration.
[0140] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0141] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0142] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0143] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0144] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
Claims
1. A method for intelligent duplicate detection in power grid engineering reports, characterized in that, include: Obtain multidimensional data from power grid engineering reports; The multidimensional data in the power grid engineering report is preprocessed to obtain a set of text regions in the power grid engineering report; Based on the set of text regions in the power grid engineering report, construct a multimodal unit set for the power grid engineering report; The multimodal unit set of the power grid engineering report is optimized to obtain a refined unit set of the power grid engineering report; The text unit features in the refined unit set of the power grid engineering report are subjected to bidirectional / bi-order attention encoding, and the image and table unit features in the refined unit set of the power grid engineering report are embedded to obtain the multimodal vector of the power grid engineering report. The multimodal vectors in the power grid engineering report are then subjected to final cross-modal embedding to obtain a unified and fused semantic vector. The unified and fused semantic vector is input into the preset plagiarism detection model to obtain a similarity score. Different duplicate alerts are triggered based on whether the duplicate similarity score falls into different preset similarity score ranges.
2. The intelligent duplicate detection method for power grid engineering reports according to claim 1, characterized in that, The step of constructing a multimodal unit set of the power grid engineering report based on the text region set of the power grid engineering report specifically includes: Extract text region set R text The initial text unit, initial table unit, and initial image unit in the text; After merging and semantically reconstructing the initial text units, cleaning and placeholder processing are performed to obtain the processed text units. The power grid table in the initial table cell is adjusted and the values are standardized to obtain the processed table cell. The primitives and edges in the initial image unit are detected to obtain the processed image unit graph; By combining text units, table units, and graph units, we obtain the multimodal unit set U of the power grid engineering report, and its formula is: U={u n |u n =(id) n ,type n ,label n ,content n ,ptr n )}, where id n The unique identifier is represented by type∈{text,table,graph}, label identifies the chapter or business semantics, ptr identifies the original text location pointer, and n represents the nth text region.
3. The intelligent duplicate detection method for power grid engineering reports according to claim 2, characterized in that, The formula for merging and semantically reconstructing the initial text units is as follows: T k =T i ⊕T j ,if IoU(R i ,R j ), where T k This represents the text unit resulting from the merging of two initial text units, where T i T j Let represent the initial text units of the i-th and j-th text regions, and ⊕ represent the string concatenation operation. Where R i R represents the i-th text region. j τ represents the j-th text region; τ represents the merging threshold.
4. The intelligent duplicate detection method for power grid engineering reports according to claim 2, characterized in that, The formula for cleaning and placeholder processing is: T' k =Clean(T k =RegexReplace(T) k ,S,__VAR__), where S represents the noise and variable field pattern set, __VAR__ represents the uniform placeholder symbol, and T' k Represents text unit T k Text units after cleaning and placeholder processing.
5. The intelligent duplicate detection method for power grid engineering reports according to claim 1, characterized in that, The step of performing bidirectional / bi-order attention encoding on the text unit features in the refined unit set of the power grid engineering report specifically includes: Extracting unit u' from the refined unit set i The cleaned character sequence T' i ; Character sequence T' i Divided into a list of tokens, Where L i =|T′ i |; The token list is embedded and mapped using a preset training term, and the formula is as follows: e I,k =Embed(t I,k ), t I,k Let L be the k-th token in the i-th unit. i This represents the number of tokens in that unit. Embedding layer d = 768; Self-attention is applied to all token vectors within the same paragraph to obtain local feature vectors. Its formula is Where Q = W Q E,K=W K E,V=W V E represents the linear mapping of query, key, and value, respectively; Represented as learnable weights; d k This represents the attention head dimension; softmax(·) represents the normalization function. For multiple paragraphs within the same chapter, cross-paragraph attention is introduced to obtain global feature vectors. The formula is: in in Indicates cross-paragraph mapping, d s Represents the global attention dimension; The local feature vector and the global feature vector are concatenated to obtain the text vector H. text (u i ′), its formula is: Where [·;·] denotes vector concatenation, Pool(.) represents average pooling operation, and the output is...
6. The intelligent duplicate detection method for power grid engineering reports according to claim 1, characterized in that, The step of embedding image and table unit features from the refined unit set of the power grid engineering report specifically includes: Extracting unit u from the refined unit set j 'Image data I j ; Image data I j Cross-modal alignment is performed using a pre-defined visual encoder to obtain the image vector H. vis (u j ′), its formula is: H vis (u j ′)=QwenVL vis (I j ),in This represents a visual feature mapping, d' = 768; Extracting unit u' from the refined unit set k Each cell Based on the preset dual-branch embedding and fusion, the table vector H is obtained according to cell m. tab (u' k The formula is: in Embed θ This represents a text embedding fine-tuned on the corpus of power grid cost and parameter tables. NormNum(v,u) represents the numerical normalization function, which normalizes the values... m Mapped to a real number vector.
7. The intelligent duplicate detection method for power grid engineering reports according to claim 1, characterized in that, The step of performing final cross-modal embedding of the multimodal vectors from the power grid engineering report to obtain a unified fused semantic vector specifically includes: The vector is first linearly projected onto the model input dimension d' = 1024, and type embedding is added to obtain the projected vector Z. i Its formula is: Z i =W proj H i +E type (t i ),in Original feature vector, d = 768, This represents the learnable projection matrix. Indicates type embedding, distinguishing between text, tables, and images. Traverse all vectors to obtain the projection sequence Z = [Z1, ..., Zn]. N ]; The projected sequences are fused and finally embedded across modalities to obtain a unified fused semantic vector H. uni (u).
8. A smart duplicate detection system for power grid engineering reports, characterized in that, The system includes a memory and a processor. The memory stores a program for an intelligent duplicate detection method in power grid engineering reports. When the processor executes the program for the intelligent duplicate detection method in power grid engineering reports, it performs the following steps: Obtain multidimensional data from power grid engineering reports; The multidimensional data in the power grid engineering report is preprocessed to obtain a set of text regions in the power grid engineering report; Based on the set of text regions in the power grid engineering report, construct a multimodal unit set for the power grid engineering report; The multimodal unit set of the power grid engineering report is optimized to obtain a refined unit set of the power grid engineering report; The text unit features in the refined unit set of the power grid engineering report are subjected to bidirectional / bi-order attention encoding, and the image and table unit features in the refined unit set of the power grid engineering report are embedded to obtain the multimodal vector of the power grid engineering report. The multimodal vectors in the power grid engineering report are then subjected to final cross-modal embedding to obtain a unified and fused semantic vector. The unified and fused semantic vector is input into the preset plagiarism detection model to obtain a similarity score. Different duplicate alerts are triggered based on whether the duplicate similarity score falls into different preset similarity score ranges.
9. The intelligent duplicate detection system for power grid engineering reports according to claim 8, characterized in that, The step of constructing a multimodal unit set of the power grid engineering report based on the text region set of the power grid engineering report specifically includes: Extract text region set R text The initial text unit, initial table unit, and initial image unit in the text; After merging and semantically reconstructing the initial text units, cleaning and placeholder processing are performed to obtain the processed text units. The power grid table in the initial table cell is adjusted and the values are standardized to obtain the processed table cell. The primitives and edges in the initial image unit are detected to obtain the processed image unit graph; By combining text units, table units, and graph units, we obtain the multimodal unit set U of the power grid engineering report, and its formula is: U={u n |u n =(id) n ,type n ,label n ,content n ,ptr n )}, where id n The unique identifier is represented by type∈{text,table,graph}, label identifies the chapter or business semantics, ptr identifies the original text location pointer, and n represents the nth text region.
10. The intelligent duplicate detection system for power grid engineering reports according to claim 9, characterized in that, The formula for merging and semantically reconstructing the initial text units is as follows: T k =T i ⊕T j ,if IoU(R i ,R j ), where T k This represents the text unit resulting from the merging of two initial text units, where T i T j Let represent the initial text units of the i-th and j-th text regions, and ⊕ represent the string concatenation operation. Where R i R represents the i-th text region. j τ represents the j-th text region; τ represents the merging threshold.