Power industry literature standardization conversion method and system based on mapping rule
Through a method based on mapping rules, a standard library of power literature rules is established, and multi-dimensional depth comparison and differential clustering analysis is carried out, which solves the problem of low accuracy of standardized conversion of literature in the existing technology, and achieves a more efficient standardized conversion effect.
Patent Information
- Application Number
- CN202510096234.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
There are significant problems with the standardization accuracy of existing power literature standardization tools, and it is difficult to accurately identify the logical correlation of the document content, resulting in the adjusted content deviating from industry standards.
Using a mapping rule-based method, we interactively obtain power literature standards, establish a standard library of power literature rules, and identify differences points and generate optimal adjustment plans through multi-dimensional depth comparison and differential clustering analysis, and optimize the literature content to meet the standards.
It improves the accuracy of standardized conversion of documents, ensures that the adjusted document content is more in line with the standards of the power industry, and improves the efficiency and effectiveness of standardized conversion.
Smart Images

Figure CN120030993A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of document standardization, and in particular relates to a method and system for converting document standardization in the electric power industry based on mapping rules. Background Art
[0002] Although the existing power document standardization tools have achieved automated processing, there are significant problems with standardization accuracy. This is mainly manifested in the process of converting the target document to meet industry standards, relying only on rule matching and simple comparison, making it difficult to deeply analyze the subtle differences in the document content. For example, when dealing with chapter structure and term replacement, existing methods often cannot accurately identify the logical associations in the document content, resulting in the adjusted document content still deviating from industry standards. In addition, when the initially adjusted document does not match the standard library well, there is a lack of a mechanism for further optimization, resulting in poor document standardization results. Summary of the invention
[0003] In order to solve the technical problem that the prior art has low accuracy in document standardization conversion, the purpose of the present invention is to provide a method and system for document standardization conversion in the electric power industry based on mapping rules, so as to improve the technical effect of document standardization conversion accuracy.
[0004] The purpose of the present invention is achieved through the following technical solutions:
[0005] A method for standardizing and converting documents in the electric power industry based on mapping rules, characterized in that the method comprises the following steps:
[0006] S100, interactively obtain power document standards, determine mapping rules based on historical standard documents, and establish a power document rule standard library;
[0007] S200, comparing and matching the target document based on the mapping rules in the power document rule standard library to generate a first standard document;
[0008] S300, performing fine-grained identification on the first standard document through multi-dimensional deep comparison and difference cluster analysis to obtain a difference point adjustment set;
[0009] S400, generating multiple adjustment schemes according to the difference point adjustment set, and determining the optimal adjustment scheme through matching degree comparison, wherein when the matching degree comparison result does not meet the preset matching degree, the multiple adjustment schemes are subjected to collision optimization to determine the optimal adjustment scheme;
[0010] S500: According to the optimal adjustment scheme, the first standard document is optimized to generate a second standard document, and the second standard document is output as a target standard document.
[0011] Step S100 includes:
[0012] Receiving the pre-uploaded electric power document standard and the historical standard document, wherein the electric power document standard at least includes document format requirements, unit usage specifications and terminology standards, and the historical standard document includes title structure, chapter structure, paragraph format, unit usage, terminology usage and citation format;
[0013] Matching the historical standard document with the power document standard to generate the mapping rule;
[0014] The electric power document rule standard library is established based on the mapping rule, and the electric power document rule standard library includes a format specification unit, a unit specification unit, a term specification unit, a chapter structure unit and a reference specification unit.
[0015] Step S200 includes:
[0016] Parsing the target document to determine the target document constituent elements, wherein the target document constituent elements include title, chapter structure, paragraph format, unit usage, and term usage;
[0017] Matching and converting the target document constituent elements based on the document rule standard library to generate multiple element conversion results;
[0018] The multiple element conversion results are integrated to obtain the first standard document.
[0019] Step S300 includes:
[0020] Extracting features from the first standard document to determine a target document feature set;
[0021] Performing the difference clustering analysis on the target document feature set in combination with the electric power document rule standard library to determine a feature difference data set;
[0022] Performing feature hierarchical analysis on the feature difference data set, determining the difference point location data and performing multi-dimensional aggregation analysis, analyzing the correlation between the difference points, and generating a multi-dimensional correlation data set;
[0023] Based on the multi-dimensional associated data set, the first standard document is contextually reconstructed to determine the difference point adjustment set.
[0024] Generating multiple adjustment schemes according to the difference point adjustment set, and determining the optimal adjustment scheme by matching degree comparison, the method further includes:
[0025] According to the difference point adjustment set, a plurality of adjustment schemes are randomly selected to determine, the first standard document is adjusted based on the plurality of adjustment schemes, and the matching degree with the power document rule standard library is calculated to determine a plurality of adjustment matching degrees;
[0026] The multiple adjustment matching degrees are compared with a preset matching degree. If there is an adjustment matching degree greater than the preset matching degree, an adjustment scheme corresponding to the maximum value is selected as the optimal adjustment scheme.
[0027] When the matching degree comparison result does not meet the preset matching degree, the multiple adjustment schemes are subjected to collision optimization, and the method includes:
[0028] If there is no adjustment matching degree greater than the preset matching degree, traverse multiple adjustment schemes and randomly select two adjustment schemes for combined collision optimization to generate a first collision adjustment scheme;
[0029] Reconstructing the content of the first standard document based on the first collision adjustment scheme, determining reconstruction result data, and calculating and determining a first adjustment matching degree;
[0030] If the first adjustment matching degree is less than or equal to the preset matching degree, the combined collision optimization and content reconstruction process is repeated until an Nth adjustment matching degree greater than the preset matching degree appears, and the Nth adjustment scheme corresponding to the Nth adjustment matching degree is used as the optimal adjustment scheme, where N is an integer greater than or equal to 2.
[0031] Calculating the degree of matching with the electric power document rule standard library, the method comprises:
[0032] Construct a matching degree calculation function, the matching degree calculation function is as follows:
[0033] M=w 1 ×S c +w 2 ×C p +w 3 ×T m +w 4 ×U m +w 5 ×R f ;
[0034] Among them, M is the matching degree, which indicates the matching degree between the document and the standard library of power document rules, S c is the chapter structure similarity, which indicates the similarity between the chapter structure of the document and the standard library. p is the consistency of paragraph content, indicating the matching degree between paragraph content and standard library, T m U is the term usage matching degree, which indicates the degree of conformity between the term usage in the literature and the standard library. m is the unit usage matching degree, indicating whether the unit usage in the literature conforms to the standard library specifications. f is the consistency of the citation format, indicating the consistency of the literature citation format with the standard library, w 1 、w2 、w 3 、w 4 、w 5 is the weight coefficient of each feature, the sum is 1;
[0035] The matching degree with the electric power document rule standard library is calculated according to the matching degree calculation function to determine the adjustment matching degree.
[0036] A mapping rule-based power industry document standardization conversion system, characterized in that the system includes:
[0037] A standard library establishment module, wherein the standard library establishment module interactively acquires power document standards, determines mapping rules based on historical standard documents, and establishes a power document rule standard library;
[0038] A first standard document generation module, which compares and matches the target document based on the mapping rules in the power document rule standard library to generate a first standard document;
[0039] A fine-grained identification module, wherein the fine-grained identification module performs fine-grained identification on the first standard document through multi-dimensional depth comparison and difference clustering analysis to obtain a difference point adjustment set;
[0040] an optimal adjustment scheme determination module, wherein the optimal adjustment scheme determination module generates a plurality of adjustment schemes according to the difference point adjustment set, and determines the optimal adjustment scheme by matching degree comparison, wherein when the matching degree comparison result does not meet the preset matching degree, the plurality of adjustment schemes are subjected to collision optimization to determine the optimal adjustment scheme;
[0041] A standard document output module is configured to optimize the first standard document according to the optimal adjustment scheme, generate a second standard document, and output the second standard document as a target standard document.
[0042] The beneficial effects of the present invention are:
[0043] The present invention interactively obtains the electric power document standard, and determines the mapping rules according to the historical standard documents, and establishes the electric power document rule standard library; compares and matches the target document based on the mapping rules in the electric power document rule standard library to generate the first standard document; through multi-dimensional deep comparison and difference cluster analysis, the first standard document is finely identified to obtain the difference point adjustment set; multiple adjustment schemes are generated according to the difference point adjustment set, and the optimal adjustment scheme is determined through matching degree comparison, wherein when the matching degree comparison result does not meet the preset matching degree, the multiple adjustment schemes are collision optimized to determine the optimal adjustment scheme; according to the optimal adjustment scheme, the first standard document is optimized to generate the second standard document, and the second standard document is output as the target standard document. The present invention solves the technical problem that the prior art has low accuracy in document standardization conversion, by establishing the electric power document rule standard library, performing preliminary matching on the target document to generate the first standard document, and then identifying the details that do not meet the standard through multi-dimensional deep comparison and difference cluster analysis, generating the difference point adjustment set, generating multiple adjustment schemes for these difference points, and then performing matching degree comparison and collision optimization, and finally generating the target standard document, so as to achieve the technical effect of improving the accuracy of document standardization conversion. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A schematic diagram of a process flow of a method for converting power industry documents based on mapping rules provided in an embodiment of the present application;
[0045] Figure 2 A schematic diagram of the structure of a power industry document standardization conversion system based on mapping rules provided in an embodiment of the present application.
[0046] Explanation of the reference numerals: standard library establishment module 11, first standard document generation module 12, fine-grained identification module 13, optimal adjustment plan determination module 14, standard document output module 15. DETAILED DESCRIPTION
[0047] A method and system for converting documents in the power industry based on mapping rules aims to solve the technical problem of low accuracy in document conversion in the prior art. By establishing a standard library of power document rules, a preliminary match is made to the target document to generate a first standard document. Then, multi-dimensional deep comparison and difference clustering analysis are used to identify details that do not meet the standards, generate a difference point adjustment set, generate multiple adjustment plans for these difference points, and then perform matching comparison and collision optimization to finally generate a target standard document, thereby achieving the technical effect of improving the accuracy of document conversion.
[0048] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0049] Embodiment 1, as Figure 1 As shown, a method for converting power industry documents into standardized documents based on mapping rules comprises:
[0050] Step S100: interactively acquire power document standards, determine mapping rules based on historical standard documents, and establish a power document rule standard library.
[0051] In this embodiment, in order to perform standardized conversion of power industry documents, the pre-uploaded power document standards and historical standard documents are first received. The power document standards include at least document format requirements, unit usage specifications and terminology standards, while the historical standard documents cover title structure, chapter structure, paragraph format, unit usage, terminology usage and citation format. Subsequently, the historical standard documents are matched and analyzed with the power document standards to generate corresponding mapping rules. Based on these mapping rules, a power document rule standard library is established.
[0052] Furthermore, in the method provided in the embodiment, the electric power document standards are interactively acquired, and mapping rules are determined according to historical standard documents to establish an electric power document rule standard library, which also includes:
[0053] Receive the pre-uploaded electric power document standard and the historical standard document, wherein the electric power document standard at least includes document format requirements, unit usage specifications and terminology standards, and the historical standard document includes title structure, chapter structure, paragraph format, unit usage, terminology usage and citation format; match the historical standard document with the electric power document standard to generate the mapping rule; establish the electric power document rule standard library based on the mapping rule, and the electric power document rule standard library includes format specification unit, unit specification unit, terminology specification unit, chapter structure unit and citation specification unit.
[0054] In this embodiment, the pre-uploaded power document standards and historical standard documents are first received and parsed. Among them, the power document standards refer to the standardization requirements documents commonly used in the industry, including format requirements, unit usage specifications and terminology standards. Historical standard documents refer to document samples that have been standardized, including information such as title structure, chapter structure, paragraph format, unit usage, terminology usage and citation format. Through data parsing and preprocessing, these contents are converted into analyzable structured data forms, such as JSON or XML, for subsequent comparison and matching.
[0055] After the data parsing is completed, the rule matching method is used to compare and analyze the historical standard documents and the power document standards item by item. First, in the format matching step, the title, chapter and paragraph formats in the historical standard documents are analyzed, and these formats are compared one by one to see if they meet the format requirements in the power document standards. For example, check whether the title hierarchy is consistent and whether the arrangement of chapters meets the structural requirements in the standard. Use rule-based text analysis tools to identify format differences and mark inconsistencies. Next is unit usage matching, in which the unit symbols in the historical standard documents are compared with the unit specifications in the power document standards. For example, if "kWh" is used in the historical standard document, and the standard requires "kilowatt-hour", record this difference. Through template-based rule matching technology, replacement suggestions for different unit differences are generated to ensure the standardization and consistency of unit usage.
[0056] Then, term usage matching is performed to analyze the professional terms and technical vocabulary used in historical standard documents to determine whether they meet the terminology specifications defined in the power document standards. Using word frequency statistics and context analysis methods, the frequency of use of each term in the document is calculated and compared with the terminology recommended in the standard. For example, if the standard recommends the use of "voltmeter" and the historical document uses "voltmeter", this difference is identified and recorded.
[0057] Based on the comparison results of the above format matching, unit usage matching and term usage matching, specific mapping rules are generated. Mapping rules are conversion schemes that adjust the content of historical documents that do not meet the standards to standard forms, including how to adjust the document format, replace unit symbols and standardize the use of terminology. By comparing each item and generating mapping rules, it is ensured that the document content meets the standard requirements in terms of format, terminology and units.
[0058] Finally, based on these mapping rules, a standard library of power document rules is established. The standard library consists of format specification unit, unit specification unit, terminology specification unit, chapter structure unit and reference specification unit. Each unit stores a specific type of mapping rules. For example, the format specification unit stores the adjustment method of the title format and paragraph structure, and the terminology specification unit stores the term replacement rules.
[0059] Step S200: Compare and match the target document based on the mapping rules in the power document rule standard library to generate a first standard document.
[0060] In this embodiment, the target document is first parsed to extract the basic components of the document, which include title, chapter structure, paragraph format, unit usage, and term usage. Through parsing, the document content is decomposed into recognizable and processable structured data units. Next, based on the mapping rules in the power document rule standard library, these parsed target document component elements are matched and converted one by one. For each element, adjustments are made according to the rules in the standard library, such as modifying the chapter format, adjusting the paragraph layout, replacing unit symbols or terms, etc., so as to generate multiple standardized element conversion results, and the element conversion results are integrated to generate the first standard document.
[0061] Furthermore, in the method provided in the embodiment, the target document is compared and matched based on the mapping rules in the power document rule standard library to generate the first standard document, and further includes:
[0062] The target document is parsed to determine the constituent elements of the target document, which include title, chapter structure, paragraph format, unit usage and terminology usage; the constituent elements of the target document are matched and converted based on the document rule standard library to generate multiple element conversion results; and the multiple element conversion results are integrated to obtain the first standard document.
[0063] In this embodiment, the target document is first parsed to determine its constituent elements. By using natural language processing and text parsing technology, the target document is broken down into multiple basic elements, including titles, chapter structures, paragraph formats, unit usage, and term usage. During the parsing process, the title hierarchy and structure of the document are identified, such as the relationship between the main title and the subtitle; the chapter structure is extracted to clarify the hierarchical relationship between the chapter and the sub-chapter; the paragraph format is analyzed, including the indentation, alignment, and line spacing of the paragraph; all unit symbols and their contextual locations that appear in the text are identified, such as "kWh", "MW", etc.; and the professional terms involved in the text are extracted to generate a term list and its usage environment.
[0064] After parsing out the constituent elements of the target document, these elements are matched and converted one by one based on the mapping rules in the standard library of power document rules to generate standardized element conversion results. Specifically, the title is first converted, and its hierarchy and style are adjusted by comparing the title format requirements to generate a standardized title structure. Then, according to the chapter structure requirements in the standard library, the chapter order and hierarchy are adjusted to meet the standard specifications and generate a standardized chapter arrangement. For the paragraph format, the format conversion rules in the standard library are used to adjust the paragraph indentation, alignment, and line spacing to ensure the consistency of the paragraph format. For the unit usage part, the rules in the unit specification are used to check and replace all non-standard unit symbols in the text to generate standardized unit representations. In terms of terminology usage, refer to the terminology comparison table in the terminology specification to replace or adjust the non-standard terms in the text to ensure that the expression of the terms is consistent with the industry standards.
[0065] After generating these standardized element conversion results, they are integrated into a complete document version, namely the first standard document. During the integration process, the adjusted titles and chapter structures are recombined to construct a complete document outline; the standardized paragraph format is applied to the specific content of each chapter to ensure the consistency of the format within the document; the adjusted unit symbols and term expressions are also replaced to the corresponding positions to ensure that the use of professional terms and units meets the requirements of the standard library. Finally, through these steps, the first standard document is generated.
[0066] Step S300: Perform fine-grained identification on the first standard document through multi-dimensional deep comparison and difference cluster analysis to obtain a difference point adjustment set.
[0067] In this embodiment, in the process of obtaining the difference point adjustment set through multi-dimensional deep comparison and difference cluster analysis, the first standard document is firstly subjected to feature extraction to generate the target document feature set. Then, in combination with the power document rule standard library, the target document feature set is subjected to difference cluster analysis to determine the feature difference data set. Subsequently, the feature difference data set is subjected to feature hierarchical analysis to identify the location of the difference points, and a multi-dimensional aggregation analysis is performed to analyze the correlation between these difference points, thereby generating a multi-dimensional correlation data set. Based on the multi-dimensional correlation data set, the first standard document is contextually reconstructed, the context of the difference points is adjusted to make the document content more consistent with the standard, and finally the difference point adjustment set is determined.
[0068] Furthermore, in the method provided in the embodiment, the first standard document is finely identified through multi-dimensional deep comparison and difference cluster analysis to determine the difference point adjustment set, and further includes:
[0069] Perform feature extraction on the first standard document to determine a target document feature set; perform the difference clustering analysis on the target document feature set in combination with the electric power document rule standard library to determine a feature difference data set; perform feature hierarchical analysis on the feature difference data set to determine difference point location data and perform multidimensional aggregation analysis to analyze the correlation between difference points and generate a multidimensional correlation data set; based on the multidimensional correlation data set, perform context reconstruction on the first standard document to determine the difference point adjustment set.
[0070] In this embodiment, the first standard document is first parsed, and the text parsing technology combined with regular expressions is used to extract each element in the document in detail to generate a target document feature set. The parsing process includes identifying information such as title hierarchy, chapter structure, paragraph format, unit usage, and term usage. Common title formats in the text (such as "Chapter X", "1.2.3") are identified through regular expressions to establish the chapter hierarchy of the document; when parsing the paragraph format, the indentation, line spacing, and alignment of each paragraph are recorded. For unit usage, the units of measurement in the text, such as "kW" and "kWh", are matched by regular expressions, and the context in which they appear is recorded. The term usage part uses a word frequency statistical method to identify the commonly used professional terms in the document and their usage environment. These parsing results are integrated to obtain the target document feature set.
[0071] Then, the target document feature set is subjected to differential clustering analysis in combination with the power document rule standard library. In this step, each feature in the target document is compared with the corresponding requirements in the standard library one by one through a rule-based comparison method, and the degree of deviation is calculated. First, the difference of each feature is analyzed, such as the deviation at the chapter level, the difference in the frequency of term use, the deviation in the line spacing of paragraphs, etc. Then these differences are classified, and the features that deviate greatly from the standard are classified into the corresponding difference categories to generate a feature difference data set. For example, if it is identified that the line spacing of some paragraphs deviates greatly from the standard requirements, these paragraphs are classified into the "paragraph format deviation group"; if the frequency of term use is significantly different from the standard requirements, a "term use deviation group" is formed.
[0072] After generating the feature difference dataset, feature hierarchical analysis and multidimensional aggregation analysis are performed. Each difference point in the feature difference dataset is analyzed layer by layer through the manually set hierarchical analysis method. For example, in the chapter level analysis, the hierarchical structure of each chapter is checked, and the specific location of the chapter level errors is located one by one; in the paragraph format analysis, the paragraphs that do not meet the standards are identified and their positions in the document are recorded. Next, these difference points are multidimensionally analyzed using manually set association rules to check the mutual relationship and scope of influence between the difference points. For example, if some chapters are found to have both terminology and unit usage deviations, they are associated together to generate a multidimensional association dataset. The multidimensional association dataset records the connections between these difference points in detail.
[0073] After completing the multidimensional association analysis, the first standard document is contextually reconstructed based on the generated multidimensional association data set to ensure that the adjustment suggestions are consistent with the overall semantics and contextual logic of the document. Through the context analysis method, the text around the difference points is checked to ensure that the adjustment suggestions do not destroy the coherence of the document. For example, when adjusting certain terms, their usage in the context is considered to maintain the semantic consistency before and after the term replacement. At the same time, combined with the specifications in the standard library of power document rules, specific adjustment plans for each difference point are derived one by one, such as adjusting the line spacing of paragraphs, modifying the hierarchy of chapters, or replacing the expression of terms. Finally, these adjustment plans are summarized to generate a difference point adjustment set.
[0074] Step S400: generating a plurality of adjustment schemes according to the difference point adjustment set, and determining an optimal adjustment scheme through matching degree comparison, wherein when the matching degree comparison result does not meet the preset matching degree, the plurality of adjustment schemes are subjected to collision optimization to determine an optimal adjustment scheme.
[0075] In this embodiment, multiple adjustment schemes are first generated based on the difference point adjustment set, and the matching degree of these schemes is compared to determine the optimal adjustment scheme. Specifically, information is extracted from the difference point adjustment set, and multiple adjustment schemes are randomly generated, which represent different adjustment strategies or processing methods. Subsequently, the matching degree of each adjustment scheme with the power document rule standard library is calculated to evaluate its compliance with the standard requirements.
[0076] During the preliminary comparison, if there are one or more adjustment solutions with a matching degree higher than the preset standard, the solution with the highest matching degree is selected as the optimal adjustment solution. However, if the matching degree of all solutions does not meet the preset standard, collision optimization will be performed, that is, two solutions with lower matching degrees are selected and combined for adjustment. A new collision optimization solution is generated by exchanging parameters or merging strategies. The matching degree of this new solution is calculated again after content reconstruction. This process is repeated continuously, and the adjustment solutions are gradually optimized until a solution with a matching degree exceeding the preset standard is generated. Finally, this optimal solution is selected as the optimal adjustment solution.
[0077] Furthermore, in the method provided in the embodiment, multiple adjustment schemes are generated according to the difference point adjustment set, and the optimal adjustment scheme is determined by matching degree comparison, and further includes:
[0078] According to the difference point adjustment set, random selection is performed to determine multiple adjustment schemes, the first standard document is adjusted based on the multiple adjustment schemes, and the degree of matching with the power document rule standard library is calculated to determine multiple adjustment matching degrees; the multiple adjustment matching degrees are compared with the preset matching degrees, and if there is an adjustment matching degree greater than the preset matching degree, the adjustment scheme corresponding to the maximum value is selected as the optimal adjustment scheme.
[0079] In this embodiment, firstly, a plurality of adjustment schemes are generated by randomly selecting based on the difference point adjustment set. The difference point adjustment set is a set of non-compliant contents identified from the first standard document and their adjustment suggestions, and each adjustment scheme corresponds to a set of specific methods for processing these difference points. For example, the adjustment scheme includes adjusting the format of a paragraph, replacing specific terms, modifying the chapter structure, etc. By randomly extracting different combinations from the difference point adjustment set, a plurality of adjustment schemes are generated to cover various possible processing strategies.
[0080] Next, the generated multiple adjustment schemes are applied to the first standard document one by one to make specific content adjustments. These adjusted document versions are compared with the power document rule standard library, and the matching degree of each adjustment scheme with the standard library is calculated. The matching degree indicates the degree of conformity between the adjusted document content and the requirements of the standard library, and is a score obtained by comparing multiple features. For example, the similarity between the adjusted chapter structure and the chapter structure in the standard library, the consistency of the paragraph format, the standardization of the terminology used, the accuracy of the unit use, etc. are calculated. These calculation results are used to quantify the extent to which each adjustment scheme meets the standardization requirements of power documents.
[0081] Based on the calculation results of the matching degree, multiple adjustment matching degrees are generated, each of which corresponds to an adjustment plan, reflecting the adjustment effect of the plan. Subsequently, all adjustment matching degrees are compared with the preset matching degree standards. The preset matching degree is a pre-set minimum qualification standard used to determine whether the adjustment plan meets the requirements of the standard library. These adjustment matching degrees are checked. If there are one or more adjustment plans whose matching degree is higher than the preset standard, the adjustment plan with the highest matching degree is selected as the optimal adjustment plan.
[0082] Furthermore, in the method provided in the embodiment, when the matching degree comparison result does not meet the preset matching degree, the multiple adjustment schemes are subjected to collision optimization, and further include:
[0083] If there is no adjustment matching degree greater than the preset matching degree, traverse multiple adjustment schemes and randomly select two adjustment schemes for combined collision optimization to generate a first collision adjustment scheme; reconstruct the content of the first standard document based on the first collision adjustment scheme, determine the reconstruction result data, and calculate and determine the first adjustment matching degree; if the first adjustment matching degree is less than or equal to the preset matching degree, repeat the combined collision optimization and content reconstruction process until an Nth adjustment matching degree greater than the preset matching degree appears, and take the Nth adjustment scheme corresponding to the Nth adjustment matching degree as the optimal adjustment scheme, where N is an integer greater than or equal to 2.
[0084] In this embodiment, if there is no adjustment match greater than the preset match, first traverse multiple adjustment schemes, and select two schemes from them for combined collision through a random sampling algorithm. After selecting two schemes, they are combined and collided using a multi-parameter fusion method to generate a new first collision adjustment scheme. Combined collision refers to integrating and adjusting the different adjustment parameters in the two schemes to generate a better adjustment scheme. Specifically, this process includes chapter level adjustment, paragraph format parameter adjustment and term replacement optimization. Among them, when performing chapter level adjustment, a hierarchical parsing and reorganization method is used to analyze the adjustment of each scheme in the chapter structure, including chapter depth (such as whether it is refined into sub-chapters) and chapter order (such as reordering). Assume that Scheme A adjusts the chapter depth to make some content more refined, while Scheme B optimizes the chapter order to make the document structure clearer. Extract these adjustment information and merge them through the reorganization method. For example, the chapter depth adjustment in Scheme A (such as "Section 2.1 is refined into Sections 2.1.1 and 2.1.2") is combined with the chapter order adjustment in Scheme B (such as "bringing Chapter 3 to before Chapter 2") to generate a new chapter structure. In the fine-tuning process, the arrangement of the merged chapters is gradually adjusted according to the reference relationship and logical order between the chapter contents to ensure the logical consistency of the context. For example, if a chapter quotes the content of the previous chapter, it is prioritized to ensure that the quoted chapter is in front to maintain the coherence of the content. When adjusting the paragraph format parameters, the parameter adjustment and rule matching method is used to integrate the paragraph format parameters in the two schemes, such as line spacing, indentation format, and alignment. Suppose Scheme A recommends adjusting the line spacing to 1.5 times, while Scheme B recommends using an indentation format of 2 spaces. Use the weighted average method to integrate these parameters into a new adjustment scheme that combines the advantages of both. After the integration is completed, the document format is scanned paragraph by paragraph, and compared with the standard library, and minor adjustments are made to the parts that do not meet the standards. For example, when adjusting the line spacing, a line spacing between 1.45 times and 1.55 times is randomly selected according to the requirements of the standard library and the visual effect between paragraphs to ensure the aesthetics and standardization of the document in format. When optimizing term replacement, the term frequency analysis and context matching methods are used to analyze the different suggestions of the two schemes. Suppose that Scheme A recommends replacing a term globally, while Scheme B recommends replacing it only in a specific chapter. First, a term frequency analysis is performed to count the frequency of occurrence of the term in each chapter and evaluate the impact of global and local replacements. Next, the context matching method is used to analyze the context of the term in each position to ensure that the replacement does not affect the semantic integrity of the document. According to the results of the context analysis, the replacement strategy is adjusted, such as giving priority to standard terms in more technical chapters and retaining the original terms in other chapters.
[0085] After the first collision adjustment scheme is generated, the content of the first standard document is reconstructed based on the scheme, and the chapters, paragraphs and terminology of the document are reorganized to generate reconstruction result data. The reconstruction result data represents the new content structure of the document after the adjustment scheme is applied. The reconstruction result data is compared with the power document rule standard library, and the first adjustment matching degree of the scheme is calculated through the pre-built matching degree calculation function.
[0086] If the first adjustment still fails to reach the preset matching degree, continue to select two new adjustment schemes for combined collision optimization, generate new schemes, and reconstruct the content and recalculate the matching degree. Using the adaptive iteration method, gradually adjust the parameter integration strategy in each iteration, explore a wider range of adjustment combinations, and continuously optimize the adjustment effect. Each iteration gradually generates new optimization schemes by adjusting the parameter range, and finally achieves a gradual improvement in matching degree.
[0087] When the matching degree of a certain solution exceeds the preset standard, that is, when the Nth adjustment matching degree is greater than the preset matching degree, the solution is selected as the optimal adjustment solution, where N is an integer greater than or equal to 2, representing at least two combined collision optimizations.
[0088] Furthermore, in the method provided in the embodiment, calculating the degree of matching with the electric power document rule standard library further includes:
[0089] Construct a matching degree calculation function, the matching degree calculation function is as follows:
[0090] M=w 1 ×S c +w 2 ×C p +w 3 ×T m +w 4 ×U m +w 5 ×R f ; Where M is the matching degree, which indicates the matching degree between the document and the standard library of power document rules, S c is the chapter structure similarity, which indicates the similarity between the chapter structure of the document and the standard library. p is the consistency of paragraph content, indicating the matching degree between paragraph content and standard library, T m U is the term usage matching degree, which indicates the degree of conformity between the term usage in the literature and the standard library. m is the unit usage matching degree, indicating whether the unit usage in the literature conforms to the standard library specifications. f is the consistency of the citation format, indicating the consistency of the literature citation format with the standard library, w 1 、w 2 、w 3 、w 4、w 5 is the weight coefficient of each feature, the sum of which is 1; the matching degree with the power document rule standard library is calculated according to the matching degree calculation function to determine the adjusted matching degree.
[0091] In this embodiment, in order to calculate the matching degree between the first standard document after adjustment and the power document rule standard library, a matching degree calculation function is constructed for calculation. The matching degree calculation function is M=w 1 ×S c +w 2 ×C p +w 3 ×T m +w 4 ×U m +w 5 ×R f .
[0092] Among them, M is the matching degree, which indicates the matching degree between the document and the standard library of power document rules, S c is the chapter structure similarity, which indicates the similarity between the chapter structure of the document and the standard library. p is the consistency of paragraph content, indicating the matching degree between paragraph content and standard library, T m U is the term usage matching degree, which indicates the degree of conformity between the term usage in the literature and the standard library. m is the unit usage matching degree, indicating whether the unit usage in the literature conforms to the standard library specifications. f It is the consistency of citation format, which means the consistency of literature citation format with the standard library.
[0093] In progress c When calculating the target document and the standard library, the chapters are first analyzed hierarchically to extract the title, sub-chapter level and chapter order of each chapter. Then, the cosine similarity calculation formula is used to compare the matching degree between the target document and the standard library at the chapter level to obtain the chapter structure similarity S. c .
[0094] In progress C p When calculating C, we extract key semantic content from the document paragraphs through word segmentation and keyword extraction technology. Then, we compare these extracted keywords with the paragraph keywords in the standard library and calculate the similarity between the two. For example, we use Jaccard similarity or TF-IDF analysis to measure the similarity of the content of the document paragraph with the standard library in terms of word usage and topic, and generate C p .
[0095] In progress mWhen calculating T, first extract all the professional terms in the target document and count the frequency of each term in different chapters. Then compare the extracted terms with the term list in the standard library and record the number of correctly used terms. Then divide the number of correctly used terms by the total number of terms in the document to get T. m .
[0096] In progress m When calculating the U, we scan the physical quantities and units in the literature and extract their usage, such as the length unit "meter (m)" or "centimeter (cm)". Then, we compare these units with the unit forms specified in the standard library to check whether they meet the specifications. Then, we divide the number of correct unit usage by the total number of units to get U m .
[0097] In R f When calculating R, the references in the document are parsed to extract details such as the author name, publication year, and reference format. Then, this information is compared with the reference format specifications in the standard library, such as whether it meets the APA format requirements. Finally, the number of references that meet the standards is divided by the total number of references in the document to obtain R. f .
[0098] w 1 、w 2 、w 3 、w 4 、w 5 is the weight coefficient of each feature, the sum is 1, w 1 、w 2 、w 3 、w 4 、w 5 Pre-set for technical experts.
[0099] The aforementioned parameters are calculated individually and finally substituted into the matching degree calculation function to calculate the matching degree with the electric power literature rule standard library, and the adjustment matching degree is determined.
[0100] Step S500: According to the optimal adjustment scheme, the first standard document is optimized to generate a second standard document, and the second standard document is output as a target standard document.
[0101] In this embodiment, the optimal adjustment scheme is applied to adjust the chapter structure, paragraph content, terminology usage, unit usage and reference format of the document. For example, in the chapter structure, the chapter order is rearranged or the chapter level is refined to make the document structure clearer and more orderly; in the paragraph content, the line spacing, indentation mode and other parameters are adjusted to make the paragraph format meet the requirements of the standard library. In terms of terminology usage, specific terms are globally replaced or locally adjusted according to the scheme suggestions to ensure the accuracy and consistency of terminology usage; in the unit usage and reference format adjustment, the non-standard unit format is replaced and the reference format is standardized to make the document meet the industry standards. Through this process, the second standard document is obtained, and finally the second standard document is output as the target standard document. The document is optimized by the optimal adjustment scheme to ensure that it meets the standards of the power industry documents in terms of format, content and terminology, and realizes high-precision and high-efficiency standardized conversion, and finally solves the problems of low accuracy and inconsistent standards in the conversion process of the power industry documents.
[0102] In summary, this embodiment has at least the following technical effects:
[0103] The present invention interactively obtains the electric power document standard, and determines the mapping rules according to the historical standard documents, and establishes the electric power document rule standard library; compares and matches the target document based on the mapping rules in the electric power document rule standard library to generate the first standard document; through multi-dimensional deep comparison and difference cluster analysis, the first standard document is finely identified to obtain the difference point adjustment set; multiple adjustment schemes are generated according to the difference point adjustment set, and the optimal adjustment scheme is determined through matching degree comparison, wherein when the matching degree comparison result does not meet the preset matching degree, the multiple adjustment schemes are collision optimized to determine the optimal adjustment scheme; according to the optimal adjustment scheme, the first standard document is optimized to generate the second standard document, and the second standard document is output as the target standard document. The present invention solves the technical problem that the prior art has low accuracy in document standardization conversion, by establishing the electric power document rule standard library, performing preliminary matching on the target document to generate the first standard document, and then identifying the details that do not meet the standard through multi-dimensional deep comparison and difference cluster analysis, generating the difference point adjustment set, generating multiple adjustment schemes for these difference points, and then performing matching degree comparison and collision optimization, and finally generating the target standard document, so as to achieve the technical effect of improving the accuracy of document standardization conversion.
[0104] Embodiment 2, an inventive concept based on the same method for converting power industry documents based on mapping rules as in the above embodiment, such as Figure 2 As shown, the power industry document standardization conversion system based on mapping rules, the system and method embodiments in this embodiment are based on the same inventive concept. Among them, the system includes:
[0105] A standard library establishment module 11, which interactively obtains power literature standards and determines mapping rules based on historical standard documents to establish a power literature rule standard library; a first standard document generation module 12, which performs comparison and matching on target documents based on the mapping rules in the power literature rule standard library to generate a first standard document; a fine-grained recognition module 13, which performs fine-grained recognition on the first standard document through multi-dimensional depth comparison and difference clustering analysis to obtain a difference point adjustment set; an optimal adjustment plan determination module 14, which generates multiple adjustment plans according to the difference point adjustment set and determines the optimal adjustment plan through matching degree comparison. When the matching degree comparison result does not meet the preset matching degree, collision optimization is performed on the multiple adjustment plans to determine the optimal adjustment plan; a standard document output module 15, which optimizes the first standard document according to the optimal adjustment plan to generate a second standard document and outputs the second standard document as the target standard document.
[0106] Further, the system is also used to implement the following functions:
[0107] Receive the pre-uploaded power literature standards and the historical standard documents, where the power literature standards at least include document format requirements, unit usage specifications, and term standards, and the historical standard documents include title structure, chapter structure, paragraph format, unit usage, term usage, and citation format; match the historical standard documents with the power literature standards to generate the mapping rules; establish the power literature rule standard library based on the mapping rules, and the power literature rule standard library includes a format specification unit, a unit specification unit, a term specification unit, a chapter structure unit, and a citation specification unit.
[0108] Further, the system is also used to implement the following functions:
[0109] Parse the target document to determine the target document composition elements, where the target document composition elements include title, chapter structure, paragraph format, unit usage, and term usage; perform matching conversion on the target document composition elements based on the literature rule standard library to generate multiple element conversion results; integrate the multiple element conversion results to obtain the first standard document.
[0110] Further, the system is also used to implement the following functions:
[0111] Perform feature extraction on the first standard document to determine a target document feature set; perform the difference clustering analysis on the target document feature set in combination with the electric power document rule standard library to determine a feature difference data set; perform feature hierarchical analysis on the feature difference data set to determine difference point location data and perform multidimensional aggregation analysis to analyze the correlation between difference points and generate a multidimensional correlation data set; based on the multidimensional correlation data set, perform context reconstruction on the first standard document to determine the difference point adjustment set.
[0112] Furthermore, the system is also used to implement the following functions:
[0113] According to the difference point adjustment set, random selection is performed to determine multiple adjustment schemes, the first standard document is adjusted based on the multiple adjustment schemes, and the degree of matching with the power document rule standard library is calculated to determine multiple adjustment matching degrees; the multiple adjustment matching degrees are compared with the preset matching degrees, and if there is an adjustment matching degree greater than the preset matching degree, the adjustment scheme corresponding to the maximum value is selected as the optimal adjustment scheme.
[0114] Furthermore, the system is also used to implement the following functions:
[0115] If there is no adjustment matching degree greater than the preset matching degree, traverse multiple adjustment schemes and randomly select two adjustment schemes for combined collision optimization to generate a first collision adjustment scheme; reconstruct the content of the first standard document based on the first collision adjustment scheme, determine the reconstruction result data, and calculate and determine the first adjustment matching degree; if the first adjustment matching degree is less than or equal to the preset matching degree, repeat the combined collision optimization and content reconstruction process until an Nth adjustment matching degree greater than the preset matching degree appears, and take the Nth adjustment scheme corresponding to the Nth adjustment matching degree as the optimal adjustment scheme, where N is an integer greater than or equal to 2.
[0116] Furthermore, the system is also used to implement the following functions:
[0117] Construct a matching degree calculation function, the matching degree calculation function is as follows:
[0118] M=w 1 ×S c +w 2 ×C p +w 3 ×T m +w 4 ×U m +w 5 ×R f ; Where M is the matching degree, which indicates the matching degree between the document and the standard library of power document rules, S cis the chapter structure similarity, which indicates the similarity between the chapter structure of the document and the standard library. p is the consistency of paragraph content, indicating the matching degree between paragraph content and standard library, T m U is the term usage matching degree, which indicates the degree of conformity between the term usage in the literature and the standard library. m is the unit usage matching degree, indicating whether the unit usage in the literature conforms to the standard library specifications. f is the consistency of the citation format, indicating the consistency of the literature citation format with the standard library, w 1 、w 2 、w 3 、w 4 、w 5 is the weight coefficient of each feature, the sum of which is 1; the matching degree with the power document rule standard library is calculated according to the matching degree calculation function to determine the adjusted matching degree.
[0119] It should be noted that the sequence of the above embodiments is for description only and does not represent the advantages and disadvantages of the embodiments. The above is a description of a specific embodiment of this specification. The processes depicted in the accompanying drawings do not necessarily require the specific order and continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0120] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
[0121] This specification and the accompanying drawings are merely exemplary illustrations of the present invention and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the scope of the present invention. Thus, the present invention is intended to include such changes and modifications if they fall within the scope of the present invention and its equivalents.
Claims
1. A method for standardizing and converting documents in the electric power industry based on mapping rules, characterized in that: The method comprises the following steps: S100, interactively obtain power document standards, determine mapping rules based on historical standard documents, and establish a power document rule standard library; S200, comparing and matching the target document based on the mapping rules in the power document rule standard library to generate a first standard document; S300, performing fine-grained identification on the first standard document through multi-dimensional deep comparison and difference cluster analysis to obtain a difference point adjustment set; S400, generating multiple adjustment schemes according to the difference point adjustment set, and determining the optimal adjustment scheme through matching degree comparison, wherein when the matching degree comparison result does not meet the preset matching degree, the multiple adjustment schemes are subjected to collision optimization to determine the optimal adjustment scheme; S500: According to the optimal adjustment scheme, the first standard document is optimized to generate a second standard document, and the second standard document is output as a target standard document.
2. The method for standardizing and converting documents in the electric power industry based on mapping rules according to claim 1, characterized in that: Step S100 includes: Receiving the pre-uploaded electric power document standard and the historical standard document, wherein the electric power document standard at least includes document format requirements, unit usage specifications and terminology standards, and the historical standard document includes title structure, chapter structure, paragraph format, unit usage, terminology usage and citation format; Matching the historical standard document with the power document standard to generate the mapping rule; The electric power document rule standard library is established based on the mapping rule, and the electric power document rule standard library includes a format specification unit, a unit specification unit, a term specification unit, a chapter structure unit and a reference specification unit.
3. The method for standardizing and converting documents in the electric power industry based on mapping rules according to claim 1, characterized in that: Step S200 includes: Parsing the target document to determine the target document constituent elements, wherein the target document constituent elements include title, chapter structure, paragraph format, unit usage, and term usage; Matching and converting the target document constituent elements based on the document rule standard library to generate multiple element conversion results; The multiple element conversion results are integrated to obtain the first standard document.
4. The method for standardizing and converting documents in the electric power industry based on mapping rules according to claim 1, characterized in that: Step S300 includes: Extracting features from the first standard document to determine a target document feature set; Performing the difference clustering analysis on the target document feature set in combination with the electric power document rule standard library to determine a feature difference data set; Performing feature hierarchical analysis on the feature difference data set, determining the difference point location data and performing multidimensional aggregation analysis, analyzing the correlation between the difference points, and generating a multidimensional correlation data set; Based on the multi-dimensional associated data set, the first standard document is contextually reconstructed to determine the difference point adjustment set.
5. The method for converting power industry documents based on mapping rules as claimed in claim 4, characterized in that: Generating multiple adjustment schemes according to the difference point adjustment set, and determining the optimal adjustment scheme by matching degree comparison, the method further includes: According to the difference point adjustment set, a plurality of adjustment schemes are randomly selected to determine, the first standard document is adjusted based on the plurality of adjustment schemes, and the matching degree with the power document rule standard library is calculated to determine a plurality of adjustment matching degrees; The multiple adjustment matching degrees are compared with a preset matching degree. If there is an adjustment matching degree greater than the preset matching degree, an adjustment scheme corresponding to the maximum value is selected as the optimal adjustment scheme.
6. The method for standardizing and converting documents in the electric power industry based on mapping rules according to claim 5, characterized in that: When the matching degree comparison result does not meet the preset matching degree, the multiple adjustment schemes are subjected to collision optimization, and the method includes: If there is no adjustment matching degree greater than the preset matching degree, traverse multiple adjustment schemes and randomly select two adjustment schemes for combined collision optimization to generate a first collision adjustment scheme; Reconstructing the content of the first standard document based on the first collision adjustment scheme, determining reconstruction result data, and calculating and determining a first adjustment matching degree; If the first adjustment matching degree is less than or equal to the preset matching degree, the combined collision optimization and content reconstruction process is repeated until an Nth adjustment matching degree greater than the preset matching degree appears, and the Nth adjustment scheme corresponding to the Nth adjustment matching degree is used as the optimal adjustment scheme, where N is an integer greater than or equal to 2.
7. The method for converting power industry documents based on mapping rules as claimed in claim 5, characterized in that: Calculating the degree of matching with the electric power document rule standard library, the method comprises: Construct a matching degree calculation function, the matching degree calculation function is as follows: M=w1×S c +w2×C p +w3×T m +w4×U m +w5×R f ; Among them, M is the matching degree, which indicates the matching degree between the document and the standard library of power document rules, S c is the chapter structure similarity, which indicates the similarity between the chapter structure of the document and the standard library. p is the consistency of paragraph content, indicating the matching degree between paragraph content and standard library, T m U is the term usage matching degree, which indicates the degree of conformity between the term usage in the literature and the standard library. m is the unit usage matching degree, indicating whether the unit usage in the literature conforms to the standard library specifications. f is the consistency of the citation format, indicating the consistency of the literature citation format with the standard library. w1, w2, w3, w4, and w5 are the weight coefficients of each feature, and the sum is 1; The matching degree with the electric power document rule standard library is calculated according to the matching degree calculation function to determine the adjustment matching degree.
8. A mapping rule-based power industry document standardization conversion system, characterized in that: The system comprises: A standard library establishment module, wherein the standard library establishment module interactively acquires power document standards, determines mapping rules based on historical standard documents, and establishes a power document rule standard library; A first standard document generation module, which compares and matches the target document based on the mapping rules in the power document rule standard library to generate a first standard document; A fine-grained identification module, wherein the fine-grained identification module performs fine-grained identification on the first standard document through multi-dimensional depth comparison and difference clustering analysis to obtain a difference point adjustment set; an optimal adjustment scheme determination module, wherein the optimal adjustment scheme determination module generates a plurality of adjustment schemes according to the difference point adjustment set, and determines the optimal adjustment scheme by matching degree comparison, wherein when the matching degree comparison result does not meet the preset matching degree, the plurality of adjustment schemes are subjected to collision optimization to determine the optimal adjustment scheme; A standard document output module is configured to optimize the first standard document according to the optimal adjustment scheme, generate a second standard document, and output the second standard document as a target standard document.