Intelligent aggregation analysis method and system for standard draft
Patent Information
- Application Number
- CN202611117055.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-27
AI Technical Summary
[0005]本发明的目的在于提供一种标准征求意见稿的智能汇总分析方法及系统,用于解决现有技术对标准草案对应的若干反馈意见的信息汇总效果差的技术问题
在公开征求标准草案的反馈意见,并相应收集得到多个原始意见稿后,先通过对多个原始意见稿进行语义聚类,将针对标准草案中的同一条款的不同原始意见稿归入同一类簇,以实现对多个原始意见稿的自动化文本分组,在此基础上,对每一个稿件类簇进行语义价值分析,通过对稿件类簇包括的原始意见稿的反馈信息进行深入挖掘,区分不同稿件类簇在标准草案完善过程中的数据重要性差异,也即区分标准草案中的不同条款的讨论优先级,据此构建意见稿汇总报告,可以较为准确地反映标准文件关联的利益相关方的真实意见,提升标准草案对应的若干反馈意见的信息汇总效果,为后续的草案完善工作提供可靠的数据支持。
Smart Images

Figure CN122635362B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of data analysis, specifically to an intelligent summary and analysis method and system for a draft standard. Background Technology
[0002] In the process of developing and revising standards, soliciting opinions is a crucial step in ensuring the scientific rigor, impartiality, and operability of the standards. The standards drafting working group typically solicits feedback from stakeholders such as industry experts, enterprises, research institutions, testing organizations, and the general public regarding the draft standards.
[0003] In practice, a draft national or industry standard often receives a large number of feedback opinions from various organizations and individuals. Currently, the collection and analysis of this feedback mainly relies on manual processing by the standard drafting working group. Specifically, staff members need to read each piece of feedback, extract the specific clause numbers, technical indicator modification suggestions, and wording adjustment opinions, manually enter them into a summary table, and categorize and organize them according to their relevance and reasonableness. During this process, staff members also need to use their personal experience to determine which opinions involve safety, environmental protection, or significant industry conflicts of interest and require priority discussion and response, while which are textual or format-related editing suggestions can be processed later. Finally, based on the results of manual sorting, a summary table of opinions is formed as the basis for subsequent discussions on the revision of the standard.
[0004] Because manual review of feedback is inefficient and is significantly constrained by the reviewer's personal biases and experience, the final summary table of comments is unlikely to accurately reflect the true opinions of stakeholders related to the standard document. This will result in subsequent revisions to the draft standard being less effective than expected. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent summary and analysis method and system for standard drafts, which solves the technical problem of poor information summary effect of existing technology for several feedback opinions corresponding to standard drafts.
[0006] In a first aspect, one embodiment of the present invention provides an intelligent summary and analysis method for standard draft for comments, the method comprising: Semantic clustering was performed on multiple original drafts to obtain multiple draft clusters, wherein the original drafts were the comments made by relevant parties on the draft standard; In the multiple manuscript clusters, semantic value analysis is performed on each manuscript cluster to obtain a value score for each manuscript cluster. The value score is used to indicate the importance of the corresponding manuscript cluster in the process of improving the draft standard. Based on the multiple manuscript clusters and the value score of each manuscript cluster, a summary report of the draft opinions is constructed.
[0007] Optionally, the step of performing semantic value analysis on each of the plurality of manuscript clusters to obtain a value score for each manuscript cluster includes: Analyze the logical depth of the corresponding clauses in the draft standard for the target manuscript cluster to obtain the document location factor of the target manuscript cluster, wherein the target manuscript cluster is any one of the multiple manuscript clusters; Analyze the content richness of the target manuscript clusters to obtain the information value factor of the target manuscript clusters; Analyze the degree of conflict between the target manuscript cluster and its corresponding clause in the draft standard to obtain the positional contradiction factor of the target manuscript cluster; Analyze the proportion of original opinion drafts included in the target manuscript cluster among the multiple original opinion drafts to determine the attention factor of the target manuscript cluster; The value score of the target manuscript cluster is determined based on the document positioning factor, information value factor, stance contradiction factor, and attention factor.
[0008] Optionally, the analysis of the logical depth of the target manuscript cluster in the corresponding clause of the draft standard to obtain the document location factor of the target manuscript cluster includes: The logical chain of the draft standard is parsed to obtain the chapter structure tree; Based on the corresponding clauses of the target manuscript cluster in the standard draft, determine the target node associated with the target manuscript cluster in the chapter structure tree; The semantic similarity between the target manuscript cluster and its corresponding clause in the draft standard is analyzed to obtain the semantic similarity value of the target manuscript cluster; The document location factor of the target manuscript cluster is determined based on the semantic similarity value of the target manuscript cluster and the node depth of the target node in the chapter structure tree.
[0009] Optionally, the step of analyzing the semantic similarity between the target manuscript cluster and its corresponding clause in the draft standard to obtain a semantic similarity value for the target manuscript cluster includes: Within the target manuscript cluster, content that overlaps with the corresponding clauses in the standard draft is removed to obtain the target clean cluster; Calculate the cosine similarity between the semantic vector of the target clean cluster and the semantic vector of the corresponding clause in the standard draft of the target manuscript cluster, and obtain the cosine similarity value of the target manuscript cluster; The absolute value of the cosine similarity of the target manuscript cluster is determined as the semantic similarity value of the target manuscript cluster.
[0010] Optionally, the semantic similarity value is positively correlated with the corresponding document location factor, and the node depth of the target node in the document structure tree is positively correlated with the document location factor of the target manuscript cluster.
[0011] Optionally, the analysis of the content richness of the target manuscript cluster to obtain the information value factor of the target manuscript cluster includes: Extract content words from all original drafts included in the target manuscript cluster to obtain multiple target content words included in the target manuscript cluster; Based on the pre-configured part-of-speech weighting rules, the TF-IDF values of multiple target content words included in the target manuscript cluster are weighted and calculated to obtain the initial information density value of the target manuscript cluster. Logical rules are identified for all original drafts included in the target manuscript cluster to obtain at least one logical rule included in the target manuscript cluster; Based on the number of logical rules included in the target manuscript cluster, the initial value of the information density of the target manuscript cluster is amplified to obtain the information value factor of the target manuscript cluster.
[0012] Optionally, the analysis of the degree of conflict between the target manuscript cluster and its corresponding clause in the draft standard yields the positional contradiction factor of the target manuscript cluster, including: If there is a parameter value conflict between the target manuscript cluster and its corresponding clause in the draft standard, and the target manuscript cluster includes conflict explanation text for the corresponding parameter value, the preset first contradiction factor is determined as the position contradiction factor of the target manuscript cluster. If there is a parameter value conflict between the target manuscript cluster and its corresponding clause in the draft standard, and the target manuscript cluster does not include the conflict explanation text for the corresponding parameter value, the preset second contradiction factor is determined as the position contradiction factor of the target manuscript cluster. If there is no conflict in parameter values between the target manuscript cluster and its corresponding clause in the draft standard, the preset third contradiction factor is determined as the position contradiction factor of the target manuscript cluster. Wherein, the first contradiction factor is greater than the second contradiction factor, and the second contradiction factor is greater than the third contradiction factor.
[0013] Secondly, another embodiment of the present invention also provides an intelligent summary and analysis system for standard draft comments, the system comprising: The manuscript clustering module is used to perform semantic clustering on multiple original draft opinions to obtain multiple manuscript clusters, wherein the original draft opinions are the opinions provided by relevant parties regarding the draft standard; The value analysis module is used to perform semantic value analysis on each of the multiple manuscript clusters to obtain a value score for each manuscript cluster, wherein the value score is used to indicate the importance of the corresponding manuscript cluster in the process of improving the draft standard. The manuscript summary module is used to construct a summary report of the draft opinions based on the multiple manuscript clusters and the value score of each manuscript cluster.
[0014] Thirdly, in another embodiment of the present invention, an electronic device is provided, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described in the first aspect.
[0015] Fourthly, in another embodiment of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the method described in the first aspect.
[0016] The present invention has the following beneficial effects: After soliciting feedback on the draft standard and collecting multiple original draft comments, semantic clustering was first performed on these original draft comments. This grouped different original draft comments on the same clause in the draft standard into the same cluster, thus achieving automated text grouping of multiple original draft comments. Based on this, semantic value analysis was performed on each draft cluster. By deeply mining the feedback information of the original draft comments included in the draft cluster, the differences in data importance of different draft clusters in the process of improving the draft standard were distinguished, that is, the discussion priority of different clauses in the draft standard was distinguished. Based on this, a summary report of the draft comments was constructed, which can more accurately reflect the true opinions of stakeholders related to the standard document, improve the information summary effect of several feedback opinions corresponding to the draft standard, and provide reliable data support for subsequent draft improvement work. Attached Figure Description
[0017] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating an intelligent summary and analysis method for standard draft comments provided in an embodiment of the present invention; Figure 2This is a schematic diagram of the structure of an intelligent summary and analysis system for standard draft comments provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0019] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of an intelligent summary analysis method and system based on a draft standard for comments proposed in accordance with the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0021] The following description, in conjunction with the accompanying drawings, details the specific scheme of the intelligent summary and analysis method and system for standard drafts provided by this invention.
[0022] In one embodiment, the present invention provides an intelligent summary and analysis method for standard draft for comments, such as... Figure 1 As shown, the method includes: Step S1: Perform semantic clustering on multiple original drafts to obtain multiple draft clusters.
[0023] The original draft comments refer to the comments provided by relevant parties regarding the draft standard.
[0024] The aforementioned draft standard is specifically a formal working document (also known as a consultation draft) prepared by the standard drafting working group based on preliminary research and experimental results, submitted to the standardization technical committee or through public channels, and used to solicit technical opinions from relevant parties. This document is in an intermediate stage within the standard system, awaiting feedback and refinement, and has not yet undergone technical review or administrative approval.
[0025] The comments from stakeholders regarding the draft standard are: original feedback documents submitted by stakeholders during the standard consultation phase, through designated channels (such as letters, email attachments, online forms, etc.), containing specific modification suggestions, technical viewpoints, or questions, and which have undergone data preprocessing.
[0026] In practice, due to differences in the expression methods of various stakeholders in the draft standard, the collected original feedback documents may have different data formats. Therefore, before performing semantic clustering on multiple original drafts, it is necessary to unify the format of the original feedback documents (such as Word files, PDF files, or web forms) to convert them into plain text format (text extraction for image documents is performed using OCR technology, and text extraction for form documents is performed using pre-configured data mapping rules). After noise reduction (removing content in the original draft that is not related to the modification suggestions, such as greetings, repetitive signatures, and formatted explanatory text) and automatic splitting (splitting the content in the same original feedback document that points to different clauses in the draft standard), the corresponding semantic clustering work is then carried out.
[0027] Because different original draft opinions vary greatly in their expression, level of detail, and focus, directly performing semantic clustering on the full text of the original draft opinions is easily affected by differences in expression habits. This can cause opinions that should be directed at the same clause to be scattered into different clusters. However, the content of the original draft opinions usually points to specific clauses in the standard draft. Therefore, this invention chooses to first group multiple original draft opinions based on the clauses they point to in the standard draft, to ensure that several original draft opinions within the same group point to the same clause. Then, semantic clustering is carried out within the group based on the content of different original draft opinions, so as to achieve refined clustering processing of multiple original draft opinions.
[0028] For example, a lightweight semantic model (such as Sentence-BERT) can be used to extract the semantic vectors of each original draft, then the cosine similarity between the semantic vectors of different original drafts can be calculated, and the semantic distance between different original drafts can be quantified accordingly. Finally, a clustering algorithm can be used to process the semantic distance between different original drafts within the same group to obtain multiple draft clusters.
[0029] Step S2: In the multiple manuscript clusters, perform semantic value analysis on each manuscript cluster to obtain the value score of each manuscript cluster.
[0030] The value score is used to indicate the importance of the data in the process of improving the draft standard for the corresponding manuscript cluster.
[0031] Specifically, the step of performing semantic value analysis on each of the multiple manuscript clusters to obtain a value score for each manuscript cluster includes: Analyze the logical depth of the corresponding clauses in the draft standard for the target manuscript cluster to obtain the document location factor of the target manuscript cluster, wherein the target manuscript cluster is any one of the multiple manuscript clusters; Analyze the content richness of the target manuscript clusters to obtain the information value factor of the target manuscript clusters; Analyze the degree of conflict between the target manuscript cluster and its corresponding clause in the draft standard to obtain the positional contradiction factor of the target manuscript cluster; Analyze the proportion of original opinion drafts included in the target manuscript cluster among the multiple original opinion drafts to determine the attention factor of the target manuscript cluster; The value score of the target manuscript cluster is determined based on the document positioning factor, information value factor, stance contradiction factor, and attention factor.
[0032] In the above settings, the multi-dimensional joint analysis method can effectively suppress the noise interference that may be introduced by single-dimensional analysis, and ensure that the importance of data for manuscript clusters in the process of improving the standard draft can be accurately assessed.
[0033] In this invention, the values of the document positioning factor, information value factor, positional contradiction factor, and attention factor are all in the range of 0-1. The value score is obtained by weighted calculation of the corresponding document positioning factor, information value factor, positional contradiction factor, and attention factor. The higher the value score, the higher the data importance of the corresponding manuscript cluster in the process of improving the standard draft.
[0034] When revising clauses based on feedback, the core is to discuss the specific content of the clauses. The key to the discussion is to debate and make decisions on issues with technical disagreements. The arguments and volume of feedback play a supporting role. Therefore, the calculation weights of the document positioning factor, information value factor, positional contradiction factor, and attention factor are set as follows: 0.35, 0.2, 0.25, and 0.15, respectively.
[0035] Specifically, the larger the value of the document location factor, the deeper the logical depth of the corresponding clause in the draft standard.
[0036] In actual production, the logical framework of multiple clauses in a draft standard usually does not generate controversy. Moreover, the broader the content of the clauses, the less likely they are to generate substantive disputes, and correspondingly, the more difficult it is to demonstrate their feasibility. Conversely, the more specific the content of the clauses, the more likely they are to trigger substantive content that will be tested, followed, and disputed during the implementation of the standard, and the easier it is to demonstrate their feasibility. Analyzing the logical depth of the corresponding clauses in the draft standard by analyzing the document positioning factor can effectively distinguish the specificity of the clauses discussed by the corresponding document clusters, and accordingly amplify the data contribution of document clusters pointing to specific clauses, while reducing the data contribution of document clusters pointing to broad clauses.
[0037] The higher the value of the information value factor, the richer the information density and the lower the content redundancy at the semantic level of the feedback opinions contained in the corresponding manuscript cluster.
[0038] The feedback collected during the consultation phase of the draft standard came from a wide range of sources and varied in quality. A large number of feedback comments only contained social expressions or emotional exclamations such as "agree," "no comment," or "received." Although such comments consumed a lot of work in processing, they did not provide any incremental information on the technical content of the draft standard. Another part of the feedback comments, although lengthy, were merely restates or paraphrases of the original text of the draft standard, without proposing new directions for modification or supplementary evidence, and their actual information contribution was close to zero.
[0039] The aforementioned opinions with low information content and high value coexist. If they are treated equally without distinction, on the one hand, the information density of the opinion summary report will be severely diluted, increasing the cognitive burden on the revision working group to screen effective opinions. On the other hand, a small number of high-value opinions containing key modification suggestions or technical arguments will be submerged in a large number of invalid texts and will be easily overlooked or underestimated.
[0040] Analyzing the richness of manuscript clusters can effectively distinguish the level of information increment carried by different manuscript clusters, and accordingly amplify the data contribution of manuscript clusters that have sufficient technical arguments, clear modification directions, and contain causal logic chains, while reducing the data contribution of manuscript clusters that have no information or are purely perfunctory.
[0041] The larger the value of the positional contradiction factor, the stronger the logical opposition or technical deviation between the feedback opinions contained in the manuscript cluster and the corresponding clauses they refer to in the draft standard.
[0042] The core value of soliciting opinions on draft standards lies in exposing differences. Regardless of whether a clause is a "requirement" with a mandatory element, a "recommendation" with room for choice, or a "statement" used only to clarify concepts, as long as the feedback sends a signal in the opposite direction regarding the substantive content of the clause, it means that there are points of disagreement in the industry's understanding of the clause that are worth further discussion.
[0043] By analyzing the degree of conflict between manuscript clusters and their corresponding clauses in the draft standard, it is possible to effectively capture the logical polarity relationship between feedback opinions and the original clauses (i.e., whether the opinions constitute positive reinforcement, neutral supplementation, or negative opposition to the original clauses in the direction of modification), and accordingly amplify the data contribution of all manuscript clusters that constitute negative opposition, thereby ensuring that the review committee prioritizes controversial issues with genuine differences of opinion.
[0044] The higher the value of the attention factor, the higher the proportion of the original drafts covered by the manuscript cluster in the total number of original drafts, that is, the higher the attention the clause receives within the scope of the opinion solicitation.
[0045] The public consultation on draft standards typically covers a wide range of people. Different stakeholders, due to differences in their technical backgrounds, industry positioning, and business needs, often exhibit vastly different levels of interest and willingness to provide feedback on the same standard clause. When only a single source provides feedback on a particular standard clause, that feedback may only represent the specific needs of a particular organization or user, raising questions about its general applicability and representativeness. Conversely, when the same clause receives similar feedback from multiple independent stakeholders (even if the specific wording of each feedback differs), it indicates that the clause faces widespread implementation obstacles or misunderstandings in practical application. This represents a common weakness in the standard text, making it more necessary to discuss and more urgent for revision during the draft standard revision process.
[0046] By analyzing the proportion of original drafts included in a manuscript cluster among the multiple original drafts, we can effectively measure the breadth of consensus coverage of different clauses within the scope of soliciting opinions, and accordingly amplify the data contribution of manuscript clusters that receive independent feedback from multiple parties, while reducing the data contribution of manuscript clusters that are proposed only by a limited number of sources.
[0047] Furthermore, the analysis of the logical depth of the target manuscript cluster in the corresponding clause of the draft standard yields the document location factor of the target manuscript cluster, including: The logical chain of the draft standard is parsed to obtain the chapter structure tree; Based on the corresponding clauses of the target manuscript cluster in the standard draft, determine the target node associated with the target manuscript cluster in the chapter structure tree; The semantic similarity between the target manuscript cluster and its corresponding clause in the draft standard is analyzed to obtain the semantic similarity value of the target manuscript cluster; The document location factor of the target manuscript cluster is determined based on the semantic similarity value of the target manuscript cluster and the node depth of the target node in the chapter structure tree.
[0048] This invention uses document intelligence technology (such as LayoutLM or PaddleStructure) to perform logical layout parsing on the standard draft to obtain the chapter structure tree.
[0049] The chapter's structure tree can be understood as structured information stored in a tree structure, used to represent the logical architecture and hierarchical relationship of the entire draft standard.
[0050] Specifically, the chapter structure data is organized along the chapter hierarchy of the draft standard. According to the inherent hierarchical structure of the standard document (including but not limited to the nested sequence of "Chapter-Section-Article-Clause-Item" or "Chapter X-Section X-Article X-Clause X-Item X"), the full text of the draft standard is broken down into several hierarchical nodes with clear subordinate relationships, and constructed into an ordered tree in the form of parent-child node association.
[0051] The draft standard has a strict hierarchical structure (such as chapters, sections, articles, clauses, and items), and the functional positioning and technical binding force of clauses at different levels in the standard system are significantly different.
[0052] However, during the consultation process, the feedback came from a wide range of sources and varied in quality, with many instances of inaccurate citations: some feedback vaguely pointed to a broad chapter in the draft standard, while others, although containing specific clause numbers, had actual semantic content that was seriously inconsistent with the technical subject matter of the clause in question.
[0053] If the corresponding clause is determined solely based on the number stated in the feedback or the chapter manually located by the user, it is very easy to mistakenly include substantive modification suggestions for clause A in the opinion set for clause B, resulting in systematic bias in subsequent value assessment due to mismatched input data.
[0054] To address this issue, this invention first parses the logical chain of the draft standard to generate a chapter structure tree, establishing clear semantic coordinates for each level of clauses. Based on this, it further calculates the semantic similarity between the overall semantics of the manuscript cluster and the original text of the corresponding clause it claims to point to. By cross-validating the user's claimed location with the actual semantic matching location, it effectively filters and eliminates misaligned feedback where the numbering and content are inconsistent. Finally, it combines the validated semantic similarity value with the inherent depth of the target node in the chapter structure tree to determine the document location factor.
[0055] This approach allows manuscript clusters that point to more specific and deeper-level clauses to receive higher positioning factor values. On the other hand, it effectively prevents value assessment distortion caused by incorrect citations or ambiguous positioning by the feedback provider, significantly improving the data credibility of document positioning factors.
[0056] Preferably, the step of analyzing the semantic similarity between the target manuscript cluster and its corresponding clause in the draft standard to obtain a semantic similarity value for the target manuscript cluster includes: Within the target manuscript cluster, content that overlaps with the corresponding clauses in the standard draft is removed to obtain the target clean cluster; Calculate the cosine similarity between the semantic vector of the target clean cluster and the semantic vector of the corresponding clause in the standard draft of the target manuscript cluster, and obtain the cosine similarity value of the target manuscript cluster; The absolute value of the cosine similarity of the target manuscript cluster is determined as the semantic similarity value of the target manuscript cluster.
[0057] The application revealed significant differences in how feedback was written. Some feedback providers tended to first restate the original draft standard text and then attach their proposed modifications; others directly offered modification suggestions without restating the original text; still others merely paraphrased the original draft standard text without proposing any substantive changes. In these situations, directly calculating the semantic similarity between the complete text of the draft standard and the corresponding clauses of the draft standard presents the following problems: For feedback with extensive restatements, the semantic similarity to the standard clauses is artificially inflated, but this high similarity does not stem from a substantive connection between the feedback and the clauses, but is merely a mechanical repetition of the original text; conversely, for concise feedback that does not include restatements and directly addresses the key points of modification, the original similarity to the standard clauses is actually lower, but its actual anchoring accuracy and modification value are often higher.
[0058] By setting up the above, before calculating the semantic similarity value, the interference of paraphrasing in the feedback can be eliminated by removing content that is repeated with the corresponding clauses of the draft standard. This improves the accuracy and reliability of the semantic verification process.
[0059] In this invention, removing content that is repeated with the corresponding clause in the standard draft of the target manuscript cluster should be understood as: removing content that is repeated with the corresponding clause in the standard draft of each original draft included in the target manuscript cluster.
[0060] The steps for calculating the cosine similarity between the semantic vector of the target cleanliness cluster and the semantic vector of the corresponding clause in the draft standard, obtaining the cosine similarity value of the target manuscript cluster, and determining the absolute value of the cosine similarity value of the target manuscript cluster as the semantic similarity value of the target manuscript cluster are as follows: In all original drafts that have undergone duplicate content removal within the target clean cluster, any two different original drafts are combined, and the absolute value of the cosine similarity of the semantic vectors of the two in the combination is calculated to obtain multiple inter-draft similarity values. Then, among the multiple inter-manuscript similarity values, the inter-manuscript similarity values associated with the same original draft opinion after removing duplicate content are summed to obtain multiple inter-manuscript similarity features that correspond one-to-one with multiple original draft opinion opinions after removing duplicate content. Then, the original draft with the greatest inter-draft similarity and after removing duplicate content, and the original draft with the least inter-draft similarity and after removing duplicate content, are both identified as representative drafts of the target draft cluster. Then, the cosine similarity between the semantic vectors of the two representative drafts of the target draft cluster and the semantic vectors of the corresponding clauses of the target draft cluster in the standard draft are calculated. Finally, the larger of the absolute values of the two cosine similarities is output as the semantic similarity value of the target draft cluster.
[0061] In the above settings, the original opinion draft with the greatest inter-manuscript similarity and after removing duplicate content is used to represent the mainstream opinion in the target manuscript cluster, so as to ensure that the semantic similarity comparison process covers the mainstream opinion in the target manuscript cluster; while the original opinion draft with the least inter-manuscript similarity and after removing duplicate content is used to represent the specific opinion in the target manuscript cluster, so as to ensure that the semantic similarity comparison process preserves the specific opinion in the target manuscript cluster.
[0062] The semantic similarity value is positively correlated with the corresponding file location factor, and the node depth of the target node in the chapter structure tree is positively correlated with the file location factor of the target manuscript cluster.
[0063] For example, the ratio of the node depth of the target node in the document structure tree to the maximum node depth of the document structure tree can be calculated first, and then the average of the ratio and the corresponding semantic similarity value can be determined as the corresponding file location factor.
[0064] The steps for analyzing the content richness of target manuscript clusters and obtaining the information value factor of the target manuscript clusters include: Extract content words from all original drafts included in the target manuscript cluster to obtain multiple target content words included in the target manuscript cluster; Based on the pre-configured part-of-speech weighting rules, the TF-IDF values of multiple target content words included in the target manuscript cluster are weighted and calculated to obtain the initial information density value of the target manuscript cluster. Logical rules are identified for all original drafts included in the target manuscript cluster to obtain at least one logical rule included in the target manuscript cluster; Based on the number of logical rules included in the target manuscript cluster, the initial value of the information density of the target manuscript cluster is amplified to obtain the information value factor of the target manuscript cluster.
[0065] Feedback messages come in various forms and vary in quality. Their information value is not directly related to the length of the text. What is more important is whether the text provides a logical explanation for the feedback and whether the logic is self-consistent.
[0066] Existing methods for assessing information content (such as directly calculating text length or word frequency) have significant technical limitations: on the one hand, texts that use a large number of function words, conjunctions, and polite expressions may be long, but the actual effective information they carry is extremely limited; on the other hand, texts containing complex causal reasoning and conditional arguments may not be the longest, but their reference value for standard revision is far greater than that of simple declarative sentences.
[0067] If information content is assessed solely based on the frequency of word occurrence, it is easy to misjudge lengthy but empty feedback as high value, while underestimating short but logically sound and concise opinions as low value, resulting in a significant deviation between the value assessment results and the actual business situation.
[0068] By performing content word extraction (using Jieba word segmentation combined with part-of-speech filtering) and part-of-speech weighted calculation, different weights are assigned to content words of different parts of speech. This results in feedback containing a large number of domain-specific terms and precise technical verbs receiving higher initial information density values, while the information contribution of feedback containing only emotional adjectives or general function words is effectively suppressed. This ensures that the initial information density value can truly reflect the information carrying capacity of feedback at the technical vocabulary level.
[0069] It should be noted that since nouns (i.e., technical terms) carry core technical concepts, they are given the highest weight in the part-of-speech weighting rules; verbs express operations or judgments, so their weight is set slightly lower than that of nouns; as for adjectives and adverbs, they only play a modifying role, so their weight is set lower than that of nouns and verbs.
[0070] For example, in the part-of-speech weighting rule, the part-of-speech weights of nouns, verbs, adjectives and adverbs can be 1, 0.75, 0.35 and 0.2 respectively.
[0071] Feedback containing causal arguments often explains the technical basis for the revisions, feedback containing conditional logic usually clarifies the scope of application, and feedback containing adversative logic reflects a deep consideration of the pros and cons. All of these have higher reference value than simple declarative sentences.
[0072] Since logical structure is of great significance in standard argumentation, this invention further identifies logical rules for manuscript clusters (using a large language model, completed through logical rule extraction by prompting engineering), and amplifies the initial value of information density based on the number of identified logical rules. By amplifying the information density of feedback opinions containing multiple logical rules, it distinguishes between straightforward factual statements and systematically argued modification suggestions, so that the latter can obtain a higher quantitative manifestation in the information value factor.
[0073] For example, logical rules include, but are not limited to, causal logic, conditional logic, transitional logic, and parallel progressive logic.
[0074] In one example, the adjustment coefficient corresponding to the initial information density value can be the product of a preset coefficient sensitivity value (such as 0.02) and the number of logical rules included in the target manuscript cluster. The amplification coefficient corresponding to the initial information density value is the sum of its corresponding adjustment coefficient and the value 1. After amplification, the initial information density value is then processed by the Sigmoid function, and the result is the corresponding information value factor.
[0075] Furthermore, the analysis of the degree of conflict between the target manuscript cluster and its corresponding clause in the draft standard yields the positional contradiction factor of the target manuscript cluster, including: If there is a parameter value conflict between the target manuscript cluster and its corresponding clause in the draft standard, and the target manuscript cluster includes conflict explanation text for the corresponding parameter value, the preset first contradiction factor is determined as the position contradiction factor of the target manuscript cluster. If there is a parameter value conflict between the target manuscript cluster and its corresponding clause in the draft standard, and the target manuscript cluster does not include the conflict explanation text for the corresponding parameter value, the preset second contradiction factor is determined as the position contradiction factor of the target manuscript cluster. If there is no conflict in parameter values between the target manuscript cluster and its corresponding clause in the draft standard, the preset third contradiction factor is determined as the position contradiction factor of the target manuscript cluster. Wherein, the first contradiction factor is greater than the second contradiction factor, and the second contradiction factor is greater than the third contradiction factor.
[0076] The conflict explanation text should be understood as: statements that explain the reasons, basis or rationality for adjusting parameter values, including but not limited to: citing external basis (such as test data, field measured values, industry practice), stating implementation difficulties (such as insufficient equipment capacity, infeasibility of process), expressing technical viewpoints (such as safety considerations, economic evaluation), etc.
[0077] Analysis revealed that the opinions raised in the feedback on the draft standard that contradicted the direction of the parameters in the original text had significant differences in their business nature.
[0078] Some feedback stemmed from in-depth research or experimental verification of the standard's implementation scenarios. While proposing reverse parameter suggestions, they provided relatively sufficient supporting evidence (such as "Since existing equipment generally cannot meet this accuracy requirement, it is recommended to relax ±1% to ±2%" or "Based on the data from the XX test report, the original threshold is too conservative"). Other feedback, while also proposing suggestions in the opposite direction to the original clause parameters, were limited to expressing subjective positions or rough tendencies (such as "This value is unreasonable" or "It should be reduced"), lacking specific explanations of the technical basis.
[0079] If the existence of conflicting parameter values is used as the sole criterion for determining the intensity of a positional conflict, then these two types of conflicting opinions, which are of vastly different natures, will be treated the same in terms of value assessment. This will result in empty objections lacking evidence receiving higher scores due to their contradictory nature, while constructive objections with solid evidence will fail to receive the priority they deserve.
[0080] By employing the aforementioned settings and parameter value conflict detection, the core controversial opinions that truly involve changes in standard technical indicators are accurately identified, avoiding the misjudgment of mere textual polishing suggestions as substantive conflicts. Based on the detection of parameter value conflicts, the system further determines whether the conflict cluster contains explanatory text addressing that parameter conflict. The manuscript clusters with parameter value conflicts are further subdivided into two subcategories: objective disagreements supported by evidence and subjective doubts without basis. A first conflict factor and a second conflict factor are assigned accordingly, with the former being set significantly higher than the latter. This allows dissenting opinions with sufficient evidence to receive a higher value assignment in the dimension of positional conflict, thereby enhancing the value of substantive controversial issues with a basis for discussion. As for manuscript clusters without any parameter value conflicts, the lowest third conflict factor is assigned.
[0081] For example, the first contradiction factor, the second contradiction factor, and the third contradiction factor can be set to 1, 0.5, and 0 respectively.
[0082] The steps to obtain the attention factor are as follows: The ratio of the proportion of original drafts included in the target manuscript cluster among the multiple original drafts to the reference proportion is calculated to obtain the attention factor of the corresponding manuscript cluster.
[0083] The reference percentage can be understood as the maximum percentage of original draft opinions included in a cluster among multiple original draft opinions. This value can be set manually based on actual business conditions, or it can be calculated based on the maximum historical percentage in historical data (by converting the number of historical original draft opinions included in the historical dataset, the maximum historical percentage, and the number of multiple original draft opinions currently involved).
[0084] In application, when the proportion of original drafts included in the target manuscript cluster among the multiple original drafts is greater than the reference proportion, the attention factor of its corresponding manuscript cluster is forcibly locked to 1.
[0085] Step S3: Construct a summary report of the draft opinions based on the multiple manuscript clusters and the value score of each manuscript cluster.
[0086] In the application, each manuscript can be classified into different value levels according to its value score and a preset scoring threshold.
[0087] Specifically, manuscripts with a value score of 0.7 or higher are marked as high-value, labeled with key controversial opinions, forced to be displayed at the top of the report, and automatically included in the candidate list of topics for discussion at the review meeting. Manuscripts with a value score between 0.4 and 0.7 are marked as medium value, given a general revision suggestion label, and placed in the regular revision suggestion area according to the chapter order they point to, for the drafting team to process one by one; Manuscripts with a value score below 0.4 are marked as low-value, labeled as "reference comments," and placed in the report appendix or general comment section, and are not considered for priority review.
[0088] Furthermore, the summary report of the draft comments will summarize the modification intentions of all original draft comments in each draft cluster and generate a standardized summary of revision directions, which includes: the location and content of the clauses to which the draft cluster refers, and the summary content of the original draft comments of the two representative claims in the cluster (i.e., the original draft comments with the greatest inter-draft similarity and after removing duplicate content and the original draft comments with the least inter-draft similarity and after removing duplicate content).
[0089] In the application, the draft summary report supports multiple output formats. When output as a Word document, the differences and revisions are marked with revision mode, making it easy for the review team to directly modify the standard text. When output as a PDF, it includes a bookmark navigation structure, facilitating printing, distribution, and long-term archiving. When output as an Excel spreadsheet, it supports multi-dimensional filtering and sorting operations, facilitating data analysis and statistical summarization. And when output as an HTML webpage, it supports foldable and expandable interactive methods, facilitating online review and multi-person collaboration.
[0090] In summary, after soliciting feedback on the draft standard and collecting multiple original draft opinions, this invention first performs semantic clustering on these original draft opinions, grouping different original draft opinions on the same clause in the draft standard into the same cluster. This achieves automated text grouping of multiple original draft opinions. Based on this, semantic value analysis is performed on each draft cluster. By deeply mining the feedback information of the original draft opinions included in the draft cluster, the differences in data importance of different draft clusters in the process of improving the draft standard are distinguished, that is, the discussion priority of different clauses in the draft standard is distinguished. Based on this, a summary report of the draft opinions is constructed, which can more accurately reflect the true opinions of stakeholders related to the standard document, improve the information summary effect of several feedback opinions corresponding to the draft standard, and provide reliable data support for subsequent draft improvement work.
[0091] In another embodiment, the present invention also provides an intelligent summary and analysis system for standard draft comments, such as... Figure 2 As shown, the system 200 includes: The manuscript clustering module 201 is used to perform semantic clustering on multiple original draft opinions to obtain multiple manuscript clusters, wherein the original draft opinions are the opinions provided by relevant parties regarding the draft standard; The value analysis module 202 is used to perform semantic value analysis on each of the multiple manuscript clusters to obtain a value score for each manuscript cluster, wherein the value score is used to indicate the importance of the corresponding manuscript cluster in the process of improving the draft standard. The manuscript summary module 203 is used to construct a summary report of the draft opinions based on the multiple manuscript clusters and the value score of each manuscript cluster.
[0092] It should be noted that the system provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the intelligent summary and analysis system for a draft standard and the intelligent summary and analysis method for a draft standard provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0093] This invention also provides an electronic device. Please refer to [link to relevant documentation]. Figure 3 The electronic device may include a processor 301, a memory 302, and a program 3021 stored in the memory 302 and capable of running on the processor 301.
[0094] When program 3021 is executed by processor 301, it can achieve the following: Figure 1 Any steps in the corresponding method embodiments and the achievement of the same beneficial effects will not be repeated here.
[0095] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by hardware related to program instructions, and the program can be stored in a readable medium.
[0096] This invention also provides a readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described functions. Figure 1 Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.
[0097] The computer-readable storage medium of this invention can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0098] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0099] The program code contained on the storage medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0100] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In cases involving a remote computer, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0101] This invention also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to achieve the intelligent summary and analysis method for standard draft comments provided in the above embodiments.
[0102] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0103] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
Claims
1. A method for intelligently summarizing and analyzing draft standards for public comment, characterized in that: The method includes: Semantic clustering was performed on multiple original drafts to obtain multiple draft clusters, wherein the original drafts were the comments made by relevant parties on the draft standard; In the multiple manuscript clusters, semantic value analysis is performed on each manuscript cluster to obtain a value score for each manuscript cluster. The value score is used to indicate the importance of the corresponding manuscript cluster in the process of improving the draft standard. Based on the multiple manuscript clusters and the value score of each manuscript cluster, a summary report of the draft opinions is constructed; The step involves performing semantic value analysis on each of the multiple manuscript clusters to obtain a value score for each manuscript cluster, including: Analyze the logical depth of the corresponding clauses in the draft standard for the target manuscript cluster to obtain the document location factor of the target manuscript cluster, wherein the target manuscript cluster is any one of the multiple manuscript clusters; Analyze the content richness of the target manuscript clusters to obtain the information value factor of the target manuscript clusters; Analyze the degree of conflict between the target manuscript cluster and its corresponding clause in the draft standard to obtain the positional contradiction factor of the target manuscript cluster; Analyze the proportion of original opinion drafts included in the target manuscript cluster among the multiple original opinion drafts to determine the attention factor of the target manuscript cluster; The value score of the target manuscript cluster is determined based on the document positioning factor, information value factor, stance contradiction factor, and attention factor. The analysis of the logical depth of the target manuscript cluster in the corresponding clause of the standard draft yields the document location factor of the target manuscript cluster, including: The logical chain of the draft standard is parsed to obtain the chapter structure tree; Based on the corresponding clauses of the target manuscript cluster in the draft standard, determine the target node associated with the target manuscript cluster in the chapter structure tree; The semantic similarity between the target manuscript cluster and its corresponding clause in the draft standard is analyzed to obtain the semantic similarity value of the target manuscript cluster; The document location factor of the target manuscript cluster is determined based on the semantic similarity value of the target manuscript cluster and the node depth of the target node in the chapter structure tree.
2. The intelligent summary and analysis method for standard drafts according to claim 1, characterized in that, The analysis of the semantic similarity between the target manuscript cluster and its corresponding clause in the draft standard, to obtain the semantic similarity value of the target manuscript cluster, includes: Within the target manuscript cluster, content that overlaps with the corresponding clauses in the standard draft is removed to obtain the target clean cluster; Calculate the cosine similarity between the semantic vector of the target clean cluster and the semantic vector of the corresponding clause in the standard draft of the target manuscript cluster, and obtain the cosine similarity value of the target manuscript cluster; The absolute value of the cosine similarity of the target manuscript cluster is determined as the semantic similarity value of the target manuscript cluster.
3. The intelligent summary and analysis method for draft standards according to claim 1, characterized in that, The semantic similarity value is positively correlated with the corresponding file location factor, and the node depth of the target node in the chapter structure tree is positively correlated with the file location factor of the target manuscript cluster.
4. The intelligent summary and analysis method for standard drafts according to claim 1, characterized in that, The analysis of the content richness of the target manuscript clusters yields the information value factors of the target manuscript clusters, including: Extract content words from all original drafts included in the target manuscript cluster to obtain multiple target content words included in the target manuscript cluster; Based on the pre-configured part-of-speech weighting rules, the TF-IDF values of multiple target content words included in the target manuscript cluster are weighted and calculated to obtain the initial information density value of the target manuscript cluster. Logical rules are identified for all original drafts included in the target manuscript cluster to obtain at least one logical rule included in the target manuscript cluster; Based on the number of logical rules included in the target manuscript cluster, the initial value of the information density of the target manuscript cluster is amplified to obtain the information value factor of the target manuscript cluster.
5. The intelligent summary and analysis method for draft standards according to claim 1, characterized in that, The analysis of the target manuscript cluster and its corresponding clause in the draft standard yields a positional conflict factor for the target manuscript cluster, including: If there is a parameter value conflict between the target manuscript cluster and its corresponding clause in the draft standard, and the target manuscript cluster includes conflict explanation text for the corresponding parameter value, the preset first contradiction factor is determined as the position contradiction factor of the target manuscript cluster. If there is a parameter value conflict between the target manuscript cluster and its corresponding clause in the draft standard, and the target manuscript cluster does not include the conflict explanation text for the corresponding parameter value, the preset second contradiction factor is determined as the position contradiction factor of the target manuscript cluster. If there is no conflict in parameter values between the target manuscript cluster and its corresponding clause in the draft standard, the preset third contradiction factor is determined as the position contradiction factor of the target manuscript cluster. Wherein, the first contradiction factor is greater than the second contradiction factor, and the second contradiction factor is greater than the third contradiction factor.
6. An intelligent summary and analysis system for draft standards, characterized in that, The system for implementing the intelligent summary and analysis method for standard drafts as described in any one of claims 1-5, the system comprising: The manuscript clustering module is used to perform semantic clustering on multiple original draft opinions to obtain multiple manuscript clusters, wherein the original draft opinions are the opinions provided by relevant parties regarding the draft standard; The value analysis module is used to perform semantic value analysis on each of the multiple manuscript clusters to obtain a value score for each manuscript cluster, wherein the value score is used to indicate the importance of the corresponding manuscript cluster in the process of improving the draft standard. The manuscript summary module is used to construct a summary report of the draft opinions based on the multiple manuscript clusters and the value score of each manuscript cluster.
7. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when executed by the processor, the computer program implements the steps of the intelligent summary analysis method for the draft standard as described in any one of claims 1 to 5.
8. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the intelligent summary analysis method for the draft standard as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Opinion information collection method and apparatus, computer device and storage medium
CN108737248A
Policy and regulation solicitation opinion intelligent release and visual analysis method and system
CN117472992A