Detection scheme automatic generation method and device based on multi-stage information extraction
By employing multi-stage information extraction and automatic assembly technologies, the problem of information fragmentation in the generation of engineering testing solutions has been solved, achieving high accuracy and consistency in testing solutions, improving the completeness and automated generation efficiency of testing solutions, and adapting to the testing needs of complex engineering scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN BEST INNOVATION INFORMATION TECH CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-01
AI Technical Summary
In the current process of generating engineering testing solutions, information understanding, standard matching, and process configuration are fragmented, lacking a unified semantic foundation and overall linkage mechanism. This leads to the quality of solution generation relying on human experience, which can easily result in problems such as information omissions, inconsistent standard selection, mismatch between approval processes and project levels, and unreasonable allocation of testing resources. In particular, it is difficult to achieve overall closed-loop control in scenarios with complex project types, diverse testing items, or frequent updates to standards and specifications.
By employing multi-stage information extraction methods, including document parsing, named entity recognition, relation extraction, and attribute classification, a structured information set is formed. Standard clauses, testing methods, and acceptance requirements are located in the standard and specification knowledge base. The testing plan content is automatically assembled, and the association judgment is performed in conjunction with the engineering level and approval process rules to achieve automatic configuration and resource matching.
It improves the accuracy and consistency of engineering testing solutions, avoids problems such as information omissions and inconsistent standard selection, enhances the completeness, accuracy and automated generation efficiency of testing solutions, and achieves adaptive response in process configuration and resource allocation, adapting to the testing needs of complex engineering scenarios.
Smart Images

Figure CN121961485A_ABST
Abstract
Description
An automatic generation method and apparatus for detection schemes based on multi-stage information extraction Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to a method and apparatus for automatically generating detection schemes based on multi-stage information extraction. Background Technology
[0002] In the field of engineering testing, testing plans are typically developed based on documents such as project contracts, technical agreements, and design specifications. These documents are often unstructured, with scattered content and inconsistent expression. Currently, the generation of testing plans relies heavily on manual methods. Technicians must read and understand numerous unstructured documents line by line, manually extracting key information such as project category, testing items, and project level. They then combine this information with personal experience or by consulting national standards and industry specifications to manually select applicable standard clauses, testing methods, and acceptance requirements, and develop the testing plan accordingly. Furthermore, they must manually determine the approval process and the allocation of testing resources based on the project's attributes. While some systems have introduced document management or template-based filling tools, their core remains at the level of information display or semi-automatic filling, failing to achieve a systematic understanding of document semantics and cross-stage collaborative processing.
[0003] In existing technologies, the generation process of testing solutions is fragmented between information understanding, specification matching, and process configuration, lacking a unified semantic foundation and overall linkage mechanism. This results in the quality of solution generation being highly dependent on human experience, easily leading to problems such as information omissions, inconsistent specification selection, mismatch between approval processes and project levels, and unreasonable allocation of testing resources. Especially in scenarios with complex project types, diverse testing items, or frequently updated standards and specifications, manual methods struggle to achieve timely and accurate closed-loop control from original documents to testing solutions, and then to process and resource allocation. This deficiency cannot be fundamentally resolved through simple template optimization or manual verification.
[0004] Therefore, a method is needed to improve the accuracy of engineering testing solutions. Summary of the Invention
[0005] This invention provides a method and apparatus for automatically generating detection schemes based on multi-stage information extraction, which can improve the accuracy of engineering information detection while achieving automated processing.
[0006] In a first aspect, this invention provides an automatic generation method for detection schemes based on multi-stage information extraction. The method includes: receiving and parsing unstructured raw documents, performing document parsing processing on the content, and converting it into text data; based on the text data, performing multi-stage information extraction processing according to a preset information extraction order to form a structured information set, wherein the multi-stage information extraction processing sequentially includes named entity recognition, relation extraction, and attribute classification; using the engineering category, detection items, and engineering level of the structured information set as combination constraints, locating the corresponding standard clauses, detection methods, and acceptance requirements in a standard and specification knowledge base, and automatically assembling and generating detection scheme content while maintaining the semantic consistency of the structured information; based on the detection scheme content and the engineering attribute results determined in the structured information set, associating the engineering level with approval process rules and associating the detection item type with detection resource rules to achieve automatic configuration of approval nodes and adaptive matching of detection personnel or detection equipment.
[0007] In a second aspect, the present invention provides an automatic detection scheme generation apparatus based on multi-stage information extraction. The apparatus is used to execute an automatic detection scheme generation method based on multi-stage information extraction as described above. The apparatus includes an acquisition module, a processing module, and an output module, wherein: the acquisition module is used to receive and parse unstructured raw documents, perform document parsing processing on the content, and convert it into text data; the processing module is used to perform multi-stage information extraction processing based on the text data according to a preset information extraction order to form a structured information set, wherein the multi-stage information extraction processing sequentially includes named entity recognition and relation extraction. The processing module is used to take the engineering category, testing item, and engineering level of the structured information set as combination constraints, locate the corresponding standard clauses, testing methods, and acceptance requirements in the standard and specification knowledge base, and automatically assemble and generate the testing scheme content while maintaining the semantic consistency of the structured information. The output module is used to determine the automatic configuration of approval nodes and the adaptive matching of testing personnel or testing equipment by associating the engineering level with the approval process rules and the testing item type with the testing resource rules based on the testing scheme content and the engineering attribute results determined in the structured information set.
[0008] In a third aspect of the invention, an electronic device is provided, including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the preceding embodiments.
[0009] In a fourth aspect of the invention, a non-transitory computer-readable storage medium is provided, the computer-readable storage medium storing instructions that, when executed, perform the method as described in any of the preceding claims.
[0010] In summary, one or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages: 1. The present invention performs multi-stage information extraction processing on unstructured engineering documents, and stabilizes key information such as engineering category, testing items and engineering level in a structured form. On this basis, it drives the positioning of standard specifications, generation of testing schemes and configuration of processes and resources through a combination constraint method, so that information understanding, specification matching and execution configuration are established on a unified semantic basis, thereby avoiding information omissions or misselections caused by human understanding bias and single-dimensional matching. At the same time, the entities, relationships and attributes formed by multi-stage information extraction maintain a continuous inheritance relationship in each processing stage, which can accurately characterize the semantic boundaries of engineering testing and constrain subsequent rule judgments, so that the testing clauses, testing methods, acceptance requirements and approval and resource configurations are precisely corresponding to the actual attributes of the engineering, thus significantly improving the overall accuracy of engineering information testing in terms of completeness, consistency and accuracy.
[0011] 2. By sequentially performing semantic normalization, named entity recognition, relation extraction, and attribute classification on text data, the engineering objects, engineering categories, inspection items, engineering levels, and related constraints are expressed in a complete, accurate, and consistent structure at the semantic level. This avoids the impact of differences in human understanding and semantic ambiguity on the results of key information extraction, improves the completeness, accuracy, and reusability of basic information acquisition for engineering inspection, and provides a stable data foundation for subsequent specification matching and scheme generation.
[0012] 3. By using project category, testing items, and project level as combined constraints, the standard and specification knowledge base is matched for consistency. After semantic verification, the content of the testing scheme is automatically assembled, so that the standard clauses, testing methods, and acceptance requirements in the testing scheme correspond precisely to the actual attributes of the project. This avoids the problem of omission or mismatch in the selection of specifications, thereby significantly improving the technical effect of the testing scheme in terms of specification applicability, logical consistency, and automated generation efficiency.
[0013] 4. By semantically mapping the content of the testing plan with the results of structured engineering attributes, and introducing engineering level, testing risk and mandatory standard attributes to automatically determine the approval process and testing resources in a rule-driven manner, the configuration of approval nodes and resource matching no longer rely on manual experience judgment. This enables the process configuration and resource allocation to adaptively respond to engineering attributes, thereby improving the accuracy, compliance and overall operational efficiency of the testing execution phase.
[0014] 5. By simultaneously identifying the main project category and the special project category during the attribute classification stage, and constructing a multi-level project category expression, the project category can truly reflect the hierarchical relationship between the project structure and the detection semantics. This avoids the problem of special detection requirements being omitted or misclassified due to the use of only a single project category, and improves the completeness of the project detection semantic expression and the adaptability to complex project scenarios.
[0015] 6. By introducing a hierarchical rule set for engineering categories and a semantic association determination mechanism, a clear hierarchical relationship is established between the main engineering categories and the special engineering categories. The hierarchical information is encapsulated in a structured manner, so that the hierarchical relationship, applicable boundaries and constraints between engineering categories can be accurately inherited and utilized in subsequent processing stages. This provides reliable hierarchical semantic support for hierarchical standard retrieval and detection requirement combination.
[0016] 7. By performing hierarchical expansion and layered retrieval of multi-level engineering categories, and by performing hierarchical constraint combination and semantic consistency verification between the main engineering category and the special engineering category, the basic testing requirements and special testing requirements can be orderly integrated in the same testing scheme without duplication or conflict. This achieves the overall technical effect of complete coverage of testing clauses, consistent constraints on testing requirements, and structured arrangement of testing schemes in complex engineering category nesting scenarios. Attached Figure Description
[0017] Figure 1 is a flowchart illustrating an automatic generation method for detection schemes based on multi-stage information extraction disclosed in an embodiment of the present invention; Figure 2 is a module diagram illustrating an automatic generation device for detection schemes based on multi-stage information extraction disclosed in an embodiment of the present invention; Figure 3 is a structural diagram illustrating an electronic device disclosed in an embodiment of the present invention.
[0018] Explanation of reference numerals in the attached drawings: 201, acquisition module; 202, processing module; 203, output module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0020] In the description of the embodiments of the present invention, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0021] In the description of the embodiments of the present invention, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0022] The current process of generating engineering testing solutions relies heavily on manual interpretation and experience-based judgment of unstructured documents. Information understanding, standard matching, and process and resource allocation are fragmented, lacking a unified semantic foundation and linkage mechanism. This results in low efficiency and poor consistency in solution preparation, and problems such as omission of key information, inconsistent standard selection, mismatch between approval process and project level, and unreasonable allocation of testing resources. In particular, in scenarios with complex project types, numerous testing items, or frequent updates to standards and specifications, these defects are further amplified, making it difficult to achieve a closed-loop control from original documents to testing solutions and then to process and resource allocation.
[0023] This embodiment discloses an automatic generation method for detection schemes based on multi-stage information extraction. Referring to Figure 1, it includes the following steps S110-S140: S110, by receiving and parsing unstructured raw documents, performing document parsing processing on the content, and converting it into text data.
[0024] This invention discloses an automatic generation method for detection schemes based on multi-stage information extraction, which is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a backend server running an automatic generation method for detection schemes based on multi-stage information extraction. The server can be implemented using a standalone server or a server cluster composed of multiple servers.
[0025] By receiving and parsing unstructured raw documents, the system first performs document reception and type determination processing on the unstructured raw documents. By identifying the document's file format, encoding method, and content carrier form, it distinguishes whether the unstructured raw document is a text document, a scanned document, or a mixed document. It then matches the corresponding parsing strategy for different types of unstructured raw documents, thereby ensuring that the subsequent document parsing processing has a definite processing path and technical constraints at the input level.
[0026] After determining the document type, text extraction is performed on the unstructured original document for text-type documents. By parsing the paragraph structure, table structure and heading hierarchy of the document, the character sequence is extracted while maintaining the original logical order and hierarchical identifiers. This ensures that the extracted text content can still reflect the sequential and subordinate relationships between the engineering descriptions, technical clauses and explanatory statements in the document.
[0027] For scanned or image-based documents, after determining the document type, character recognition processing is performed on the unstructured original document. By segmenting the document page, locating text regions, and recognizing characters, the engineering specifications, technical agreements, and contract terms in image form are converted into calculable character sequences. During the character recognition process, the page position and row / column order of the characters are recorded simultaneously to preserve the semantic arrangement features in the original document.
[0028] After completing text extraction or character recognition, the resulting character sequence is cleaned by removing non-semantic characters, standardizing symbolic representation, correcting line breaks and page breaks, and eliminating duplicate paragraphs. This ensures that the character sequence meets the basic requirements of subsequent natural language processing in terms of content integrity and semantic continuity, thereby avoiding interference from noise information in subsequent information extraction.
[0029] After the text cleaning process is completed, the character sequence is processed by text structuring. By identifying paragraph boundaries, sentence boundaries, and clause structure features common in the field of engineering inspection, the character sequence is semantically segmented and reordered, so that the processed text content forms a continuous and stable text data representation in terms of structure.
[0030] By sequentially performing document reception and type determination processing, text extraction or character recognition processing, text cleaning processing, and text structuring and organizing processing, the original unstructured raw documents with diverse expressions and inconsistent structures are stably converted into semantically continuous, clearly structured text data that can be directly used by subsequent multi-stage information extraction and processing.
[0031] S120, based on text data, performs multi-stage information extraction processing according to a preset information extraction order to form a structured information set.
[0032] In one possible implementation, based on text data, multi-stage information extraction processing is performed according to a preset information extraction order to form a structured information set. Specifically, this includes: performing semantic normalization processing on the text data to obtain processed data; performing named entity recognition processing on the processed data according to a preset information extraction order, and annotating the names of engineering objects, engineering categories, inspection items, standard references, and engineering attribute description fragments to form a named entity set; and performing relation extraction processing based on the named entity set, by analyzing the co-occurrence positions, semantic references, and predefined engineering... The entity association model is tested by constructing a relationship network, which includes the hierarchical relationship between engineering categories and testing items, the reference relationship between testing items and standard clauses, the limiting relationship between engineering objects and engineering levels, and the constraint relationship between engineering attributes and testing requirements. Based on the relationship network, attribute classification processing is performed on named entities. By combining the semantic role, association strength, and upstream and downstream relationship types of named entities in the relationship network, the attributes of engineering category, engineering level, testing items, specifications, and resource requirements are refined, distinguished, and unified, thus forming a structured information set containing named entities, entity relationships, and entity attributes.
[0033] Specifically, when performing semantic regularization processing on text data for engineering inspection semantics, the text data is first calibrated at the semantic granularity within a unified engineering inspection semantic space. This is achieved by standardizing paragraph boundaries, sentence boundaries, and frequently occurring technical expression patterns in the engineering inspection field, so that the same engineering semantics can be mapped into stable semantic units under different expressions. Semantic regularization processing is used to eliminate differences in expression style, sentence structure, and terminology usage between contract texts, technical agreement texts, and design specification texts. This is achieved by standardizing and replacing specialized terms in the engineering inspection field, semantically merging synonymous expressions, and semantically completing engineering descriptions across sentences or paragraphs. This results in processed data that meets the requirements for subsequent information extraction in terms of both semantic continuity and expression standardization.
[0034] After obtaining the processed data, named entity recognition processing is performed on the data according to a preset information extraction order. By analyzing the text fragments sequentially within the engineering detection semantic domain, text content with clear engineering semantic orientation is identified as named entities. Among them, the engineering object name is used to represent the specific engineering entity being detected, the engineering category name is used to limit the major category attribute to which the engineering belongs, the detection project name is used to indicate the specific detection content, the standard specification reference identifier is used to mark the source of the standard on which the engineering detection is based, and the engineering attribute description fragment is used to characterize the attribute characteristics such as the engineering level, scale, or importance. By annotating the above different types of text fragments into entities, the key information that originally existed in the form of natural language is transformed into named entities with clear semantic boundaries and category identifiers, forming a set of named entities that retains the original contextual position information to support subsequent semantic relationship analysis.
[0035] After forming a named entity set, relation extraction processing is performed based on the named entity set. By jointly analyzing the relative positions of named entities in the processed data, contextual dependencies, and predefined entity association patterns in the engineering testing domain, semantic relationships between named entities are identified and established. Among them, co-occurrence position is used to determine the possibility of association between entities within the same semantic unit, semantic orientation is used to identify the master-slave or limiting relationship between entities, and engineering testing entity association patterns are used to constrain which entity types are allowed to establish relationships. Through the above analysis, the subordinate relationship between engineering categories and testing projects is constructed to limit the applicable engineering scope of testing projects, the reference relationship between testing projects and standard clauses is constructed to clarify the source of testing basis, the limiting relationship between engineering objects and engineering levels is constructed to characterize the importance of engineering projects, and the constraint relationship between engineering attributes and testing requirements is constructed to limit the testing execution conditions, thereby forming a relational network that can reflect the overall logical structure of engineering testing.
[0036] After constructing the relationship network, attribute classification processing is performed on named entities based on the network. By combining the semantic role of the named entity in the relationship network, the strength of its association with other entities, and its upstream and downstream position in the network, the engineering attributes carried by the named entities are refined, differentiated, and unified. Among them, the semantic role is used to distinguish the functional positioning of the entity in the engineering inspection semantics, the association strength is used to measure the importance of the entity attribute in the engineering inspection, and the upstream and downstream relationship type is used to determine the pre- or constraint role of the attribute in the inspection process. Through this attribute classification processing, engineering category attributes, engineering grade attributes, inspection project attributes, specification attributes, and resource requirement attributes are classified separately, and the attribute results formed by the same entity in different contexts are unified and integrated. In the end, a structured information set containing named entities, entity relationships, and entity attributes is formed, so that the engineering inspection semantics are solidified in a stable and reusable structured form and can be directly used for subsequent inspection scheme generation and process configuration.
[0037] S130 uses the engineering category, testing items, and engineering level of the structured information set as combination constraints, locates the corresponding standard clauses, testing methods, and acceptance requirements in the standard and specification knowledge base, and automatically assembles and generates the testing scheme content while maintaining the semantic consistency of the structured information.
[0038] In one possible implementation, the project category, testing items, and project level of the structured information set are used as combined constraints. The corresponding standard clauses, testing methods, and acceptance requirements are located in the standard and specification knowledge base. While maintaining the semantic consistency of the structured information, the testing plan content is automatically assembled and generated. Specifically, this includes: performing combined constraint construction processing on the project category, testing items, and project level in the structured information set to form combined constraints, wherein the project category is limited to the scope of application of the specification, the testing items are limited to the technical content constraints, and the project level is limited to the stringency constraints of the specification; based on the combined constraints, specification location processing is performed in the standard and specification knowledge base. By matching the project category with the standard's applicable scope labels, the testing items with the standard clause themes, and the project level with the level limitation clauses, a set of standard clauses that simultaneously satisfy the combined constraints is obtained. Association parsing is performed on this set of standard clauses to extract the corresponding testing methods and acceptance requirements. Semantic consistency verification is then performed on the standard clauses, testing methods, and acceptance requirements based on the structured information set, forming a set of standard elements. Automatic assembly is then performed on this set of standard elements to generate the testing plan content, where standard clauses are used as the standard's basis, testing methods as the testing implementation content, and acceptance requirements as the result judgment content, all in a unified arrangement.
[0039] Specifically, when performing combined constraint construction on engineering categories, testing items, and engineering grades in a structured information set, constraint vectors are first established for each engineering category, testing item, and engineering grade within the same semantic coordinate system. The engineering category is mapped to a specification scope constraint to limit the set of retrievalable specification domains in the standard specification knowledge base. Testing items are mapped to technical content constraints to limit the retrieval topics of clause themes, testing method paragraphs, and acceptance requirement paragraphs. Engineering grades are mapped to specification severity constraints to limit the applicable branches of clauses with grade limitations, thus forming combined constraint conditions that can be uniformly calculated. Among these, the specification scope constraint controls the boundary of the retrieval space, the technical content constraint controls the focus of the retrieval topic, and the specification severity constraint controls the selection priority of different grade branches under the same specification. To avoid mutual cancellation or weight drift between different constraints, the combined constraint conditions are represented as a weighted fusion constraint score, and the score result is used to drive the subsequent candidate screening and ranking for specification positioning. This ensures that the combined constraint conditions maintain a stable constraint effect on the semantics of the same target engineering project throughout the retrieval process. The calculation formula is as follows:
[0040] in, This represents the combined constraint score; a larger value indicates a better match between the candidate specification clause and the combined constraint condition. , , This represents the fusion weight of the three types of constraints. The value range is a real number greater than 0 and can be normalized according to business experience or validation data. The semantic representation of the project category can be obtained from the standardized encoding or vector representation of the project category name in the domain thesaurus; The semantic representation of the detected item can be obtained from the vector representation formed by the name of the detected item and its synonym merging results; The semantic representation of the engineering level can be obtained from the vector representation formed by the level label and the level interval mapping result; The semantic representation of the label indicating the scope of application of the specification. The semantic representation of the subject identifier of the normative clause. The semantic meaning of the level restriction clause; The similarity function measures the closeness of the semantic representations on both sides and outputs a similarity value. Vector cosine similarity, edit distance normalized similarity, or hybrid similarity can be used to simultaneously consider terminology consistency and semantic closeness. The above formula calculates the similarity between the three types of constraints and their corresponding knowledge base tags and then performs weighted fusion. This ensures that the combined constraints are subject to the joint constraints of the scope of application, topic content, and level of severity during the retrieval process, thereby reducing the probability of mismatches caused by a single constraint.
[0041] When performing specification location processing in the standard specification knowledge base based on combined constraints, three types of indexes are pre-established for each standard clause in the knowledge base: a specification scope tag index, a specification clause subject identifier index, and a level limitation clause index. During retrieval, the candidate specification domain set is first coarsely screened based on the specification scope constraint, then the subject identifiers of the candidate standard clauses are matched for subject consistency based on the technical content constraint, and finally the level limitation clauses in the candidate standard clauses are matched for level consistency based on the specification stringency constraint. This ensures that the candidate standard clauses must simultaneously satisfy the engineering category, testing item, and engineering level. Only those three types of consistency conditions are included in the standard clause set. Among them, the standard positioning process emphasizes consistency matching rather than just keyword matching. Consistency matching is used to ensure that the engineering category and the standard scope label correspond to each other under the same engineering classification system, the testing item and the standard clause theme identifier correspond to each other under the same testing object and testing purpose, and the engineering level and the level limitation clause correspond to each other under the same level semantic branch. Furthermore, the candidate results can be sorted and truncated by combining the aforementioned combined constraint scores, so as to retain the standard clause set that is closest to the combined constraint conditions under the premise of controllable candidate size, thereby providing a stable and interpretable standard basis range for subsequent association analysis processing.
[0042] When performing association parsing on the standard clause set, each standard clause is treated as a structured carrier containing a standard basis section, a testing method section, and an acceptance requirement section. By jointly parsing the internal citation markers, clause number levels, key verb patterns, and result judgment sentences of the clause, the testing methods corresponding to the testing items are identified, and their implementation conditions, operational points, and data recording requirements are extracted. At the same time, the acceptance requirements corresponding to the testing methods are identified, and their threshold conditions, pass / fail judgment rules, and re-inspection trigger conditions are extracted. The goal of association parsing is to maintain the traceable citation relationship between the testing methods and acceptance requirements to the same standard clause, avoiding cross-splitting where the testing method comes from one clause and the acceptance requirement comes from another. The testing method is used to represent the testing implementation path, the acceptance requirement is used to represent the result judgment criteria, and the standard clause is used to represent the source of the standard basis. By retaining the clause number, chapter number, and citation chain marker during parsing, each testing method and acceptance requirement is bound to a unique standard clause identifier, thereby maintaining a complete closed loop of standard citation in subsequent semantic consistency verification and automatic assembly processing.
[0043] When performing semantic consistency verification on standard clauses, testing methods, and acceptance requirements based on a structured information set, the semantics of engineering categories, testing items, and engineering levels in the structured information set are aligned and verified with the semantic representations of candidate standard clauses, testing methods, and acceptance requirements. The verification results are used to eliminate candidate elements with conflicting scopes of application, drifting testing objects, or inconsistent level branches, thus forming a set of specification elements that can be directly used to generate testing scheme content. The semantic consistency verification process not only verifies text similarity but also verifies the consistency of constraint logic. For example, if the engineering level in the structured information set points to a high-severity branch, the acceptance requirements must not fall into the low-severity threshold set; if the testing item in the structured information set specifies the testing object, the testing method must not shift to other objects or other testing purposes. To quantify the verification results and support threshold filtering, consistency is represented as a multi-dimensional consistency score, and a score not lower than a preset threshold is used as a condition for entering the specification element set. The calculation formula is as follows:
[0044] in, This represents the semantic consistency score; a higher value indicates a greater consistency between the candidate specification elements and the structured information set. , , , This represents the weighting coefficient, which is a real number greater than 0 and can be set according to business verification. The semantic representation of a standard provision can be generated jointly from the provision's subject, scope of application, and the context of its number. The semantic representation of the detection method can be generated by combining the method name, key operation descriptions, and applicable conditions. The semantic representation of acceptance requirements can be generated by combining judgment statements, threshold descriptions, and level limitation descriptions. This indicates a logical conflict penalty term. This represents a conflict set, used to collect conflict types such as scope conflicts, object inconsistencies, hierarchical branch inconsistencies, and broken reference chains. The formula can be accumulated based on the number and severity of conflicts. By weighting and summing the three types of semantic alignment similarity and subtracting the logical conflict penalty, the verification can reflect both the semantic closeness and the consistency of the constraint logic, thereby reducing the risk of misselection due to "semantic similarity but logical inconsistency" caused by relying solely on similarity.
[0045] When automatically assembling a set of standard elements to generate test plan content, the set of standard elements is first aggregated according to the dimension of test items. Within each aggregated unit of a test item, the standard clauses are used as the standard basis, the test methods are used as the test implementation content, and the acceptance requirements are used as the result judgment content for unified arrangement. At the same time, the engineering category, engineering level, and engineering attribute results in the structured information set are written into the corresponding positions of the test plan content as global context constraints, so that the test plan content maintains semantic and reference consistency with the structured information set at the expression level. Among them, the automatic assembly process emphasizes unified arrangement rules and traceable reference relationships. Unified arrangement rules are used to ensure that different test items use the same structural expression in the plan, which facilitates the parsing and review in the subsequent approval and execution stages. Traceable reference relationships are used to ensure that each test implementation content and result judgment content can be traced back to its corresponding standard clause identifier and knowledge base version identifier, thereby supporting difference positioning and plan regeneration when the standard specification knowledge base is updated. By maintaining the consistency of terminology, explanation order, and reference chain during the assembly process, the generated test plan content can directly meet the rule judgment requirements of automatic configuration of subsequent approval nodes and adaptive matching of test resources.
[0046] S140, based on the content of the testing plan and the engineering attribute results determined in the structured information set, determines the automatic configuration of approval nodes and the adaptive matching of testing personnel or testing equipment by associating the engineering level with the approval process rules and the testing project type with the testing resource rules.
[0047] In one possible implementation, based on the content of the testing plan and the engineering attribute results determined in the structured information set, the system associates the engineering level with approval process rules and the testing item type with testing resource rules to achieve automatic configuration of approval nodes and adaptive matching of testing personnel or equipment. Specifically, this includes: performing engineering attribute mapping processing on the testing plan content based on the engineering attribute results, semantically aligning the testing items, testing methods, and acceptance requirements in the testing plan content with the engineering category, engineering level, and testing item attributes in the structured information set, so that each testing unit in the testing plan content is bound to a corresponding engineering attribute identifier; performing approval process rule association determination processing based on the engineering level, matching the engineering level with the level range or level label in the preset approval process rules; and combining the risk attributes and mandatory attributes of the testing items in the testing plan content. The approval process rules are jointly evaluated to determine the number of approval levels, types of approval nodes, and approval sequence constraints, generating an approval node configuration result that matches the project level. Based on the testing item types determined in the testing plan and the testing item attributes in the structured information set, the testing resource rules are correlated and evaluated. This involves matching the testing item types with the predefined personnel qualification conditions, equipment capability parameters, and resource occupancy constraints in the testing resource rules, while simultaneously verifying the project level's requirements for resource level or accuracy level, to obtain a set of candidate testing personnel or a set of candidate testing equipment. Adaptive filtering is then performed on the candidate set of testing personnel or equipment. By combining the requirements in the testing plan, the qualification level, historical execution records, and current availability of the candidate resources are comprehensively evaluated to determine the testing personnel or testing equipment configuration result that is consistent with the testing plan content and project attribute results.
[0048] Specifically, when performing engineering attribute mapping processing on the content of the testing plan, testing units are first established for the testing items, testing methods, and acceptance requirements in the testing plan content. The semantic correspondence between the testing units and the structured information set is then described within the same engineering testing semantic space. This is achieved by aligning testing items with their thematic consistency, aligning testing methods with the applicable context of the specifications defined by the engineering category, and aligning acceptance requirements with the severity branch defined by the engineering level. This ensures that each testing unit is bound to a unique engineering attribute identifier. The engineering attribute results are used to characterize the target project in terms of engineering category, engineering level, and other relevant parameters. In addition to global constraints on the attributes of the detection items, engineering attribute mapping is used to project these global constraints onto each detection unit. Semantic alignment is used to ensure that the terminology, object orientation, and level branches of the detection units do not drift. Engineering attribute identifiers are used to record the binding relationships between detection units and engineering categories, engineering levels, and detection item attributes in a structured manner, so that subsequent approval process rule association judgments and detection resource rule association judgments can be carried out on the same engineering attribute identifier and maintain consistency between the two. To quantify the semantic alignment effect and set filtering conditions for the mapping confidence, the engineering attribute mapping confidence is represented as a multi-source similarity fusion, and the calculation formula is as follows:
[0049] in, This represents the confidence level of the engineering attribute mapping; the larger the value, the more reliable the binding between the detection unit and the engineering attribute result. , , represents the weighting coefficient, which is a real number greater than 0 and can be normalized according to the focus of engineering testing business; P represents the semantic representation of the testing item attribute, which can be generated by the testing item name, synonym merging results and testing object description; C represents the semantic representation of the engineering category, which can be generated by the mapping result of the engineering category name and the scope of application of the standard; G represents the semantic representation of the engineering level, which can be generated by the level label, level interval mapping result and severity branch identifier. This represents the semantic representation of the detection items in the detection unit. The semantic representation of the detection method in the detection unit. The semantic representation of the acceptance requirements in the testing unit; This represents a semantic similarity function that outputs a similarity value by comparing the closeness of semantic representations. The formula simultaneously constrains the three types of elements—detection items, detection methods, and acceptance requirements—to align with the detection item attributes, project categories, and project levels, respectively. This ensures that the mapping of project attributes relies not only on matching a single field but also on the overall consistency of the detection unit, thereby reducing the chain reaction of deviations caused by mapping errors in subsequent process configurations.
[0050] When performing approval process rule association judgment based on project level, the preset approval process rules are represented as a set of searchable rule entries, and a level range or level label is configured for each rule entry, so that the project level can be directly matched with the triggering conditions of the rule entry. Among them, the level range is used to express the continuous hierarchical relationship of the project level, the level label is used to express the discrete branch relationship of the project level, and the approval process rule is used to express the triggering conditions and constraints that the approval node configuration needs to meet. By mapping the project level to the level range or level label and retrieving the corresponding rule entry, an initial rule hit set is formed, which serves as the basis for subsequent joint judgment of the number of approval levels, approval node types, and approval order constraints. This ensures that the approval process rule association judgment has a definite rule entry in the project level dimension and maintains the level semantics consistent with the structured information set.
[0051] When jointly determining the approval process rules by combining the risk attributes and mandatory attributes of the testing items in the testing plan, the risk attributes and mandatory attributes are converted into calculable risk factors and mandatory factors, respectively. Based on the initial rule hit set, the number of approval levels, approval node types, and approval sequence constraints in the approval process rules are further screened and parameterized. This ensures that the approval node configuration simultaneously satisfies the level constraints of the engineering grade, the risk constraints of the risk attributes, and the compliance constraints of the mandatory attributes. Specifically, the risk attributes describe the degree of impact of the testing items on safety consequences, quality risks, or construction risks; the mandatory attributes describe whether the testing items correspond to mandatory provisions or mandatory acceptance conditions; the number of approval levels represents the depth of the approval decision-making chain; the approval node types represent the roles involved in the approval; and the approval sequence constraints represent the sequential dependencies between different role nodes. To achieve interpretable and adjustable joint determination, the approval intensity is represented as a comprehensive approval score, which is mapped to the number of approval levels and the set of approval node types. Simultaneously, the node dependencies are solidified by the approval sequence constraints. The calculation formula is as follows:
[0052] in, This represents the overall approval score; a higher value indicates a greater level of approval required. , , This represents the weighting coefficient, which is a real number greater than 0 and can be set according to organizational management strategies. R represents the engineering grade mapping function, which is obtained by mapping the grade label or grade range of the engineering grade to the grade score. The higher the grade, the larger the corresponding score. R represents the risk attribute set, which can be formed by combining factors such as the risk category of the inspection project, the probability of historical defects, and the complexity of on-site conditions. The risk score function is obtained by weighted aggregation of the risk attribute set; M represents the set of mandatory attributes of the standard, which can be formed by merging mandatory clause identifiers, mandatory acceptance condition identifiers, and mandatory re-inspection trigger identifiers. The mandatory score function is obtained by weighted aggregation of the set of mandatory attributes. This formula maps the project level, risk attributes and mandatory attributes of the standard to a comprehensive approval score, so that the configuration result of the approval node is no longer determined by the single factor of project level. Instead, it automatically increases the approval depth and approval role coverage for high-risk or mandatory testing items while maintaining the consistency of the level, thereby improving the reliability of the matching between process configuration and project attributes.
[0053] When performing association determination on the testing resource rules to obtain a set of candidate testing personnel or candidate testing equipment, the testing project types in the testing plan content and the testing project attributes in the structured information set are first merged to obtain a testing project type identifier with a unique semantic boundary. The matching relationship between the testing personnel qualification conditions, equipment capability parameters, and resource occupancy constraints and the testing project type is then checked in the testing resource rules. Simultaneously, the engineering level's limitation requirements on the resource level or accuracy level are introduced as an additional filtering condition, thereby filtering and forming a set of candidate testing personnel or candidate testing equipment from the resource pool. The testing project type is used to abstract the methodological category attribution of the testing project, such as non-destructive testing. The test type, entity sampling type, or on-site monitoring type; personnel qualification requirements; equipment capability parameters; and resource availability constraints. These constraints limit the types of certificates, authorization scope, and training records that testing personnel must possess. Equipment capability parameters define the capability boundaries that testing equipment must meet, such as the range of objects it can cover, the set of supported testing methods, and the achievable measurement accuracy levels. Resource availability constraints limit resource availability within time windows, geographical areas, and concurrent tasks. Resource level or accuracy level constraints ensure that the resource capabilities corresponding to high-level projects are not lower than a preset threshold. To quantify the matching degree between candidate resources and testing project types and support sorting selection, the resource matching score is represented as a fusion of multi-condition consistency and availability. The calculation formula is as follows:
[0054] in, This represents the resource matching score; a higher value indicates a better match between the candidate resources. , , , This represents the weighting coefficient, which is a real number greater than 0 and can be set according to resource management strategies; T represents the testing item type identifier; Q represents the set of personnel qualification conditions. This represents a consistency function between the type of testing item and the personnel qualifications. It outputs a matching value by determining whether the personnel qualifications cover the qualification range required by the type of testing item; A represents the set of equipment capability parameters. This function represents the consistency between the test item type and the equipment capability parameters. It outputs a matching value by determining whether the equipment capability parameters cover the capability boundaries required by the test item type; G represents the engineering level, and L represents the resource level or accuracy level. This represents the level adaptation function, which determines whether the resource level or precision level meets the minimum threshold corresponding to the project level and outputs the adaptation value according to the degree of excess; U represents the set of resource occupancy states. The occupancy penalty function is obtained by aggregating the proportion of resources occupied within the target time window, the number of conflicting tasks, and the geographical scheduling cost to obtain the penalty value. This formula incorporates qualification coverage, capability coverage, level adaptation, and availability occupancy into the score, so that the candidate set of testing personnel or candidate set of testing equipment not only meets the hard rule filtering, but also forms an interpretable ranking basis among multiple feasible resources.
[0055] When performing adaptive screening on a set of candidate testing personnel or equipment, the requirements in the testing plan are abstracted into a requirement vector, and the qualification level, historical execution record, and current availability status of candidate resources are abstracted into a resource profile. By comprehensively judging the requirement vector and resource profile, the configuration of testing personnel or equipment consistent with the testing plan content and engineering attributes is determined. Specifically, the qualification level characterizes the upper limit of a candidate resource's capabilities within its qualification coverage and level; the historical execution record characterizes the candidate resource's completion quality, anomaly handling capabilities, and compliance stability under the same or similar testing project types; and the current availability status characterizes the candidate resource's schedulability and conflict risk within the target time window. Adaptive screening is the process of simultaneously meeting hard thresholds and optimizing multiple objective indicators within the candidate set. To ensure the comprehensive judgment is repeatable and adjustable, the final selection is represented as maximizing the comprehensive adaptation score, and the contribution of qualification level, historical execution record, and current availability status is explicitly reflected in the score. The calculation formula is as follows:
[0056] in, This indicates an adaptive filtering score; a higher value indicates a better candidate resource. , , , This represents the weighting coefficient, which is a real number greater than 0 and can be set according to project management strategies. The qualification level score is obtained by comparing the qualification level of candidate resources with the minimum qualification threshold required for the type of testing project and mapping the excess to a score; H represents the set of historical execution records. The historical performance score is obtained by weighting and aggregating indicators such as on-time rate, first-time pass rate, re-inspection trigger rate, and compliance audit pass rate. The availability score is calculated by aggregating the proportion of available time slots, the number of scheduling conflicts, and geographical accessibility of candidate resources within the target time window; Z represents the set of risk factors. The risk penalty item is calculated by accumulating risk factors such as recent abnormal events, near-expiration of qualifications, near-expiration of equipment calibration, or maintenance alarms for candidate resources. This formula further maximizes the comprehensive adaptation score while meeting the rule filtering requirements, so that the output configuration results of testing personnel or testing equipment can simultaneously meet the implementation requirements of the testing plan, the stringency requirements of the engineering level, and the availability constraints of resource management. This achieves a closed-loop linkage between automatic configuration of approval nodes and adaptive matching of testing resources on the same semantic basis.
[0057] Furthermore, in testing scenarios where engineering categories have a hierarchical nesting relationship, the same engineering project often simultaneously meets the judgment conditions of both the main engineering category and the special engineering category. Different levels of engineering categories correspond to independent and partially overlapping testing clauses, testing methods, and acceptance requirements in the standard specification system. If the hierarchical relationship and applicable boundaries of engineering categories are not uniformly modeled and semantically constrained, the testing scheme generation process may easily rely solely on a single category for specification matching, resulting in the omission of mandatory testing clauses corresponding to special engineering projects, or the problem of duplicate reference of testing clauses when the main engineering and special engineering specifications are in effect simultaneously. This leads to risks in the completeness, consistency, and executability of the testing scheme, and this problem is particularly prominent when the engineering structure is complex, the number of specification clauses is large, or the specifications are frequently updated.
[0058] In one possible implementation, based on text data, multi-stage information extraction processing is performed according to a preset information extraction order. Specifically, this includes: attribute classification, which involves performing multi-level engineering category determination processing on the engineering category when the same text data simultaneously meets the determination conditions of the main engineering category and the special engineering category, so as to simultaneously identify the main engineering category and the special engineering category; establishing a hierarchical relationship between the main engineering category and the special engineering category, so that the engineering category forms a multi-level engineering category expression; and forming a multi-level structured information set together with the detection items and engineering level.
[0059] Specifically, in attribute classification processing, when semantic clues pointing to both the main project category and the special project category appear simultaneously in the same text data, a multi-level project category determination process is performed on the project category. By performing parallel determination on the project description fragments in the semantic space of the engineering inspection domain, the main project category and the special project category can be identified synchronously in the same text data. The main project category is used to represent the higher-level classification attributes of the project at the overall structure and function level, while the special project category is used to represent the lower-level classification attributes of the project at the local structure, special process, or specific risk dimension. The multi-level project category determination is used to avoid covering all engineering inspection semantics with only a single project category. By independently matching and jointly verifying different engineering semantic triggering conditions in the text, both the main project category and the special project category are retained as valid project category results, thereby preventing the special project inspection requirements from being implicitly ignored in the subsequent specification matching process.
[0060] After completing the synchronous identification of the main project category and the special project category, a hierarchical relationship between the main project category and the special project category is established. By setting a superior node identifier for the main project category and a subordinate node identifier for the special project category in the project category system, the two form a clear parent-child relationship or inclusion relationship in the same project category hierarchy. The hierarchical relationship is used to express the coverage and constraint boundary of the main project category to the special project category, avoiding the introduction of specification conflicts or duplicate matching by treating different levels of project categories as peer categories. The hierarchical relationship clarifies which testing requirements should be uniformly applied by the main project category and which testing requirements are only triggered when the conditions of the special project category are met, thus providing an interpretable hierarchical semantic basis for subsequent standard specification positioning and testing clause assembly.
[0061] After establishing the hierarchical relationship between engineering categories, the multi-level engineering category expression is jointly structured with the testing items and engineering grades. By simultaneously retaining the main engineering category nodes, special engineering category nodes, and their hierarchical relationships in the structured information set, and binding the testing items with the corresponding engineering category hierarchical nodes, while introducing the engineering grade as a cross-level global constraint attribute, the engineering categories, testing items, and engineering grades form a multi-level, inheritable, and definable structured information set. The multi-level structured information set is used to support hierarchical expansion or hierarchical constraint processing strategies in the subsequent specification matching and scheme generation process, so that the general testing requirements corresponding to the main engineering category and the special testing requirements corresponding to the special engineering category can coexist in the same testing scheme without conflict, thereby achieving complete coverage and semantic consistency of the testing scheme content in scenarios where engineering categories have hierarchical nesting relationships.
[0062] In one possible implementation, a hierarchical relationship is established between the main project category and the special project category, enabling the project categories to form a multi-level project category expression. Specifically, this includes: performing hierarchical role determination processing on each project category in the candidate set of project categories based on a preset hierarchical rule set to determine the hierarchical role of the main project category or special project category corresponding to each project category, and generating a corresponding hierarchical role identifier for each project category; based on the hierarchical role identifier and combined with the semantic relationship between project-related entities in the text data, performing hierarchical association matching processing on the main project category and the special project category, and establishing a hierarchical relationship between the main project category and the special project category when the preset hierarchical association conditions are met by comprehensively judging the scope of the project object, the functional orientation, and the consistency of the detection object; performing hierarchical structured encapsulation processing on the main project category and the special project category with the established hierarchical relationship as the lower-level project category node, and recording the membership identifier, hierarchical depth identifier, and association constraint identifier between project categories in the multi-level structured information set, so that the project categories form a multi-level project category expression.
[0063] Specifically, when performing hierarchical role determination processing on each project category in the project category candidate set based on the preset project category hierarchy rule set, each project category in the project category candidate set is first mapped to the category space defined by the project category hierarchy rule set. Then, a consistency check is performed on the project category's higher-level coverage, lower-level subdivision features, and triggering semantic cues in the category space to determine the hierarchical role corresponding to the project category. The project category candidate set originates from the merged results of named entities related to project categories after named entity recognition and relation extraction. The project category hierarchy rule set is used to solidify the determination conditions of higher-level and lower-level categories in the project category system in a rule-based manner. The hierarchical role is used to distinguish project categories in multi-level project category expressions. In this functional positioning, the hierarchical role of the main project category is used to represent the higher-level category role that can cover the overall semantics of the project, while the hierarchical role of the special project category is used to represent the lower-level category role that only exists when specific structural, technological, or risk conditions are met. To ensure that the hierarchical role determination results can be directly referenced by subsequent hierarchical association matching processing, a hierarchical role identifier is generated for each project category. The hierarchical role identifier is used to record the role type, role confidence, and rule hit basis of the main project category or special project category, so that the role does not drift and terminology consistency is maintained in subsequent processing for the same project category. When it is necessary to make a joint decision among the rule hit count, rule weight, and semantic similarity, the hierarchical role confidence is represented as a fusion score, and the calculation formula is:
[0064] in, This represents the confidence level of the hierarchical role; a higher value indicates a more reliable determination of the hierarchical role for the project category. 'n' represents the number of rule entries involved in the determination within the hierarchical rule set for the project category. This indicates whether the i-th rule was hit; the value is 1 if the rule was hit and 0 if the rule was not hit. The weight of the i-th rule is used to reflect the contribution of the rule to the determination of the role at the level of the main project category or the role at the level of the special project category. The value range is a real number greater than 0. The semantic representation of the project category in the text data can be generated by the context window embedding vector or normalized encoding of the project category named entity; The semantic representation of the rule semantic template can be jointly generated from the set of trigger phrases, function-oriented descriptions, and coverage descriptions defined in the rule entries; This represents a semantic similarity function that outputs a similarity value by measuring the closeness of the semantic representations on both sides. The weight coefficient of the semantic similarity term is represented by a real number greater than 0. This formula integrates rule hit information with semantic similarity information, so that the determination of hierarchical roles depends on interpretable rule hit criteria and can remain robust when there are variations in text expression.
[0065] When performing hierarchical association matching processing on the main project category and the special project category based on hierarchical role identification and combined with the semantic relationship between project-related entities in the text data, the following steps are taken: First, locate entities such as project objects, project categories, functional descriptions, structural parts, and detection objects in the semantic relationship between project-related entities. Then, bind the main project category and the special project category to the project object scope and functional orientation with the strongest semantic support, respectively. Then, comprehensively judge the consistency of project object scope, functional orientation, and detection objects to confirm whether the superior coverage relationship of the main project category to the special project category is valid. Among them, the semantic relationship is formed in the relationship extraction stage and is used to express the limiting relationship between project objects and project categories, the inclusion relationship between project objects and structural parts, and the pointing relationship between structural parts and detection objects. The project object scope is used to describe the spatial range or component range covered by the project category, and the functional orientation is used to describe the... The engineering function or process purpose corresponding to the engineering category, and the consistency of the detection object are used to determine whether the detection object pointed to by the special engineering category falls within the scope of the engineering object defined by the main engineering category and does not conflict with the functional context of the main engineering category; the hierarchical association condition is used to constrain when a hierarchical association relationship is allowed to be established, such as requiring that the scope of the engineering object of the special engineering category is covered by the scope of the engineering object of the main engineering category, that the function of the special engineering category is a sub-branch of the function of the main engineering category, and that the detection object of the special engineering category does not drift across domains with the detection object of the main engineering category. When the preset hierarchical association conditions are met, a hierarchical association relationship is established between the main engineering category and the special engineering category, and the chain of evidence entities of the association is recorded to support subsequent auditing and retrospection; in order to make the comprehensive judgment quantifiable and support threshold triggering, the hierarchical association matching score is represented as a fusion of coverage, orientation consistency and detection object consistency, and the calculation formula is:
[0066] in, This represents the hierarchical association matching score. The higher the value, the higher the credibility of establishing a hierarchical association between the main project category and the special project category. , , , This represents the weighting coefficient, and its value ranges from a real number greater than 0. The scope of engineering objects corresponding to a specific engineering category can be generated by merging the associated set of engineering object entities, the set of structural parts entities, and the scope limitation description. This represents the scope of engineering objects corresponding to the main engineering category, and its generation method is the same as... Consistent; This represents the coverage function, which outputs a coverage value by measuring the degree to which the scope of a specific project category is included by the scope of the main project category. The function of indicating the category of a specific project points to its semantic representation. The function of indicating the main project category points to the semantic representation. Indicates the similarity in functional orientation; This represents the set of testing objects associated with a specific project category. This represents the set of inspection objects associated with the main project category. This represents a consistency function for detecting objects, which outputs a consistency value by judging the inclusion relationship, synonym merge consistency, and object domain consistency among sets of objects. This represents a conflict set, used to collect conflict types such as cross-domain object conflicts, function pointer conflicts, and scope uncontainment conflicts. The conflict penalty function is defined as the penalty value obtained by accumulating the number and severity of conflict sets. This formula, by penalizing conflicts on the basis of satisfying coverage and consistency, enables hierarchical association matching to stably distinguish between true hierarchical relationships and superficial co-occurrence relationships in complex text expressions.
[0067] When performing hierarchical structured encapsulation on main project categories and special project categories, the main project category is encapsulated as a higher-level project category node, and the special project categories with established hierarchical relationships are encapsulated as lower-level project category nodes. Simultaneously, the hierarchical structured information set records the membership identifiers, hierarchical depth identifiers, and association constraint identifiers between project categories, enabling the project categories to form a multi-level project category expression that can be directly utilized for subsequent specification location and deduplication / merging processes. The higher-level project category node carries the global coverage semantics of the main project category and serves as the parent node of the lower-level project category nodes. The membership identifier clarifies that the lower-level project category nodes belong to the higher-level project category. The node's hierarchical direction and relationship type, and the hierarchical depth identifier are used to record the hierarchical distance between lower-level engineering category nodes and upper-level engineering category nodes, and support multi-level nesting expansion. The association constraint identifier is used to record the coverage conditions, functional orientation conditions, and consistency conditions of the test object formed in the hierarchical association matching process. This enables subsequent standard specification knowledge base matching to distinguish between general test requirements triggered by the main engineering category and special test requirements triggered by the special engineering category according to the association constraint identifier. When generating test plan content, conflict resolution and duplication suppression are performed on the reference of clauses of different engineering categories, thereby maintaining complete coverage of test clauses and avoiding duplicate reference of test clauses in the scenario of hierarchical nesting of engineering categories.
[0068] In one possible implementation, the engineering category, testing items, and engineering grade of the structured information set are used as combination constraints. The corresponding standard clauses, testing methods, and acceptance requirements are located in the standard and specification knowledge base. While maintaining the semantic consistency of the structured information, the testing scheme content is automatically assembled and generated. Specifically, this includes: based on the multi-level structured information set, performing hierarchical expansion processing on the multi-level engineering category expressions; constructing an engineering category hierarchical tree structure by parsing hierarchical relationships, the engineering category hierarchical tree structure includes upper-level engineering category nodes and at least one lower-level engineering category node; based on the engineering category hierarchical tree structure, using the engineering category, testing items, and engineering grade as combination constraints, performing hierarchical retrieval processing on the standard and specification knowledge base, where basic standard clauses, general testing methods, and overall acceptance requirements are located for upper-level engineering category nodes, and specific testing clauses, supplementary testing methods, and specific acceptance requirements are located for lower-level engineering category nodes; and performing hierarchical retrieval processing on the main engineering category corresponding to the hierarchical relationships. The testing requirements and the testing requirements corresponding to the specific project categories are processed using hierarchical constraints. Testing requirements corresponding to the main project category are used as higher-level constraints, and testing requirements corresponding to the specific project category are used as lower-level supplementary constraints. Consistency constraints are applied to testing frequency, testing depth, and acceptance criteria based on the project level to form a set of testing requirements. This set of testing requirements undergoes semantic consistency verification. It is matched item by item with the scope of project objects, semantics of testing items, and project level limitations in the multi-level structured information set. Testing requirements inconsistent with project attributes are eliminated, and semantically overlapping testing requirements are uniformly referred to and merged. After semantic consistency verification, based on the hierarchically constrained and semantically consistent set of testing requirements, the testing requirements are structured and arranged according to the testing order, testing method adaptation relationship, and acceptance logic relationship, thereby generating a testing scheme content consistent with the multi-level project category expression, testing items, and project level.
[0069] Specifically, when performing hierarchical expansion processing on the expression of multi-level engineering categories based on a multi-level structured information set, the upper-level and lower-level engineering category nodes are first located in the multi-level structured information set, and the membership identifier, hierarchical depth identifier, and association constraint identifier are read. By performing directional consistency verification on the membership identifier and hierarchical merging according to the hierarchical depth identifier, a hierarchical tree structure of engineering categories is constructed. Among them, the hierarchical expansion processing is used to restore the engineering categories that originally existed in the form of discrete nodes into a traversable hierarchical structure. The hierarchical tree structure of engineering categories is used to express the superior coverage relationship of the main engineering category to the special engineering category in a tree topology. The upper-level engineering category node is used to carry the semantics of the main engineering category and serve as the root node or superior parent node. The lower-level engineering category node is used to carry the semantics of the special engineering category and serve as the child node or subordinate node. The hierarchical association relationship is used to define the parent-child connection and applicable boundary between nodes. At least one layer of lower-level engineering category nodes is used to ensure that the hierarchical structure can express the introduction position of the special constraint, thereby providing a clear category traversal order and constraint inheritance path for subsequent hierarchical retrieval processing.
[0070] When performing hierarchical retrieval processing on the standard and specification knowledge base based on the hierarchical tree structure of engineering categories, engineering categories, testing items, and engineering levels are collectively represented as combined constraints. The hierarchical traversal results of the engineering category hierarchical tree structure are used as the retrieval layering strategy. Basic searches are performed on upper-level engineering category nodes to locate basic standard provisions, general testing methods, and overall acceptance requirements. Specialized searches are performed on lower-level engineering category nodes to locate specialized testing provisions, supplementary testing methods, and specialized acceptance requirements. The hierarchical retrieval processing emphasizes that the retrieval targets for the same combined constraint differ across engineering category nodes at different levels. The basic standard provisions corresponding to the upper-level engineering category nodes are used for coverage... The system covers the overall general requirements of the project, general testing methods to cover the main testing implementation path, overall acceptance requirements to cover the overall judgment logic of the project, and specific testing clauses corresponding to the lower-level project category nodes to cover the specific requirements of specific structural parts or specific processes. Supplementary testing methods are used to supplement specific implementation paths that cannot be covered by general testing methods, and specific acceptance requirements are used to supplement the judgment branches of overall acceptance requirements in specific scenarios. To avoid redundancy caused by the same clause hitting the same level repeatedly, each search result is bound to its generation level, and the level attribution arbitration is performed on clauses that hit the same level repeatedly, so that the same standard element maintains a unique and traceable attribution in the project category hierarchical tree structure.
[0071] When performing hierarchical constraint combination processing on the testing requirements corresponding to the main project category and the testing requirements corresponding to the special project category based on the hierarchical association relationship, the basic standard clauses, general testing methods, and overall acceptance requirements retrieved from the upper-level project category node are first abstracted into the testing requirements corresponding to the main project category. Similarly, the special testing clauses, supplementary testing methods, and special acceptance requirements retrieved from the lower-level project category node are abstracted into the testing requirements corresponding to the special project category. Then, through the hierarchical association relationship, the testing requirements corresponding to the main project category are set as upper-level constraints, and the testing requirements corresponding to the special project category are set as lower-level supplementary constraints. This ensures that the lower-level supplementary constraints only apply to the project object scope and functional requirements defined by the association constraint identifier. The conditions take effect immediately upon application. The higher-level constraints provide basic testing boundaries and minimum compliance requirements, while the lower-level supplementary constraints introduce additional testing content and stricter judgment branches for specific scenarios. The engineering level provides cross-level consistency constraints in the combination of hierarchical constraints. By inheriting and propagating the engineering level constraints and applying them to the higher-level and lower-level supplementary constraints, the testing frequency, testing depth, and acceptance judgment conditions are unified, avoiding inconsistencies in judgment standards caused by higher-level constraints using low-severity branches while lower-level supplementary constraints use high-severity branches. When it is necessary to perform level-consistency conversion on testing frequency and testing depth, the level consistency coefficient is introduced into the parameterized expression of the testing requirements, and the calculation formula is as follows:
[0072] in, This represents the level consistency coefficient; a larger value indicates that the detection requirements need a higher degree of enhancement in terms of frequency or depth. This represents the strengthening coefficient, which is a real number greater than 0 and is used to control the influence of the engineering grade on the strengthening range required for testing. The ordinal mapping value representing the project level is obtained by mapping the project level label to an ordered ordinal number; This represents the lowest level in the engineering grade system. This represents the highest level in the engineering grade system. The formula linearly maps the relative position of the engineering grade in the grade system to a strengthening coefficient, so that the upper-level constraints and lower-level supplementary constraints in the hierarchical constraint combination processing can be uniformly adjusted under the same strengthening scale, thereby maintaining the consistency of the stringency across levels when forming the set of testing requirements.
[0073] When performing semantic consistency verification on the set of detection requirements, each detection requirement is decomposed into engineering object scope constraints, detection item semantic constraints, and engineering level limitations. These are then matched item by item with the engineering object scope, detection item semantics, and engineering level limitations in the multi-level structured information set. Detection requirements inconsistent with engineering attributes are eliminated based on the matching criteria. Item-by-item matching verifies whether a detection requirement falls within the engineering object scope, corresponds to the identified detection item semantics, and meets the engineering level limitations. Inconsistent engineering attributes indicate that the detection requirement references an inapplicable object domain, deviates from the target detection item semantics, or uses a mismatched level branch. Detection requirements with semantic overlap are uniformly referred to and combined. During processing, the detection objects, detection methods, and acceptance criteria in the detection requirement set are first synonymized and aligned with the reference chain. Semantic overlap refers to duplicate requirements formed by different sources of clauses when the detection objects are the same and the acceptance criteria are equivalent or containment relationships are established. Unified referencing is used to generate a unique detection requirement identifier for overlapping detection requirements and retain their multi-source specification reference evidence. Merging processing is used to merge equivalent requirements into a single detection requirement and merge containment relationship requirements into a higher-level requirement plus a lower-level difference item, thereby eliminating the redundancy caused by duplicate detection clauses while ensuring traceability of references. When it is necessary to quantify the degree of semantic overlap to trigger the merging threshold, the overlap is expressed as a set similarity and combined with the equivalence of the judgment criteria for adjudication. The calculation formula is as follows:
[0074] in, This indicates the semantic overlap of the detection requirements; a larger value means that the two detection requirements should be uniformly referred to and merged. , , This represents the weighting coefficient, and its value ranges from a real number greater than 0. and The sets of engineering object ranges representing the two testing requirements can be formed by merging the sets of engineering object entities and the sets of structural part entities. The Jaccard similarity, representing the scope of engineering objects, is obtained by calculating the ratio of the intersection size to the union size; M and These represent the semantic representations of the detection methods for the two detection requirements, respectively. Indicates the similarity between detection methods; A and These represent the acceptance criteria for the two testing requirements. This represents the equivalence function of the decision condition, which outputs the equivalent value by judging whether the threshold expression is consistent, whether the decision logic is consistent, and whether the level limitation branch is consistent. This formula, by simultaneously measuring the equivalence of the object scope, method semantics, and decision condition, enables the merging process to avoid erroneously merging detection requirements with different detection purposes simply because of object overlap.
[0075] After semantic consistency verification, when performing structured orchestration on the set of detection requirements that are hierarchically constrained and semantically consistent, the set of detection requirements is first sorted according to the detection order. By placing detection requirements with pre-dependencies before those with post-dependencies, the set of detection requirements forms a feasible sequence in the execution path. Then, based on the detection method adaptation relationship, detection requirements are bound to detection methods one-to-one or one-to-many, ensuring that each detection requirement receives a detection method configuration consistent with its detection object scope and engineering level constraints. Here, the detection order expresses the logical sequence of detection tasks in the field, the detection method adaptation relationship expresses the applicability constraints of the detection method to the detection object scope, environmental conditions, and level of severity, and the acceptance logic relationship is used... To express the combination of multiple testing requirements in result judgment, the judgment links of overall acceptance requirements and special acceptance requirements are combined according to hierarchical association. This ensures that special acceptance requirements are only superimposed on the judgment link of overall acceptance requirements when special conditions are triggered. The merged judgment link is then bound to the corresponding testing requirement identifier, thereby generating testing scheme content consistent with the expression of multi-level engineering categories, testing items, and engineering levels. To ensure that the structured arrangement results are verifiable and reusable, the position of each testing requirement in the scheme is represented as a comprehensive topology sequence number. The topology sequence number, the binding relationship between the method and the judgment link are written into the structured field of the testing scheme content. This allows subsequent approval process configuration and testing resource matching to directly reference this structured field without repeatedly parsing the testing scheme content.
[0076] This embodiment also discloses an automatic detection scheme generation device based on multi-stage information extraction. Referring to Figure 2, it includes an acquisition module 201, a processing module 202, and an output module 203. The device is used to execute any of the above-described automatic detection scheme generation methods based on multi-stage information extraction, wherein: the acquisition module 201 is used to receive and parse unstructured raw documents, perform document parsing processing on the content, and convert it into text data; the processing module 202 is used to perform multi-stage information extraction processing based on the text data according to a preset information extraction order to form a structured information set. The multi-stage information extraction processing includes named entity recognition and relation extraction in sequence. The system includes attribute classification; processing module 202, which uses the engineering category, testing items, and engineering level of the structured information set as combination constraints, locates the corresponding standard clauses, testing methods, and acceptance requirements in the standard and specification knowledge base, and automatically assembles and generates the testing plan content while maintaining the semantic consistency of the structured information; output module 203, which, based on the testing plan content and the engineering attribute results determined in the structured information set, determines the automatic configuration of approval nodes and the adaptive matching of testing personnel or testing equipment by associating the engineering level with the approval process rules and the testing item type with the testing resource rules.
[0077] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0078] This embodiment also discloses an electronic device. Referring to FIG3, the electronic device may include: at least one processor 301, at least one communication bus 302, user interface 303, network interface 304, and at least one memory 305.
[0079] The communication bus 302 is used to enable communication between these components.
[0080] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0081] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0082] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0083] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data involved in the various method embodiments described above, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. As a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface 303 module, and an application program for an automatic generation method of a detection scheme based on multi-stage information extraction.
[0084] In the electronic device shown in Figure 3, the user interface 303 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 301 can be used to call an application program stored in the memory 305 that is an automatic generation method for a detection scheme based on multi-stage information extraction. When executed by one or more processors 301, the electronic device executes one or more methods as described in the above embodiments.
[0085] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0086] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0087] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0088] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0089] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0090] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned memory 305 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.
[0091] The present invention also discloses a non-transitory computer-readable storage medium storing instructions. When executed by one or more processors 301, these instructions cause an electronic device to perform one or more methods as described in the above embodiments.
[0092] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truths. This invention is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A method for automatically generating detection schemes based on multi-stage information extraction, characterized in that, The method includes: receiving and parsing unstructured raw documents, performing document parsing processing on the content, and converting it into text data; based on the text data, performing multi-stage information extraction processing according to a preset information extraction order to form a structured information set, wherein the multi-stage information extraction processing sequentially includes named entity recognition, relation extraction, and attribute classification; using the engineering category, testing items, and engineering level of the structured information set as combination constraints, locating the corresponding standard clauses, testing methods, and acceptance requirements in the standard and specification knowledge base, and automatically assembling and generating testing scheme content while maintaining the semantic consistency of the structured information; based on the testing scheme content and the engineering attribute results determined in the structured information set, associating the engineering level with approval process rules and associating the testing item type with testing resource rules to achieve automatic configuration of approval nodes and adaptive matching of testing personnel or testing equipment.
2. The automatic generation method for detection schemes based on multi-stage information extraction according to claim 1, characterized in that, The process involves performing multi-stage information extraction processing based on the text data according to a preset information extraction order to form a structured information set. Specifically, this includes: performing semantic normalization processing on the text data to obtain processed data; performing named entity recognition processing on the processed data according to a preset information extraction order, and annotating the names of engineering objects, engineering categories, inspection items, standard references, and engineering attribute descriptions to form a named entity set; and performing relation extraction processing based on the named entity set, analyzing the co-occurrence positions, semantic orientations, and predefined engineering inspection entities in the text data. The entity association model constructs the hierarchical relationships between project categories and testing items, the reference relationships between testing items and standard clauses, the limiting relationships between project objects and project levels, and the constraint relationships between project attributes and testing requirements to form a relationship network. Based on the relationship network, attribute classification processing is performed on the named entities. By combining the semantic role, association strength, and upstream and downstream relationship types of the named entities in the relationship network, the project category attributes, project level attributes, testing item attributes, standard attributes, and resource requirement attributes are refined, distinguished, and uniformly merged, thereby forming a structured information set containing named entities, entity relationships, and entity attributes.
3. The automatic generation method for detection schemes based on multi-stage information extraction according to claim 1, characterized in that, The process involves using the project category, testing items, and project level of the structured information set as combined constraints, locating the corresponding standard clauses, testing methods, and acceptance requirements in the standard and specification knowledge base, and automatically assembling and generating testing scheme content while maintaining the semantic consistency of the structured information. Specifically, this includes: performing combined constraint construction processing on the project category, testing items, and project level in the structured information set to form combined constraints, wherein the project category is limited to the scope of application of the specification, the testing items are limited to the technical content constraints, and the project level is limited to the stringency constraints of the specification; based on the combined constraints, performing specification location processing in the standard and specification knowledge base, by matching the project category with the specification... The scope of application labels, testing items, and standard clause theme identifiers, as well as engineering level and level limitation clauses, are matched for consistency to obtain a set of standard clauses that simultaneously satisfy the combined constraints. Association parsing is performed on the set of standard clauses to extract the corresponding testing methods and acceptance requirements. Semantic consistency verification is performed on the standard clauses, testing methods, and acceptance requirements based on the structured information set to form a set of standard elements. Automatic assembly is then performed on the set of standard elements to generate the testing plan content, wherein the standard clauses are used as the standard basis, the testing methods as the testing implementation content, and the acceptance requirements as the result judgment content, all in a unified arrangement.
4. The automatic generation method for detection schemes based on multi-stage information extraction according to claim 1, characterized in that, The process, based on the testing plan content and the engineering attribute results determined in the structured information set, associates the engineering level with approval process rules and the testing item type with testing resource rules to achieve automatic configuration of approval nodes and adaptive matching of testing personnel or equipment. Specifically, this includes: performing engineering attribute mapping processing on the testing plan content based on the engineering attribute results; semantically aligning the testing items, testing methods, and acceptance requirements in the testing plan content with the engineering category, engineering level, and testing item attributes in the structured information set, so that each testing unit in the testing plan content is bound to a corresponding engineering attribute identifier; performing approval process rule association determination processing based on the engineering level; matching the engineering level with the level range or level label in the preset approval process rules; and combining the risk attributes and mandatory attribute of the testing items in the testing plan content to automatically configure the approval node and adaptively match the testing personnel or equipment. The number of approval levels, approval node types, and approval order constraints in the process rules are jointly determined to generate an approval node configuration result that matches the project level. Based on the testing item types determined in the testing plan and the testing item attributes in the structured information set, association determination processing is performed on the testing resource rules. This involves matching the testing item types with the predefined personnel qualification conditions, equipment capability parameters, and resource occupancy constraints in the testing resource rules, while simultaneously verifying the project level's limitations on resource level or accuracy level, to obtain a set of candidate testing personnel or a set of candidate testing equipment. Adaptive filtering processing is then performed on the set of candidate testing personnel or the set of candidate testing equipment. By combining the requirements in the testing plan, the qualification level, historical execution records, and current availability of the candidate resources are comprehensively determined to obtain a testing personnel or testing equipment configuration result that is consistent with the testing plan content and the project attribute results.
5. The automatic generation method for detection schemes based on multi-stage information extraction according to claim 1, characterized in that, The process of performing multi-stage information extraction processing based on the text data and according to a preset information extraction order specifically includes: the attribute classification includes performing multi-level engineering category determination processing on the engineering category when the same text data simultaneously meets the determination conditions of the main engineering category and the special engineering category, so as to simultaneously identify the main engineering category and the special engineering category; establishing a hierarchical association relationship between the main engineering category and the special engineering category, so that the engineering category forms a multi-level engineering category expression; and forming a multi-level structured information set together with the detection item and the engineering level.
6. The automatic generation method for detection schemes based on multi-stage information extraction according to claim 5, characterized in that, The establishment of a hierarchical relationship between the main project category and the special project category, enabling the project categories to form a multi-level project category expression, specifically includes: performing hierarchical role determination processing on each project category in the candidate set of project categories based on a preset hierarchical rule set of project categories to determine the hierarchical role of the main project category or special project category corresponding to each project category, and generating a corresponding hierarchical role identifier for each project category; based on the hierarchical role identifier, combined with the semantic relationship between project-related entities in the text data, performing hierarchical association matching processing on the main project category and the special project category, and establishing a hierarchical relationship between the main project category and the special project category when the preset hierarchical association conditions are met by comprehensively judging the scope of the project object, the functional orientation, and the consistency of the detection object; performing hierarchical structured encapsulation processing on the main project category and the special project category, taking the main project category as the upper-level project category node, taking the special project category with the established hierarchical relationship as the lower-level project category node, and recording the membership identifier, hierarchical depth identifier, and association constraint identifier between project categories in the multi-level structured information set, so that the project categories form the multi-level project category expression.
7. The automatic generation method for detection schemes based on multi-stage information extraction according to claim 6, characterized in that, The process of using the engineering category, testing item, and engineering grade of the structured information set as combined constraints to locate the corresponding standard clauses, testing methods, and acceptance requirements in the standard and specification knowledge base, and automatically assembling and generating testing scheme content while maintaining the semantic consistency of the structured information, specifically includes: based on the multi-level structured information set, performing hierarchical expansion processing on the multi-level engineering category expressions therein, constructing an engineering category hierarchical tree structure by parsing the hierarchical relationship, the engineering category hierarchical tree structure containing the upper-level engineering category nodes and at least one layer of the lower-level engineering category nodes; based on the engineering category hierarchical tree structure, using the engineering category, the testing item, and the engineering grade as combined constraints, performing hierarchical retrieval processing on the standard and specification knowledge base, wherein for the upper-level engineering category nodes, basic standard clauses, general testing methods, and overall acceptance requirements are located, and for the lower-level engineering category nodes, specific testing clauses, supplementary testing methods, and specific acceptance requirements are located; and according to the hierarchical relationship, the main engineering category is corresponding to The testing requirements and the testing requirements corresponding to the specific project categories are subjected to hierarchical constraint combination processing. This involves using the testing requirements corresponding to the main project category as the higher-level constraint and the testing requirements corresponding to the specific project category as the lower-level supplementary constraint. Consistency constraints are then applied to the testing frequency, testing depth, and acceptance criteria based on the project level to form a set of testing requirements. This set of testing requirements undergoes semantic consistency verification processing. This involves matching the set of testing requirements with the scope of project objects, semantics of testing items, and project level limitations in the multi-level structured information set item by item, eliminating testing requirements inconsistent with project attributes, and unifying and merging semantically overlapping testing requirements. After semantic consistency verification, based on the hierarchically constrained and semantically consistent set of testing requirements, the testing requirements are structured and arranged according to the testing order, testing method adaptation relationship, and acceptance logic relationship, thereby generating testing scheme content consistent with the multi-level project category expression, the testing items, and the project level.
8. An automatic detection scheme generation device based on multi-stage information extraction, characterized in that, The device is used to execute an automatic generation method for detection schemes based on multi-stage information extraction as described in any one of claims 1-7. The device includes an acquisition module, a processing module, and an output module, wherein: the acquisition module is used to receive and parse unstructured raw documents, perform document parsing processing on the content, and convert it into text data; the processing module is used to perform multi-stage information extraction processing based on the text data according to a preset information extraction order to form a structured information set, wherein the multi-stage information extraction processing includes named entity recognition, relation extraction, and attribute classification in sequence; the processing module is used to use the engineering category, detection item, and engineering level of the structured information set as combination constraints, locate the corresponding standard clauses, detection methods, and acceptance requirements in the standard specification knowledge base, and automatically assemble and generate detection scheme content while maintaining the semantic consistency of the structured information; the output module is used to, based on the detection scheme content and the engineering attribute results determined in the structured information set, determine the automatic configuration of approval nodes and the adaptive matching of detection personnel or detection equipment by associating the engineering level with approval process rules and the detection item type with detection resource rules.
9. An electronic device, characterized in that, The device includes a processor, a communication bus, a user interface, a network interface, and a memory. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The communication bus is used to enable communication between the components within the electronic device. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.